Most teams know they should refactor. They also know they should floss. The problem isn't awareness — it's justification. Refactors compete for the same calendar as features, and features have advocates. Refactors have guilt.
That framing is the real problem. A refactor isn't penance for past sins. It's a capital investment. And like any investment, it deserves a thesis: what return do you expect, by when, and how will you measure it?
We delayed a significant infrastructure refactor for nearly a year. Then we did it, and it paid for itself in six weeks. Here's what changed our minds, and what we'd tell teams stuck in the same loop.
We didn't wake up one morning and decide the codebase was ugly. We noticed three numbers moving in the wrong direction at the same time.
Deploy frequency was dropping. Not because we shipped less code — engineers were writing plenty. But the time from "merge" to "running in production" kept stretching. What used to take minutes started taking the better part of an afternoon. People began batching deploys to avoid the pain, which made each deploy riskier, which made people batch more. A vicious cycle with a legible trend line.
Incident volume was climbing on a per-deploy basis. We weren't just deploying less often — each deploy was more likely to cause trouble. The ratio mattered more than the raw count. It told us the system had become harder to change safely, not that we'd gotten worse at our jobs.
Onboarding time for new engineers was stretching past a month. When a competent hire needs four-plus weeks before they can ship a small change with confidence, the system is carrying knowledge debt. Documentation can't fix architecture that requires oral tradition to operate.
Any one of these signals alone is easy to rationalize. Together, they told a clear story: the cost of not refactoring was compounding.
The decision to refactor became simple once we forced ourselves to write down what "done" looked like in outcome terms. Not "cleaner code" or "better separation of concerns" — those describe the work, not the payoff.
We committed to three targets:
Each target was measurable. Each mapped to a cost the business already felt. Each had a deadline: six weeks after we resumed normal feature work.
This is the step most refactor proposals skip. They describe the mess. They don't quantify the cleanup's value. That's why leadership says "not now" — not because they don't believe the mess exists, but because they can't compare the cost of fixing it to the cost of the next feature.
We carved out three weeks of focused work. No feature flags, no split attention. Three weeks sounds expensive. It is — until you calculate what you were already spending.
The key discipline was scope. We wrote a list of what was in bounds and, more importantly, a longer list of what was out. Every refactor wants to become a rewrite. Rewrites rarely finish. We treated scope creep as the primary risk, not timeline.
Daily check-ins were short: what moved, what's blocked, are we still inside the boundary we drew? When someone found an adjacent improvement that was tempting but out of scope, we logged it and moved on. Some of those items are still on the list. That's fine.
Six weeks after we returned to normal development:
The system got out of the way.
Not every refactor clears this bar. If you can't name the return in terms the business cares about — deploy speed, incident cost, hiring velocity, customer-facing reliability — you might be proposing a preference, not an investment. Preferences are fine. They don't earn a protected window on the roadmap.
The honest test: would you fund this work if it were a product feature with the same expected return? If yes, do it. If you're not sure, get sharper on the numbers until the answer is clear.
Tax is money you pay because you must. Investment is money you spend because you expect more back. The difference is whether you've done the math.
We spent three weeks. We got faster deploys, fewer incidents, and an onboarding experience that stopped embarrassing us. The numbers justified the work within six weeks. The compounding benefits — engineers moving faster every week since — are still accumulating.
Name the return before you start. Protect the window. Measure the result. That's it.
Be the first to comment.
0 comments
Loading comments...