Part 2 of 10 5 min read

Never rewrite from scratch

Rewrites Might Be Better Than Massive Refactors

The old codebase used to be a liability you couldn't afford to throw away, and now it's the spec an agent ports from.

I used to believe you should never rewrite an app.

That came from Joel Spolsky's "Things You Should Never Do," written in 2000. Netscape threw away its code, spent years rebuilding it, and handed the market to Internet Explorer. His argument was simple and correct: old code is ugly because it's full of bug fixes. Every weird conditional is a lesson somebody paid for. Throw it away and you throw away the lessons.

Whenever someone pitched a rewrite, I sent them that essay. Refactor in place. Small steps. Keep the tests green. Strangle the old thing slowly.

Now I rewrite things all the time, and it's easier than refactoring them.

The premise changed

Joel's argument rested on two costs. Rewrites are expensive, because humans have to type the whole thing again. And rewrites lose knowledge, because the humans doing the typing don't know why that weird conditional exists.

Both costs assumed a human was doing the rewrite.

An agent doesn't lose the knowledge. It reads it. Every edge case, every odd validation, every "if the user signed up before 2019 do this" branch is sitting in the repo for the agent to plan against. The lessons don't disappear. They get ported.

The old code is the spec. It's a better spec than anything a product manager ever handed me, because it's the one that's been running in production.

And typing, the expensive part, is basically free now.

Same job, two ways

Take an aging Rails 5 app. jQuery everywhere, Turbolinks, a dozen gems that stopped getting updates in 2019, a Sprockets pipeline nobody wants to touch. The goal is a modern Rails 8 app with Hotwire and Tailwind that does exactly what the old one does.

The refactor I used to run. Upgrade Rails one minor version at a time, fixing deprecations at each step. Swap dead gems for live ones. Put the new UI behind feature flags and migrate page by page, strangler style, so the old path and the new path both ship to production. For months, half the app is Stimulus and half is $(document).on('turbolinks:load'). Every flag is a branch you test both sides of.

That hybrid is also the worst possible context for an agent. Every file is a negotiation between two conventions, plus whatever half-finished migration the last session left behind.

Ruby made this worse, honestly. An essay I wrote years ago about why we taught Ruby first celebrated Matz giving us many ways to do the same thing because "people are different." Ten years of different people later, a legacy Rails app has five ways to do the same thing, and a refactor has to preserve all five.

The rewrite I run now. Before a line of the new app exists, the agent writes end to end tests against the old one. Not unit tests: those are coupled to the internals you're throwing away. Browser tests. Visit this page, click this, fill in that, expect to see this. Behavior, not structure. (More on why I lean on these now in part 9.)

Once that suite is green against the old app, the agent builds the new one from scratch with one set of conventions, and the same suite runs against it. The port is done when it passes.

markdown
## Phase 1: Capture behavior (old app)
- Crawl every route, record URL, status, key content
- System specs for signup, checkout, admin reports, emails sent
- Run against the Rails 5 app. All green before Phase 2.

## Phase 2: Port
- rails new, Rails 8, Hotwire, Tailwind
- Same routes, same URLs, same emails
- Point the Phase 1 suite at the new app

## Done when
- Phase 1 suite passes against the new app with zero spec edits
bash
/rg:plan rewrite legacy app on Rails 8, e2e specs against the old app first
/rg:work plans/feature-rewrite-legacy-app.md
bin/agent-rspec run spec/system/legacy_parity

That last line is the whole game. The agent keeps going until parity is green. I read test results, not the port. If a spec fails, either the port is wrong or I just learned something about the old app I didn't know. I did exactly this moving avi.nyc from Jekyll to Rails, with both repos open side by side.

Incremental refactorRewrite against the old app
TimelineMonths. Every step ships to production.Days to a couple of weeks for an app this size.
RiskSpread thin across dozens of deploys, each one small but each one live.Concentrated in one cutover, but measured by a suite that ran against production behavior first.
Codebase along the wayTwo versions of everything alive at once, flags everywhere.Old app untouched, new app clean. Never a hybrid.
What you learnHow the old code is structured.What the old app actually does, written down as tests you keep.
What can go wrongStalls halfway and the hybrid becomes permanent.Behavior the suite didn't capture, data that won't import, integrations nobody knew about.

The refactor's real advantage is the risk column. You never have a big scary day. If your team can't tolerate a cutover, or the app is making money every minute and you can't run both side by side for a week, incremental still wins.

But the most common way I've seen refactors fail is that they never finish. The flags stay. The hybrid becomes the architecture. A rewrite either passes the suite or it doesn't.

Where rewrites still bite

Data is the hard part. Code you can regenerate. Production data you can't. Ten years of unexpected nulls and two columns that mean the same thing depending on the row's age are now the new app's problem. Write the migration scripts early, run them against a real production snapshot, and put assertions on the output. A green suite against seed data means nothing if the real data doesn't import.

Undocumented integrations are the other trap. The webhook a partner hits that nobody remembers. The cron job emailing finance a CSV every Monday. The API client that depends on an exact response shape. None of that shows up when you crawl the UI. Grep the logs, check inbound traffic, ask whoever's been around longest.

And if the app is huge, with years of business rules across dozens of domains, don't port it all at once. Carve it into pieces and rewrite one at a time. That's still a rewrite. It's just not a Netscape.

Kill the nostalgia

Joel was right for his era, and for most of mine. I'd still send that essay to anyone planning to rewrite by hand.

But when I look at a crusty app now, I don't ask how to refactor it safely. I ask what the tests would look like against its current behavior. Once those exist, the rewrite is the easy part.

This is part of a series where I'm questioning everything I used to believe about building software. Start at Rethink Everything. Next up: why I pick languages for their strengths now, not for what I already know, in part 3.

The whole series

  1. 1Code is written for humans to readStop Writing and Reading Code
  2. 2Never rewrite from scratchRewrites Might Be Better Than Massive Refactors
  3. 3Use the language your team already knowsLanguage Strengths Over Language Familiarity
  4. 4Build it once for the web and wrap itGo Native Everywhere
  5. 5Keep the monolith majesticI Hated Microservices. Agents Love Them.
  6. 6Pay someone else to run your serversOwn Your Infrastructure
  7. 7Don't repeat yourselfDuplicate Code So Your Agents Stop Colliding
  8. 8Types are ceremonyTypeScript, I Don't Hate You Anymore
  9. 9Respect the testing pyramidMore E2E Tests, Fewer Unit Tests
  10. 10Refactor for readabilityStop Refactoring for Humans. Refactor for the Agent.