Part 9 of 10 5 min read

Respect the testing pyramid

More E2E Tests, Fewer Unit Tests

The testing pyramid was built for humans maintaining flaky browser tests. Agents flip it: test the flows your users actually care about.

You know the testing pyramid. Lots of fast unit tests at the bottom, some integration tests in the middle, a tiny, reluctant sliver of browser tests at the top. I drew it on whiteboards. I said "system specs are the last resort" with a straight face.

Dijkstra said the tools we use have a profound (and devious!) influence on our thinking habits. I used to quote that about programming languages. It applies just as well to the pyramid. It shaped how a whole generation thinks about testing, and it was built around one scarce resource: human time.

Human time isn't the bottleneck anymore.

Sixty green specs and a broken checkout

Here's what happens when you tell an agent "add tests for this."

I had it build a checkout flow. It made a CheckoutService, a PriceCalculator, a CouponApplier, and a confirmation mailer. Then it wrote sixty specs like this:

ruby
RSpec.describe CouponApplier do
  it "applies the discount" do
    cart = instance_double(Cart, subtotal: 100)
    coupon = instance_double(Coupon, percent_off: 20)
    expect(described_class.new(cart, coupon).total).to eq(80)
  end
end

RSpec.describe CheckoutService do
  it "creates a session with the applied total" do
    applier = instance_double(CouponApplier, total: 80)
    allow(CouponApplier).to receive(:new).and_return(applier)
    expect(StripeGateway).to receive(:create_session).with(amount: 8000)
    described_class.new(cart, "LAUNCH20").call
  end
end

All green.

Checkout was broken. The form posted the coupon as coupon_code, the controller permitted :code, so the coupon silently never applied and Stripe charged full price. None of the sixty specs could catch it, because none of them touched a form, a controller, or a browser. Each one tested a little island that worked perfectly on its own. The mocks guaranteed it: the CheckoutService spec literally stubbed in the $80 it was supposed to be verifying.

That's not a test. That's a tautology. The code does what the code does.

The one spec that catches it:

ruby
RSpec.describe "Checkout", type: :system do
  it "applies a coupon and charges the discounted price" do
    sign_in users(:avi)
    visit product_path(products(:course))

    click_on "Buy now"
    fill_in "Coupon code", with: "LAUNCH20"
    click_on "Apply"

    expect(page).to have_content("$80.00")
  end
end

Capybara, Cuprite driving headless Chrome, fixtures for the data. It reads like the thing a user does because it is the thing a user does. It fails on the $80.00 line with the page showing $100.00, and the agent traces that to the strong params in about a minute.

Side by side:

60 mocked unit specs1 system spec
Caught the coupon bugNoYes
Runtime~2 seconds~4 seconds
Survives moving CouponApplier into CartNo, a dozen breakYes, untouched
Tells you where the bug isPrecisely, for bugs it can seeRoughly, the agent narrows it down
Written byThe agent, to match its own codeThe flow a user actually walks through

The unit specs win one row, and it's a real one. When they fail, they point at the exact line. But they only fail on bugs inside the islands, and the expensive bugs live between them.

The pyramid's assumptions, then and now

The pyramid wasn't wrong. It was a budget, and the currency was developer afternoons.

AssumptionThenNow
Who fixes a flaky or brittle testA developer, losing an afternoonThe agent, in the background
Cost of a slow suiteA person staring at a progress bar before pushingAgent time in a parallel worktree nobody is watching
Button renamed, eleven specs go redA morning of find and replaceThe agent reads the failure, updates the selector, reruns
What a green unit suite tells youThe code you wrote does what you meantThe code the agent wrote does what the agent wrote
What a failing system spec tells youProbably flakiness, hit rerunA user can't do the thing

That fourth row is the one that changed my mind. When a human writes a unit test, there's a gap between intent and implementation, and the test sits in that gap. When an agent writes the code and the tests in the same breath, there is no gap. The tests restate the implementation.

As for speed: a six minute system suite instead of forty seconds is six minutes of an agent's time, not mine. I've got several workstreams running at once, each running its own specs whenever it wants. I'll take that trade every day for a suite that catches broken checkouts.

Tests are the contract now

In Stop Reading Code I said I don't read most of the code agents write anymore. The obvious follow up: then how do you know it works?

This is how. The e2e suite is the contract.

Unit tests are coupled to the implementation. When the agent decides CouponApplier should really be a method on Cart, the unit specs break and tell me nothing about whether the product works. Worse, the agent "fixes" them by rewriting them to match the new code, and now they prove nothing again.

System specs don't care how the code is shaped. They care whether a person can sign up, buy the course, and get the email. So I can let an agent restructure everything, or throw it out and rewrite it, and the same specs tell me whether we still have a product at the end.

Where unit tests still earn their keep

Unit tests aren't dead. Some code is pure logic with a big input space, and driving a browser through fifteen edge cases is the wrong tool.

Money math is the obvious one. Proration, tax rounding, splitting a payment across installments where the pennies have to add up. I want a table of inputs and expected outputs running in milliseconds, with no browser in the way. Same for parsers: a markdown converter, a CSV importer, anything where the interesting cases are weird strings. Same for anything genuinely algorithmic, like a scheduling rule or a permissions matrix.

If the logic is pure and the edge cases are the point, unit test it, and don't mock anything. If the logic is wiring (forms, params, controllers, jobs, mailers talking to each other), a unit test can only confirm the wiring you already believe in.

What I tell the agents

I don't leave this to chance, because the default is the pile of mocked specs. My project instructions say it outright: essential coverage only. A system spec for every user facing flow, first. Request specs for APIs and anything without a UI. Model specs only for real business logic, validations, and scopes that matter. Unit tests for pure logic like money and parsing. No view specs, no helper specs, no testing private methods, no mocking your own code.

The suite is smaller and it tells me more. When it's green, checkout works.

This is part of Rethink Everything. Next up: optimize for the agent.

The whole series

  1. 1Code is written for humans to readStop Writing and Reading Code
  2. 2Never rewrite from scratchRewrites Might Be Better Than Massive Refactors
  3. 3Use the language your team already knowsLanguage Strengths Over Language Familiarity
  4. 4Build it once for the web and wrap itGo Native Everywhere
  5. 5Keep the monolith majesticI Hated Microservices. Agents Love Them.
  6. 6Pay someone else to run your serversOwn Your Infrastructure
  7. 7Don't repeat yourselfDuplicate Code So Your Agents Stop Colliding
  8. 8Types are ceremonyTypeScript, I Don't Hate You Anymore
  9. 9Respect the testing pyramidMore E2E Tests, Fewer Unit Tests
  10. 10Refactor for readabilityStop Refactoring for Humans. Refactor for the Agent.