Respect the testing pyramid
More E2E Tests, Fewer Unit Tests
The testing pyramid was built for humans maintaining flaky browser tests. Agents flip it: test the flows your users actually care about.
You know the testing pyramid. Lots of fast unit tests at the bottom, some integration tests in the middle, a tiny, reluctant sliver of browser tests at the top. I drew it on whiteboards. I said "system specs are the last resort" with a straight face.
Dijkstra said the tools we use have a profound (and devious!) influence on our thinking habits. I used to quote that about programming languages. It applies just as well to the pyramid. It shaped how a whole generation thinks about testing, and it was built around one scarce resource: human time.
Human time isn't the bottleneck anymore.
Sixty green specs and a broken checkout
Here's what happens when you tell an agent "add tests for this."
I had it build a checkout flow. It made a CheckoutService, a PriceCalculator, a CouponApplier, and a confirmation mailer. Then it wrote sixty specs like this:
RSpec.describe CouponApplier do
it "applies the discount" do
cart = instance_double(Cart, subtotal: 100)
coupon = instance_double(Coupon, percent_off: 20)
expect(described_class.new(cart, coupon).total).to eq(80)
end
end
RSpec.describe CheckoutService do
it "creates a session with the applied total" do
applier = instance_double(CouponApplier, total: 80)
allow(CouponApplier).to receive(:new).and_return(applier)
expect(StripeGateway).to receive(:create_session).with(amount: 8000)
described_class.new(cart, "LAUNCH20").call
end
endAll green.
Checkout was broken. The form posted the coupon as coupon_code, the controller permitted :code, so the coupon silently never applied and Stripe charged full price. None of the sixty specs could catch it, because none of them touched a form, a controller, or a browser. Each one tested a little island that worked perfectly on its own. The mocks guaranteed it: the CheckoutService spec literally stubbed in the $80 it was supposed to be verifying.
That's not a test. That's a tautology. The code does what the code does.
The one spec that catches it:
RSpec.describe "Checkout", type: :system do
it "applies a coupon and charges the discounted price" do
sign_in users(:avi)
visit product_path(products(:course))
click_on "Buy now"
fill_in "Coupon code", with: "LAUNCH20"
click_on "Apply"
expect(page).to have_content("$80.00")
end
endCapybara, Cuprite driving headless Chrome, fixtures for the data. It reads like the thing a user does because it is the thing a user does. It fails on the $80.00 line with the page showing $100.00, and the agent traces that to the strong params in about a minute.
Side by side:
| 60 mocked unit specs | 1 system spec | |
|---|---|---|
| Caught the coupon bug | No | Yes |
| Runtime | ~2 seconds | ~4 seconds |
Survives moving CouponApplier into Cart | No, a dozen break | Yes, untouched |
| Tells you where the bug is | Precisely, for bugs it can see | Roughly, the agent narrows it down |
| Written by | The agent, to match its own code | The flow a user actually walks through |
The unit specs win one row, and it's a real one. When they fail, they point at the exact line. But they only fail on bugs inside the islands, and the expensive bugs live between them.
The pyramid's assumptions, then and now
The pyramid wasn't wrong. It was a budget, and the currency was developer afternoons.
| Assumption | Then | Now |
|---|---|---|
| Who fixes a flaky or brittle test | A developer, losing an afternoon | The agent, in the background |
| Cost of a slow suite | A person staring at a progress bar before pushing | Agent time in a parallel worktree nobody is watching |
| Button renamed, eleven specs go red | A morning of find and replace | The agent reads the failure, updates the selector, reruns |
| What a green unit suite tells you | The code you wrote does what you meant | The code the agent wrote does what the agent wrote |
| What a failing system spec tells you | Probably flakiness, hit rerun | A user can't do the thing |
That fourth row is the one that changed my mind. When a human writes a unit test, there's a gap between intent and implementation, and the test sits in that gap. When an agent writes the code and the tests in the same breath, there is no gap. The tests restate the implementation.
As for speed: a six minute system suite instead of forty seconds is six minutes of an agent's time, not mine. I've got several workstreams running at once, each running its own specs whenever it wants. I'll take that trade every day for a suite that catches broken checkouts.
Tests are the contract now
In Stop Reading Code I said I don't read most of the code agents write anymore. The obvious follow up: then how do you know it works?
This is how. The e2e suite is the contract.
Unit tests are coupled to the implementation. When the agent decides CouponApplier should really be a method on Cart, the unit specs break and tell me nothing about whether the product works. Worse, the agent "fixes" them by rewriting them to match the new code, and now they prove nothing again.
System specs don't care how the code is shaped. They care whether a person can sign up, buy the course, and get the email. So I can let an agent restructure everything, or throw it out and rewrite it, and the same specs tell me whether we still have a product at the end.
Where unit tests still earn their keep
Unit tests aren't dead. Some code is pure logic with a big input space, and driving a browser through fifteen edge cases is the wrong tool.
Money math is the obvious one. Proration, tax rounding, splitting a payment across installments where the pennies have to add up. I want a table of inputs and expected outputs running in milliseconds, with no browser in the way. Same for parsers: a markdown converter, a CSV importer, anything where the interesting cases are weird strings. Same for anything genuinely algorithmic, like a scheduling rule or a permissions matrix.
If the logic is pure and the edge cases are the point, unit test it, and don't mock anything. If the logic is wiring (forms, params, controllers, jobs, mailers talking to each other), a unit test can only confirm the wiring you already believe in.
What I tell the agents
I don't leave this to chance, because the default is the pile of mocked specs. My project instructions say it outright: essential coverage only. A system spec for every user facing flow, first. Request specs for APIs and anything without a UI. Model specs only for real business logic, validations, and scopes that matter. Unit tests for pure logic like money and parsing. No view specs, no helper specs, no testing private methods, no mocking your own code.
The suite is smaller and it tells me more. When it's green, checkout works.
This is part of Rethink Everything. Next up: optimize for the agent.