Most advice about coding with AI assumes one model, one chat window and one person typing. That is a pair-programming session. It does not scale past the first afternoon.
What has worked for us since February 2026 is a routing rule, written down and enforced by the tooling.
The rule
A planning model does the higher-level work: the specification, the architecture decisions, the brief for each task, the review of every diff before it is committed, the diagnosis when a fix fails.
Coding models do every code change: the tests, the refactors, the bug fixes, written from the lead’s brief, one task at a time, in a fresh context.
Research models do the grep-and-report work and, more usefully, the refuting: given a claim, find the evidence against it.
A failed fix never goes back to the implementer for another guess. It goes to the lead for a diagnosis. That single rule removes most of the loops people complain about.
What makes it hold
A brief is a contract: the files, the success criterion, the gates to run. The implementer sees the brief and the files it names, nothing else.
The gates are not optional. Unit tests run before every push; the current project has more than 8,000 of them in 640-odd files. Static analysis runs with a baseline that may only shrink. A visual audit compares the rendered page with the design tokens it was supposed to use.
Memory is verified before it is trusted. A session starts by checking its remembered claims against the git history and the deploy log, and every plan’s file paths and line numbers are checked before the plan runs. A model that trusts a stale note is worse than one with no notes.
What it cost, what it bought
Roughly 3,500 commits and 290 design documents in eight months, by one person. The design documents are not overhead. They are what the planning model reads before it briefs, and what the next session reads before it continues.
The same method rebuilt this website from a page builder to a native block theme: an inventory, a pixel baseline, a plan per phase, a coding model per task, a lead that reads every diff.
Why the split holds
The roles are different jobs, and that is why it works, not because one model is smarter than another. A planner that also codes stops reviewing, and a coder that also plans starts guessing. Keeping them apart is a management decision. It happens to be the one that matters.
