Review the plan, not the diff: a checkpoint for coding agents
The cheapest place to catch a bad change from a coding agent is the plan, not the pull request. Make the agent write a short plan first, have a person or a separate reviewer agent approve it against a fixed checklist, and only then let it write code. A wrong approach costs a few minutes to reject at plan stage and hours of review, rework and token spend once it is a 1,500-line diff.
This idea is getting attention again. This week Cockroach Labs wrote up five months of running coding agents as a hospital-style team, and one of their stated lessons was that reviewing the plan caught the most expensive problems. Below is how to set up plan review for your own team, what a good plan contains, what reviewers should check, and how to tell whether the checkpoint is paying for itself.
Why the diff is the wrong place to start reviewing
Code review was designed for human authors who already shared context with the reviewer. Coding agents break that assumption in three ways.
Volume outruns attention
An agent can produce in ten minutes what a reviewer needs an hour to read carefully. When every change arrives as a large diff, reviewers skim, and skimming is how a wrong approach gets merged with clean-looking code around it.
The expensive mistakes are design mistakes
Typos and missing null checks are cheap. The costly errors are choosing the wrong layer to change, duplicating logic that already exists, changing a public interface without noticing who calls it, or "fixing" a failing test by loosening the assertion. All of these are visible in a plan before any code exists.
Rework burns tokens twice
If a reviewer rejects the approach after the code is written, the agent throws away the work and starts again. On agentic tasks, most of the cost is in exploration and iteration, so catching the wrong direction early saves the bulk of the spend. We covered how to see that spend per engineer in budgeting AI coding agent spend.
What a reviewable plan looks like
A plan is only useful if a reviewer can approve or reject it in five minutes. Ask the agent for a fixed template rather than free text. A template that works well for SaaS codebases:
- The problem in one or two sentences, with a link to the issue and the reproduction steps the agent actually ran.
- The root cause as the agent understands it, plus the file and function where it lives.
- The proposed change: which files will be touched and why, in a short list.
- What will not change, especially public APIs, database schemas and shared config.
- The tests: which existing tests cover this, which new tests will be added, and an explicit statement that no existing test will be deleted or weakened.
- Risk and rollout: anything that touches data, auth, billing or migrations, and how the change can be reverted.
- Estimated size of the diff, with a plan to split it if it will exceed your size limit.
- Open questions the agent could not answer from the code.
The open-questions section is the most valuable line in the template. An agent that lists nothing there on a non-trivial task is usually guessing.
A plan review checklist
Reviewers should check the plan against a short, fixed list. Fixed criteria make reviews faster and make it possible to measure whether reviewers agree with each other.
Is the diagnosis right?
Does the root cause match the reproduction? A plan that fixes a symptom, such as adding a retry around a call that fails because of bad input, should be rejected even if the code would work.
Is this the right place to change?
Check whether a helper, service or pattern for this already exists in the codebase. Agents often write a new utility instead of finding the old one. Check whether the change belongs in the caller or the callee.
Does it protect the tests?
Any plan that edits an existing assertion needs a stated reason a human agrees with. "The test was wrong" is sometimes true, and it is also the most common way an agent makes a failure disappear without fixing it.
Is the blast radius stated and acceptable?
Changes to schemas, auth, payments, data deletion or shared configuration should be flagged and sent to a person with the authority to approve them. For a SaaS product, a short list of protected paths that always require human sign-off is worth writing down once.
Is it small enough to review?
Set a size ceiling for one change and make the agent split work that exceeds it. The exact number matters less than having one. Cockroach Labs described a limit of around 1,000 lines; many teams pick a lower figure for application code.
Who reviews the plan: people, agents, or both
A practical split for most teams:
- Low-risk plans that touch no protected paths and change fewer than a few hundred lines can be reviewed by a separate reviewer agent working from the checklist, with a person spot-checking a sample each week.
- Plans that touch protected paths, change public interfaces or modify tests always go to a person.
- Any plan where the reviewer agent and the authoring agent disagree twice in a row goes to a person instead of looping. Endless review cycles are a real failure mode and need a circuit breaker.
The reviewer agent should be configured differently from the author: a different prompt, ideally a different model, and no access to the author's reasoning beyond the plan itself. A reviewer that sees the author's chain of thought tends to agree with it. If you use a model as the reviewer, validate it the way you would any automated grader, as described in our post on using an LLM judge as a release gate.
Measuring whether plan review pays for itself
Plan review adds latency to every task, so measure it rather than assume it helps. Track four numbers for a month:
- Plan rejection rate, and the reason for each rejection mapped to the checklist item that caught it.
- Time from plan submitted to plan approved. If this is longer than the coding step, the process has become bureaucratic.
- Post-merge outcomes: reverts, hotfixes and bugs traced back to agent changes, split by whether the plan was human-reviewed or agent-reviewed.
- Rework after code review: how often an approved plan still produced a pull request that needed a different approach.
A healthy pattern is a modest rejection rate concentrated in the diagnosis and right-place checks, short approval times, and very few approved plans that later needed a new approach. If almost nothing is ever rejected, either the agent is excellent on your codebase or reviewers are rubber-stamping. Sample twenty approved plans and re-review them blind to find out which.
The same scoring discipline that applies to code applies to plans. Our rubric for grading AI-generated code can be cut down into a plan rubric with anchored pass and fail examples, and the reviewer-agreement checks in measuring inter-rater agreement for LLM evals tell you whether two reviewers apply it the same way.
Common ways plan review goes wrong
- Plans that are just the code in prose. If the plan is longer than the diff will be, the template is too detailed. Ask for decisions, not line-by-line steps.
- Approving plans for code you have not read in months. Reviewers need enough codebase knowledge to judge the right place for a change. Route plans to the owners of the area.
- No precedent log. When a reviewer rejects an approach for a reason that will recur, record it in a short file the agent reads on every task, so the same plan is not proposed again next week.
- Prompt bloat. Every incident tempts the team to add another rule to the agent prompt. Audit the prompt and the precedent log regularly and remove rules that no longer fire.
For teams that do not have senior reviewers with time to spare, independent expert review of plans and patches is one of the things Boundev's expert AI evaluation network provides. Our comparison of AI coding tools for engineering teams covers where these agents fit in the wider toolchain.
FAQ
Does plan review slow down coding agents too much?
It adds a step, but a plan takes minutes to read while a rejected pull request costs a full rework cycle. Measure approval time against coding time; if approval regularly takes longer, simplify the template or let a reviewer agent handle low-risk plans.
Can an AI agent review another agent's plan?
Yes, for low-risk changes, if the reviewer uses a different prompt and ideally a different model, sees only the plan and not the author's reasoning, and is spot-checked by people. Protected areas such as auth, billing, schemas and test changes should still go to a person.
What should a coding agent's plan include?
The problem and reproduction, the root cause and location, the files it will change, what it will not change, the tests it will add or rely on, risks and rollback, the expected diff size, and open questions it could not resolve from the code.
How big should a single agent change be?
Small enough that a reviewer can read it carefully in one sitting. Pick a line limit that fits your codebase, enforce it in the plan stage, and have the agent split larger work into a sequence of plans.
Put expert judgment to work.
Talk to our team about evaluating AI-generated code, comparing responses, and building clear rubrics.