How to review code an agent wrote
24 September 2026 | 7 min
The bottleneck moved from writing code to deciding whether code is correct, and most teams are still reviewing the old way.
When an agent writes a large share of the changes, the constraint stops being how fast anyone types and becomes how fast anyone can be confident. Teams that do not notice end up with a review queue instead of a backlog, and then with a review queue nobody reads properly.
Read the claim before the code
Start with what the agent said it did, and check that against what it ran. If there is no record of anything being executed, treat the summary as an intention rather than a result. This is thirty seconds and it reorders everything after it.
Read the edges, not the middle
Agent code is usually locally plausible. The defects cluster where the change meets code the agent did not write:
- Callers. Every place the changed thing is used, especially in files the agent never opened.
- Error paths. The happy path is nearly always right. The branch taken when something fails is where invented behaviour lives.
- Deletions. An agent removing code to make a test pass is the single highest-risk edit in any diff.
- Tests it added. A test written alongside the code it tests can be shaped to pass. Break the behaviour and confirm the test fails.
What to stop spending review on
Style, naming and formatting are what reviewers reach for because they are easy to see, and they are the least valuable thing to catch by hand. Let a linter and a formatter own them, and spend the attention you save on the four categories above.
The failure that only shows up at volume
Nobody decides to rubber-stamp. It happens because there is more to review than there is attention, and approving is the only way to keep up. The signal is not a bad review, it is a gradual fall in how long reviews take while the number of changes rises.
Watch change failure rate before and after you adopted an agent. It is the number that says whether review is keeping up, and it moves quietly.
Tests become load-bearing
They used to be a safety net. When an agent iterates until the suite is green, they become the specification it is working against, and a test that passes whether or not the behaviour exists is no longer merely useless. It is a wrong instruction.
The cheapest habit that fixes this: when you add a test, break the thing it guards and confirm it goes red. If it stays green, delete it and say why.
A checklist worth keeping
- Was anything actually run, and is the output visible?
- Which callers does this change affect, and were they opened?
- What happens on the failure branch?
- Was anything deleted, and why?
- Does each new test fail when its behaviour is removed?
Where AstraCode fits
Most of the list above is work you do because the tool did not. AstraCode runs your tests before it reports, shows you what it ran rather than a summary of it, reads across the whole repository so the callers are in scope, and hands you a diff to approve rather than edits already on disk. Free tier, no card.
Review a diff, not a fait accompli
Changes arrive as something to approve, with the test output that backs them. Free to start.
Also on the blog
- When an AI agent says it is done and it is not
- Making an AI agent follow your project's conventions
- We benchmarked eleven models in our own editor. The cheapest one won.
- Fable 5.1 cut cache reads to $0.25. Here is what that saves on a real agent run.
- What vibe coding is, and when it stops working
- What Google Antigravity is, and what it costs
- What GPT-6 Astra actually costs to run a coding agent
- GPT-6 Astra's pricing cliff at 272K tokens, and why agent runs fall off it
- How we're benchmarking GPT-6 Astra for coding (and why scores won't tell you)
- How to use GPT-6 Astra in your editor, and when not to