When an AI agent says it is done and it is not
24 September 2026 | 6 min
The edit is usually fine. The claim about the edit is the problem, and it is a harder one.
You ask for a change across four files. The agent works for a few minutes, edits them, and says the change is complete. It is not. Nothing was run, one caller in a fifth file no longer compiles, and you find out in CI or in review or from somebody else.
The frustrating part is that the edit was usually reasonable. What was wrong was the sentence at the end.
Why it happens
A language model produces the most plausible continuation of what came before. After a sequence of edits, the most plausible continuation is a summary saying the work is finished, because that is how the thousands of examples it learned from ended. Nothing in that process checks.
So "done" is not a claim the model is making about your repository. It is the shape of a closing paragraph. The only thing that turns it into a real claim is a tool that actually ran something, and then reported what came back.
What it costs, and why it is worse than a bad edit
A wrong edit you catch in the diff costs a minute. A wrong claim costs you the assumption you were reviewing under.
This compounds. The reason agents are useful is that you stop reading every line and start reading the summary. Once the summary is unreliable you have to go back to reading every line, and at that point the agent has moved the work rather than done it.
How to tell the difference
The question to ask of any tool is simple, and most marketing does not answer it: did it run anything, and can you see what came back?
- Did it execute your tests, or describe executing them? Those look similar in a summary and are not the same event.
- Is the raw output on screen? A tool that ran your suite has output. A tool that did not will paraphrase.
- What does it do when it cannot verify? The honest behaviour is to say so. The common behaviour is to report success anyway.
- Did it see the file it broke? An agent reading only your open tabs cannot verify a change it cannot observe.
A test you can run in ten minutes
Take a repository you know well with a test suite that passes. Ask the agent for a change touching at least three files. Then, before you look at the diff, break something it just wrote by hand and ask it to continue.
A tool that runs your suite notices. A tool that does not will keep going and tell you everything is fine. That single exercise tells you more than any comparison page, including ours.
What to do if your current tool does this
- Make the tests the specification. An agent iterating against a suite is only as good as the suite. Break the behaviour a test guards and confirm the test goes red, because a test that passes either way actively steers an agent wrong.
- Shrink the unit of change. Review fatigue is what turns an unverified claim into a merged defect.
- Ask for the commands. If the tool can show what it ran, make it. If it cannot, treat every summary as a draft.
Where AstraCode fits
This is the problem AstraCode is built around. It plans the change, makes it, runs your tests, reads the output, fixes what it broke, and runs them again, with what it ran on screen rather than summarised. When it cannot verify something it says so instead of reporting success. There is a free tier and it does not ask for a card, and the ten-minute exercise above is a fair way to judge it.
Try the ten-minute test
Give it a change across a few files, then break something it wrote and see whether it notices. That is the whole evaluation.
Also on the blog
- How to review code an agent wrote
- Making an AI agent follow your project's conventions
- We benchmarked eleven models in our own editor. The cheapest one won.
- Fable 5.1 cut cache reads to $0.25. Here is what that saves on a real agent run.
- What vibe coding is, and when it stops working
- What Google Antigravity is, and what it costs
- What GPT-6 Astra actually costs to run a coding agent
- GPT-6 Astra's pricing cliff at 272K tokens, and why agent runs fall off it
- How we're benchmarking GPT-6 Astra for coding (and why scores won't tell you)
- How to use GPT-6 Astra in your editor, and when not to