We benchmarked eleven models in our own editor. The cheapest one won.
24 September 2026 | 8 min
Step count, not price per token, decides what an agent run costs. And predicting which model a task needs loses to simply noticing when the work failed.
Every AI editor asks you to pick a model. We wanted to know whether that choice is worth making, so we ran eleven of them through the harness we use to decide what our own agent runs. Four tasks from this repository, each graded twice: once by the visible checks, and once by a held-out check the agent never sees.
The short version is that the cheapest model we tested was unbeaten on cost per finished task, by a factor of seven to ninety. That is not the interesting part. The interesting part is why.
Expensive models did not take fewer steps. They took more.
The argument for a frontier model is that it thinks harder and finishes in fewer moves. We had written that down ourselves, confidently, before anybody measured it. On this workload it is not true.
| Model | Passed | Median steps | Cost per passing task |
|---|---|---|---|
| grok-4.6 | 4 of 4 | 10 | $0.28 |
| gpt-6-astra | 0 of 1 | 12 | did not finish |
| claude-opus-5 | 3 of 3 | 13 | $0.77 |
| gpt-5.6-luna | 18 of 27 | 15 | $0.06 |
| claude-sonnet-5 | 4 of 4 | 20 | $0.89 |
| gpt-5.6-terra | 2 of 3 | 32 | $3.52 |
| gpt-5 | 2 of 2 | 46 | $1.41 |
| qwen3-coder-next | 0 of 2 | 48 | did not finish |
| z-ai/glm-5.3 | 1 of 1 | 130 | $3.41 |
One task, so read it as a shape rather than a ranking. GLM 5.3 solved it, in 130 steps and twenty-one minutes, where the cheap model takes about fifteen steps and twenty-three seconds. Its headline claim was halving the output tokens its predecessor needed. On our repository it burned more than anything else we have ever measured.
Price per token predicts almost nothing. A model at ten times the rate that also takes four times the steps costs forty times as much, and the step count is the part no pricing page tells you.
Routing by predicted difficulty has no headroom
The obvious idea is to send hard requests to a strong model and easy ones to a cheap one. We tested it three ways and it failed three times.
- An oracle with perfect hindsight, allowed to pick the cheapest model that passes each task, picks the cheap model on every single task we have data for. There is nothing for a classifier to win.
- Our own earlier per-step router, measured against not routing at all, came out 0.4% worse per passing task.
- A purpose-built classifier, shown only the prompt, predicted which tasks the cheap model would struggle with at a rank correlation of about 0.75. Better than chance, and it still lost to doing nothing clever, because it commits before the run and pays for the expensive model on attempts the cheap one would have completed.
There is a deeper reason no better classifier fixes this. Our cheap model passes one task 67% of the time <em>on the same prompt</em>. The variation that matters is between attempts, not between requests, and nothing reading your message can see it.
Noticing beats predicting, and noticing is free
So we stopped guessing. A run now finishes on the model it started on, runs your project's own checks, and only when those go red does it hand the work to a stronger model and carry on. The information arrives thirty seconds later and it is exact rather than probabilistic.
| Policy | Work completed | Relative cost |
|---|---|---|
| Cheap model alone | 0.64 | 1x |
| Strong model throughout | 0.96 | 22x |
| Cheap model, escalating on a failed check | ~1.00 | 7x |
Escalating beats running the strong model throughout on both axes: it finishes more work and costs a third as much. You pay for the expensive model on the runs that need it, which is the only time it was ever worth paying for.
Two things we got wrong on the way
The first candidate we got excited about looked like a bargain on one task: ten steps, twenty-eight cents, cheaper and faster than every frontier model. Across four tasks it costs seven to thirty-four times our default. One task made it look like a find. Four made it look like a good strong model, which is what it is, and it is now the model our escalation reaches for.
The second was a reranker. Our code search asks a model to order twenty candidate snippets, and it was quietly declining to answer about a quarter of the time, leaving searches unranked with nothing to show for it. We evaluated a specialist model for the job. It was better. Then we asked our existing model the same question in the same shape, one score per snippet instead of one ordering of twenty, and it matched the specialist on every measure and answered every query. The gain was the shape of the question, not the model. We bought nothing and fixed it.
What this means if you use AstraCode
There is no model picker any more unless you bring your own API key, and there is nothing to configure. The mode is called AstraOne. It runs the model that measured best, checks the result against your tests, and escalates itself when those tests fail. Every plan has that, including the free one.
What we are not claiming is that this beats every model available. One model in our own table matched it on completion. Four tasks written by us are not a basis for a statement about the market, and a superiority number with no published method is exactly what we find unconvincing when other vendors publish it. Every figure above comes out of the harness in our repository, and the tasks, the grading and the held-out checks are in it.
Let it pick, and let it fix its own work
AstraOne runs the model that measured best, checks the result against your own tests, and brings in a stronger one when they fail. On every plan, including the free one. No card.
Also on the blog
- When an AI agent says it is done and it is not
- How to review code an agent wrote
- Making an AI agent follow your project's conventions
- Fable 5.1 cut cache reads to $0.25. Here is what that saves on a real agent run.
- What vibe coding is, and when it stops working
- What Google Antigravity is, and what it costs
- What GPT-6 Astra actually costs to run a coding agent
- GPT-6 Astra's pricing cliff at 272K tokens, and why agent runs fall off it
- How we're benchmarking GPT-6 Astra for coding (and why scores won't tell you)
- How to use GPT-6 Astra in your editor, and when not to