Blog

We benchmarked eleven models in our own editor. The cheapest one won.

24 September 2026 | 8 min

Step count, not price per token, decides what an agent run costs. And predicting which model a task needs loses to simply noticing when the work failed.

Every AI editor asks you to pick a model. We wanted to know whether that choice is worth making, so we ran eleven of them through the harness we use to decide what our own agent runs. Four tasks from this repository, each graded twice: once by the visible checks, and once by a held-out check the agent never sees.

The short version is that the cheapest model we tested was unbeaten on cost per finished task, by a factor of seven to ninety. That is not the interesting part. The interesting part is why.

Expensive models did not take fewer steps. They took more.

The argument for a frontier model is that it thinks harder and finishes in fewer moves. We had written that down ourselves, confidently, before anybody measured it. On this workload it is not true.

ModelPassedMedian stepsCost per passing task
grok-4.64 of 410$0.28
gpt-6-astra0 of 112did not finish
claude-opus-53 of 313$0.77
gpt-5.6-luna18 of 2715$0.06
claude-sonnet-54 of 420$0.89
gpt-5.6-terra2 of 332$3.52
gpt-52 of 246$1.41
qwen3-coder-next0 of 248did not finish
z-ai/glm-5.31 of 1130$3.41

One task, so read it as a shape rather than a ranking. GLM 5.3 solved it, in 130 steps and twenty-one minutes, where the cheap model takes about fifteen steps and twenty-three seconds. Its headline claim was halving the output tokens its predecessor needed. On our repository it burned more than anything else we have ever measured.

Price per token predicts almost nothing. A model at ten times the rate that also takes four times the steps costs forty times as much, and the step count is the part no pricing page tells you.

Routing by predicted difficulty has no headroom

The obvious idea is to send hard requests to a strong model and easy ones to a cheap one. We tested it three ways and it failed three times.

There is a deeper reason no better classifier fixes this. Our cheap model passes one task 67% of the time <em>on the same prompt</em>. The variation that matters is between attempts, not between requests, and nothing reading your message can see it.

Noticing beats predicting, and noticing is free

So we stopped guessing. A run now finishes on the model it started on, runs your project's own checks, and only when those go red does it hand the work to a stronger model and carry on. The information arrives thirty seconds later and it is exact rather than probabilistic.

PolicyWork completedRelative cost
Cheap model alone0.641x
Strong model throughout0.9622x
Cheap model, escalating on a failed check~1.007x

Escalating beats running the strong model throughout on both axes: it finishes more work and costs a third as much. You pay for the expensive model on the runs that need it, which is the only time it was ever worth paying for.

Two things we got wrong on the way

The first candidate we got excited about looked like a bargain on one task: ten steps, twenty-eight cents, cheaper and faster than every frontier model. Across four tasks it costs seven to thirty-four times our default. One task made it look like a find. Four made it look like a good strong model, which is what it is, and it is now the model our escalation reaches for.

The second was a reranker. Our code search asks a model to order twenty candidate snippets, and it was quietly declining to answer about a quarter of the time, leaving searches unranked with nothing to show for it. We evaluated a specialist model for the job. It was better. Then we asked our existing model the same question in the same shape, one score per snippet instead of one ordering of twenty, and it matched the specialist on every measure and answered every query. The gain was the shape of the question, not the model. We bought nothing and fixed it.

What this means if you use AstraCode

There is no model picker any more unless you bring your own API key, and there is nothing to configure. The mode is called AstraOne. It runs the model that measured best, checks the result against your tests, and escalates itself when those tests fail. Every plan has that, including the free one.

What we are not claiming is that this beats every model available. One model in our own table matched it on completion. Four tasks written by us are not a basis for a statement about the market, and a superiority number with no published method is exactly what we find unconvincing when other vendors publish it. Every figure above comes out of the harness in our repository, and the tasks, the grading and the held-out checks are in it.

Let it pick, and let it fix its own work

AstraOne runs the model that measured best, checks the result against your own tests, and brings in a stronger one when they fail. On every plan, including the free one. No card.

Start free | Download

Also on the blog