Blog

How we're benchmarking GPT-6 Astra for coding (and why scores won't tell you)

7 September 2026 | 6 min

Three real tasks, protected test files, negative testing, and a rule that a model only ships if it passes every attempt. Results when we have them, not before.

GPT-6 Astra launched with strong coding numbers and a claim from OpenAI's president that it marks the start of the AGI era. Within a day the roundups were out. Almost none of them had run it on a repository.

We are going to, using the harness we already use to decide what AstraOne runs. This post is the methodology, published before the results, so that when the numbers arrive nobody has to take our word for how they were produced.

Three tasks, chosen to be different from each other

The rules that make the numbers mean anything

These matter more than the task list. A benchmark whose start state drifts measures drift.

What we report, and what we will not

The column that decides things is dollars per passing task, total spend divided by successes, so a model that fails attempts pays for them. Cost per token is not cost per finished work, and the gap between those two is where every surprising result we have had came from.

We will not publish a single run as a result. Our own rule is a minimum of three attempts per cell and a pinned commit, and the last time we broke it we said so in the write-up rather than quietly rounding.

What this caught last time

Two things, both of which contradicted what we expected.

The per-step router, the feature the whole category shipped this year, measured 0.4% worse than not routing at all. Not worse within a confidence interval we could argue with. Just no better, at the cost of a pile of machinery. So we deleted it. AstraOne runs one measured model end to end, and brings in a stronger one only when a check fails.

And a cheaper candidate passed two runs of six, failing the rest by editing the code, leaving the suite red, and reporting success. Only the pass command noticed. That failure mode is the entire argument for protected paths, and it is invisible to any benchmark that trusts a model's own account of what it did.

When

The harness passes its dry run against Astra now; the paid sweep needs provider credit and a bench gateway, and Astra's own rollout is still reaching accounts. When it runs, the numbers go up here whichever way they fall, including if Astra wins, which would mean changing what AstraOne runs. That is what having a benchmark is for.

Use the model that won, not the one with the best launch post

AstraOne runs whatever passes this benchmark, and the benchmark is re-run when a model ships. Free to start, no card.

Start free | Download

Also on the blog