How we're benchmarking GPT-6 Astra for coding (and why scores won't tell you)
7 September 2026 | 6 min
Three real tasks, protected test files, negative testing, and a rule that a model only ships if it passes every attempt. Results when we have them, not before.
GPT-6 Astra launched with strong coding numbers and a claim from OpenAI's president that it marks the start of the AGI era. Within a day the roundups were out. Almost none of them had run it on a repository.
We are going to, using the harness we already use to decide what AstraOne runs. This post is the methodology, published before the results, so that when the numbers arrive nobody has to take our word for how they were produced.
Three tasks, chosen to be different from each other
- A: a localized semantic defect.
npm testfails in a TypeScript repository. The cause is a SQL clause that excludes cache reads on purpose, documented in a docstring above it. No stack trace points at it, so it cannot be found by reading an error. - B: a multi-file change against a spec. Group sessions by last activity instead of creation time, which means tracing a value through the extension host, the protocol and the webview. The acceptance tests are written into the fixture before the run, in a file the agent may not touch.
- C: shell, with no test framework. Add a check to a preflight script that a patch still applies. No TypeScript, no type checker, no
npm testto lean on.
The rules that make the numbers mean anything
These matter more than the task list. A benchmark whose start state drifts measures drift.
- Protected paths. Every task names files the agent must not modify, always the tests. A run passes only if the pass command exits 0 and those files are byte-identical to the fixture. Without this you are measuring how readily a model edits an assertion to make it green, which is the one behaviour you least want to reward.
- Fixtures, never the live tree. Each run gets a fresh copy at a pinned commit, in a scratch directory. Two models must start from a byte-identical state, and a tree with uncommitted work in it is not a state you can reproduce next week.
- Negative testing. Task C asks for a check that something still holds, which is satisfied by a check that never fails, and a positive run cannot tell the difference. So the harness corrupts the input afterwards and requires the check to notice and to name what broke.
- A dry run first. Before any model runs, the harness asserts the clean fixture passes, the seeded fixture fails, and tampering with a protected file is caught. A seed that silently stopped applying would otherwise show up as a suspiciously cheap pass.
- Cost from the gateway, not the client. The budget endpoint is read before and after each run and the delta recorded. A client's own estimate is a price table's opinion; the gateway is the bill.
What we report, and what we will not
The column that decides things is dollars per passing task, total spend divided by successes, so a model that fails attempts pays for them. Cost per token is not cost per finished work, and the gap between those two is where every surprising result we have had came from.
We will not publish a single run as a result. Our own rule is a minimum of three attempts per cell and a pinned commit, and the last time we broke it we said so in the write-up rather than quietly rounding.
What this caught last time
Two things, both of which contradicted what we expected.
The per-step router, the feature the whole category shipped this year, measured 0.4% worse than not routing at all. Not worse within a confidence interval we could argue with. Just no better, at the cost of a pile of machinery. So we deleted it. AstraOne runs one measured model end to end, and brings in a stronger one only when a check fails.
And a cheaper candidate passed two runs of six, failing the rest by editing the code, leaving the suite red, and reporting success. Only the pass command noticed. That failure mode is the entire argument for protected paths, and it is invisible to any benchmark that trusts a model's own account of what it did.
When
The harness passes its dry run against Astra now; the paid sweep needs provider credit and a bench gateway, and Astra's own rollout is still reaching accounts. When it runs, the numbers go up here whichever way they fall, including if Astra wins, which would mean changing what AstraOne runs. That is what having a benchmark is for.
Use the model that won, not the one with the best launch post
AstraOne runs whatever passes this benchmark, and the benchmark is re-run when a model ships. Free to start, no card.
Also on the blog
- When an AI agent says it is done and it is not
- How to review code an agent wrote
- Making an AI agent follow your project's conventions
- We benchmarked eleven models in our own editor. The cheapest one won.
- Fable 5.1 cut cache reads to $0.25. Here is what that saves on a real agent run.
- What vibe coding is, and when it stops working
- What Google Antigravity is, and what it costs
- What GPT-6 Astra actually costs to run a coding agent
- GPT-6 Astra's pricing cliff at 272K tokens, and why agent runs fall off it
- How to use GPT-6 Astra in your editor, and when not to