Blog

What GPT-6 Astra actually costs to run a coding agent

7 September 2026 | 5 min

Per-token prices tell you almost nothing about what an agent run costs. Here is the same run, priced across five models, from a token profile we actually measured.

OpenAI released GPT-6 Astra on 3 September at $10 per million input tokens and $50 per million output, with cached input at $1. Those are the numbers everyone quoted. They are also close to useless on their own, because nobody sends one million tokens, they run an agent, and an agent run has a shape.

The shape of one agent run

We have a benchmark harness that runs real repository tasks against a fixture and records what the gateway actually billed. Task A is a localized defect in a TypeScript codebase: the existing test suite catches it, the test files are protected so the agent cannot make it pass by editing the assertion, and the fix is discoverable from a docstring. It is the smallest of our three tasks.

One run of it, measured on Opus across 14 requests, looks like this:

Token classCountBilled at
Fresh input35,894full input rate
Cache writes39,1622x input
Cache reads333,8220.1x input
Output2,740output rate

Two things jump out. Output is under 1% of the traffic, an agent reads far more than it writes, so the output price you compared models on barely matters. And the prompt totals 408,878 tokens in a single task.

Priced across five models

Same profile, each model's published rates, the same formula our gateway bills with:

ModelModelled cost, one run
gpt-6-astra$3.16past its 272K cliff, so $20/$75
claude-opus-5$0.81measured at $0.76 on the same task
gpt-5.6-sol$0.65promotional rate to 2026-11-21
claude-sonnet-5$0.32measured at $0.32
gpt-5.6-luna$0.03what AstraOne runs
Horizontal bar chart of the modelled cost of one 14-request agent run on each model: gpt-6-astra $3.16, claude-opus-5 $0.81, gpt-5.6-sol $0.65, claude-sonnet-5 $0.32, gpt-5.6-luna $0.03.
The same numbers as the table above, drawn to scale. Generated from RUN_COST by ops/make-blog-figures.py, so it cannot disagree with it.

These are modelled, not measured on Astra. We have not run Astra through the benchmark yet. The Opus row is the control: it models at $0.81 against $0.76 actually billed for the same task, a 6% gap that comes from pricing every cache write at the 1-hour rate. Treat all of these as approximate.

Astra is not 2.5x Opus. It is nearly 4x

On per-token list price Astra is twice Opus on input and twice on output. On this run it is 3.9x, and the reason is the second number in its pricing page: past 272,000 prompt tokens the whole request reprices to $20 input and $75 output. Not the excess, the whole request.

408,878 is one and a half times that line. A single task, on the smallest benchmark we run, sits well past it. At the base rate the run models at $1.61; the cliff adds another $1.54, a 96% increase, for tokens the agent was always going to send.

We wrote that up separately, because it applies to how you build agents and not just to what you pay: GPT-6 Astra's pricing cliff at 272K tokens.

Is it worth it?

We genuinely do not know yet, and neither does anyone quoting benchmark scores four days after launch. Cost per token is not cost per finished task. A model that costs 4x and finishes in a third of the steps is cheaper; a model that costs 4x and takes the same number of steps is four times more expensive, and that is a thing you have to measure.

We measured exactly that question once before and the answer was not the one we expected: the cheap model took fewer steps than the frontier one, not more, and the per-step router we built to hedge between them came out 0.4% worse than not routing at all. So we deleted the router. Here is how that benchmark works, and it is what we will put Astra through.

What to do in the meantime

A cliff AstraCode already knows about

Astra is priced correctly in AstraCode, cliff and all, so a run on your own key is metered at the rate that actually applied. Free to start, no card.

Start free | Download

Also on the blog