GPT-6 Astra's pricing cliff at 272K tokens, and why agent runs fall off it
7 September 2026 | 4 min
It is the second row of the pricing table, it nearly doubles the bill, and the workload most likely to trigger it is the one Astra is being sold for.
GPT-6 Astra advertises a 1,050,000 token context window. It also reprices past 272,000 prompt tokens: input goes from $10 to $20 per million, output from $50 to $75. The important word is which tokens.
The higher rate applies to the entire request, not to the tokens past the line. A 272,001-token prompt does not cost a fraction more than a 271,999-token one. It costs roughly twice as much.
Why this hits agents specifically
A chat turn is small. An agent turn is not: it carries the system prompt, the tool definitions, every file it has read, and the whole transcript so far, and it carries them again on every step, which is why prompt caching exists at all.
We measured one run of the smallest task in our benchmark, a single localized defect in a TypeScript repository, fixed in 14 requests. Its prompt totalled 408,878 tokens. That is not a long session or a large monorepo. That is one small task, and it clears the cliff by 50%.
What it does to the bill
Priced through the same formula our gateway bills with, that run models at $1.61 at Astra's base rate and $3.16 at the long-context rate. The cliff is 96% of the base cost, added for tokens the agent was always going to send. Modelled, not measured on Astra, see the full cost breakdown for the token profile and the control.
Three ways this bites quietly
- Your meter under-reports. Most cost models carry one rate per model. If yours does, every long request is billed at half what it cost, and you find out at the end of the month. Ours had this bug for grok-4.6 before it had the fix.
- Compaction thresholds derived from the context window are wrong for this. Compacting at 70% of 1.05M means compacting at 735K, long past the point where every request costs double. The window and the price line are different numbers and they are not close.
- Caching does not save you. Cached reads bill at a tenth of input, but they still count toward the prompt total that triggers the cliff. A well-cached run is a cheap run that reprices anyway.
What we did about it
We already had one of these to handle: grok-4.6 reprices the whole request at 200K. That was a special case inlined in two files, which is exactly how the second one gets missed, so adding Astra meant turning both into a table of cliffs that the gateway and the editor read the same way.
We also had to give Astra its real window. Our context budget matched gpt- ids to 400K, so Astra would have compacted at 280K, discarding three quarters of a window we were paying for, and its cache with it.
None of that is clever. It is just the difference between a meter that tracks the bill and one that tells you what you would like to hear.
Priced at the rate that actually applied
Long-context cliffs included, which is what keeps AstraOne's choice of model honest rather than optimistic. Free to start, no card.
Also on the blog
- When an AI agent says it is done and it is not
- How to review code an agent wrote
- Making an AI agent follow your project's conventions
- We benchmarked eleven models in our own editor. The cheapest one won.
- Fable 5.1 cut cache reads to $0.25. Here is what that saves on a real agent run.
- What vibe coding is, and when it stops working
- What Google Antigravity is, and what it costs
- What GPT-6 Astra actually costs to run a coding agent
- How we're benchmarking GPT-6 Astra for coding (and why scores won't tell you)
- How to use GPT-6 Astra in your editor, and when not to