Blog

GPT-6 Astra's pricing cliff at 272K tokens, and why agent runs fall off it

7 September 2026 | 4 min

It is the second row of the pricing table, it nearly doubles the bill, and the workload most likely to trigger it is the one Astra is being sold for.

GPT-6 Astra advertises a 1,050,000 token context window. It also reprices past 272,000 prompt tokens: input goes from $10 to $20 per million, output from $50 to $75. The important word is which tokens.

The higher rate applies to the entire request, not to the tokens past the line. A 272,001-token prompt does not cost a fraction more than a 271,999-token one. It costs roughly twice as much.

Diagram showing that once a prompt passes 272,000 tokens, GPT-6 Astra prices the entire request at $20 per million input and $75 output instead of $10 and $50. A measured agent run carries 408,878 prompt tokens, past the line.
A step, not a slope. The higher rate applies to the whole request.

Why this hits agents specifically

A chat turn is small. An agent turn is not: it carries the system prompt, the tool definitions, every file it has read, and the whole transcript so far, and it carries them again on every step, which is why prompt caching exists at all.

We measured one run of the smallest task in our benchmark, a single localized defect in a TypeScript repository, fixed in 14 requests. Its prompt totalled 408,878 tokens. That is not a long session or a large monorepo. That is one small task, and it clears the cliff by 50%.

What it does to the bill

Priced through the same formula our gateway bills with, that run models at $1.61 at Astra's base rate and $3.16 at the long-context rate. The cliff is 96% of the base cost, added for tokens the agent was always going to send. Modelled, not measured on Astra, see the full cost breakdown for the token profile and the control.

Three ways this bites quietly

What we did about it

We already had one of these to handle: grok-4.6 reprices the whole request at 200K. That was a special case inlined in two files, which is exactly how the second one gets missed, so adding Astra meant turning both into a table of cliffs that the gateway and the editor read the same way.

We also had to give Astra its real window. Our context budget matched gpt- ids to 400K, so Astra would have compacted at 280K, discarding three quarters of a window we were paying for, and its cache with it.

None of that is clever. It is just the difference between a meter that tracks the bill and one that tells you what you would like to hear.

Priced at the rate that actually applied

Long-context cliffs included, which is what keeps AstraOne's choice of model honest rather than optimistic. Free to start, no card.

Start free | Download

Also on the blog