Fable 5.1 cut cache reads to $0.25. Here is what that saves on a real agent run.
15 September 2026 | 4 min
Input and output cost exactly what they did. The one number that moved is the one an agent run is mostly made of.
Anthropic released Claude Fable 5.1 on 1 September 2026. The headline rates did not move: $10 per million input tokens, $50 per million output, the same as Fable 5. Cache writes stay at $12.50. The change is cache reads, which went from $1.00 to $0.25 per million, a quarter of what they were.
That sounds like a footnote. For an agent it is most of the bill.
Why cache reads are the number that matters
An agent run is a loop: read a file, run a command, edit, run the tests, read the output. Every turn sends the whole conversation back to the model. Without a prompt cache that is the same 300,000 tokens billed again and again; with one, everything already seen is billed at the cache-read rate, and only the new part costs full price.
Here is the token profile of one real run from our benchmark: an agent fixing a defect across 14 requests, measured, not estimated. 82% of every prompt token it sent was a cache read.
| Token class | Count | Share of prompt |
|---|---|---|
| Fresh input | 35,894 | 9% |
| Cache writes | 39,162 | 10% |
| Cache reads | 333,822 | 82% |
| Output | 2,740 | not a prompt token |
What the cut saves on that run
Priced at each model's published rates, cache prices included, the same run costs $1.32 on Fable 5 and $1.07 on Fable 5.1. That is $0.25 less, or 19%, for changing nothing but the model name.
Look at where the money goes. On Fable 5, cache reads were 25% of the run. On Fable 5.1 they are 8%. The biggest line on the bill is now the cache writes, at 46%, which is the cost of putting new context into the cache in the first place.
| Fable 5 | Fable 5.1 | |
|---|---|---|
| Fresh input | $0.36 | $0.36 |
| Cache writes | $0.49 | $0.49 |
| Cache reads | $0.33 | $0.08 |
| Output | $0.14 | $0.14 |
| One run | $1.32 | $1.07 |
The cut only pays where the cache is already working. A chat turn with a fresh prompt has almost no cache reads, so it costs what it cost last week. If your tool rewrites the system prompt every turn, or does not cache at all, Fable 5.1 is exactly as expensive as Fable 5.
Against the rest of the family
The same run at published rates: $0.66 on Opus 5, $0.26 on Sonnet 5, $0.13 on Haiku 4.5. Fable 5.1 is now about 1.6x Opus on an agent run, down from 2.0x, and still 4x Sonnet.
PUBLISHED by ops/make-blog-figures.py, so it cannot disagree with the table.So the top model got cheaper for the workload it is sold for, and the gap to the mid-tier narrowed. It did not close. A run that a smaller model can finish is still a run you should not send to the largest one, and in our benchmark the smaller model finished the task in roughly half the steps (what vibe coding is, and when it stops working).
What this changes in AstraCode
Nothing you have to do. AstraCode prices every step at the provider's current rate, so a step Fable serves costs what Fable 5.1 charges today. AstraOne chooses the model for each run: it runs the model measured to finish real repository work, and brings in a stronger one only when a check fails, so frontier prices are paid only where a smaller model could not finish.
If you bring your own Anthropic key on a Pro plan or above, the saving lands on your own bill directly.
Also this month
- GPT-6 Astra, OpenAI, 3 September: $10/$50 with a 1M context and a repricing past 272K tokens that an agent run crosses easily. We wrote about the cliff and what a run costs.
- Muse Spark 1.3, Meta, 2 September: a frontier model aimed at coding and long agent runs, listed by llm-stats at $0.12 per million tokens.
- Gemini 3.8 Flash, Google, 2 September: the new everyday model.
- DeepSeek V4.1 Flash, DeepSeek, 10 September.
Modelled, not measured on Fable 5.1. The token profile is a real, measured run; the dollar figures are that profile priced at each provider's published rates as of 15 September 2026. Our earlier posts priced runs with the gateway's own billing formula, which treats cache writes and reads slightly differently, so their Opus and Sonnet figures differ from these by a few cents.
- Claude pricing, Anthropic, retrieved 15 September 2026.
- AI news, September 2026, llm-stats.com.
- Model release timeline, LLM Gateway.
Let AstraOne choose for you
AstraOne runs the model measured to finish the job and brings in a stronger one only when a check fails, so the frontier price is paid only where it earns its place. Free to start, no card.
Also on the blog
- When an AI agent says it is done and it is not
- How to review code an agent wrote
- Making an AI agent follow your project's conventions
- We benchmarked eleven models in our own editor. The cheapest one won.
- What vibe coding is, and when it stops working
- What Google Antigravity is, and what it costs
- What GPT-6 Astra actually costs to run a coding agent
- GPT-6 Astra's pricing cliff at 272K tokens, and why agent runs fall off it
- How we're benchmarking GPT-6 Astra for coding (and why scores won't tell you)
- How to use GPT-6 Astra in your editor, and when not to