Writing
What we measured, and what it cost.
When an AI agent says it is done and it is not
24 September 2026 | 6 min
The edit is usually fine. The claim about the edit is the problem, and it is a harder one.
How to review code an agent wrote
24 September 2026 | 7 min
The bottleneck moved from writing code to deciding whether code is correct, and most teams are still reviewing the old way.
Making an AI agent follow your project's conventions
24 September 2026 | 6 min
Every agent writes like the average of the internet until you give it a reason not to. Most attempts at giving it one do not work.
We benchmarked eleven models in our own editor. The cheapest one won.
24 September 2026 | 8 min
Step count, not price per token, decides what an agent run costs. And predicting which model a task needs loses to simply noticing when the work failed.
Fable 5.1 cut cache reads to $0.25. Here is what that saves on a real agent run.
15 September 2026 | 4 min
Input and output cost exactly what they did. The one number that moved is the one an agent run is mostly made of.
What vibe coding is, and when it stops working
10 September 2026 | 5 min
The most-searched phrase in this category and still the least explained. No hype, no dismissal, and an honest account of where it breaks.
What Google Antigravity is, and what it costs
10 September 2026 | 6 min
Antigravity is the most-searched new thing in this category. Here is what it actually is, and a comparison that does not pretend we win every row.
What GPT-6 Astra actually costs to run a coding agent
7 September 2026 | 5 min
Per-token prices tell you almost nothing about what an agent run costs. Here is the same run, priced across five models, from a token profile we actually measured.
GPT-6 Astra's pricing cliff at 272K tokens, and why agent runs fall off it
7 September 2026 | 4 min
It is the second row of the pricing table, it nearly doubles the bill, and the workload most likely to trigger it is the one Astra is being sold for.
How we're benchmarking GPT-6 Astra for coding (and why scores won't tell you)
7 September 2026 | 6 min
Three real tasks, protected test files, negative testing, and a rule that a model only ships if it passes every attempt. Results when we have them, not before.
How to use GPT-6 Astra in your editor, and when not to
13 September 2026 | 4 min
Everyone has read the benchmarks. Almost nobody has said what it costs to actually point it at your repository for an afternoon.