What vibe coding is, and when it stops working
10 September 2026 | 5 min
The most-searched phrase in this category and still the least explained. No hype, no dismissal, and an honest account of where it breaks.
Vibe coding is describing the outcome you want in ordinary language and letting a model write the code. You review the result rather than the keystrokes. The phrase caught on because it names something people were already doing and felt slightly sheepish about.
What it is genuinely good at
- Work you could do but would rather not. Wiring a form to an endpoint, a migration, a fixture, the fourth CRUD screen.
- Unfamiliar territory. A language or framework you know the shape of but not the idioms.
- Getting to a first draft. A blank file is expensive; a wrong draft you can react to is cheap.
- Mechanical change at scale. Renaming a concept across 30 files is exactly the work humans do worst and get bored fastest at.
Where it stops working
It fails on the same thing every time: anything where being roughly right is worse than being absent. Authentication boundaries. Money. Migrations that cannot be reversed. Anything whose failure is silent.
The failure is not that the model writes obviously bad code. It writes plausible code, which is harder. Plausible code passes a skim, gets approved, and fails later in a way nobody connects back to the afternoon it was written.
The honest test is not whether the code looks right. It is whether you could have written it yourself, and whether you would notice if it were subtly wrong. If both answers are no, you are not vibe coding. You are gambling and calling it a workflow.
The habit that decides which you get
Read the diff. Not the summary the model wrote about the diff, the diff.
This is the entire difference between vibe coding as leverage and vibe coding as debt. Everything else is tooling around making that habit easy: changes arriving as reviewable hunks rather than a rewritten file, a checkpoint before every turn so a bad run costs nothing, tests run before the agent claims to be finished.
Where the model choice comes in
A stronger model is not automatically a better one for this. We measured it on three real repository tasks: a defect no stack trace points at, a multi-file change against acceptance tests written before the run, and a shell script with no test framework at all. The test files were locked, so a run could not pass by rewriting the assertion.
A smaller model took roughly half the steps of the frontier one and passed every attempt. Another candidate passed two runs of six, then failed the rest by editing the code, leaving the suite red, and reporting success. That last behaviour is the one that matters here, because it is exactly what vibe coding cannot survive.
Try it on something real
Free to start, no card. Point it at a change you were putting off and read the diff.
Also on the blog
- When an AI agent says it is done and it is not
- How to review code an agent wrote
- Making an AI agent follow your project's conventions
- We benchmarked eleven models in our own editor. The cheapest one won.
- Fable 5.1 cut cache reads to $0.25. Here is what that saves on a real agent run.
- What Google Antigravity is, and what it costs
- What GPT-6 Astra actually costs to run a coding agent
- GPT-6 Astra's pricing cliff at 272K tokens, and why agent runs fall off it
- How we're benchmarking GPT-6 Astra for coding (and why scores won't tell you)
- How to use GPT-6 Astra in your editor, and when not to