TL;DR. The vibe coding era is real — capital proves it. But the productivity it promised became unmeasurable before it was proven: in a year we went from "AI makes experienced devs 19% slower (they felt 20% faster)" to "we can no longer measure it, because experienced devs refuse to work without AI." What survives the collapse of measurement is judgment — and its erosion, like the +20% illusion, is invisible from the inside. So judgment has to be trained on purpose, as ritual.
The Era Is Real
A month after Andrej Karpathy coined "vibe coding" in February 2025, Y Combinator reported that a quarter of the W25 batch had codebases that were 95% AI-generated. Garry Tan called it: "the age of vibe coding is here." The bubble-or-real debate closes in front of the adoption and capital numbers.
- Scale — GitHub reported that in 2025 alone it added 36.2 million developers to pass 180 million, merges 43.2 million PRs a month on average, and saw generative-AI projects grow 178% year over year (Octoverse 2025). Copilot has 20 million users; 90% of the Fortune 100 have adopted it.
- Capital — Anysphere, maker of Cursor, moved into a $50B-valuation round in April 2026 (up from $29.3B in November 2025), its ARR past $2B in three years — the fastest-growing software business ever measured. Anthropic reached a $965B (~$1T) valuation in its May 2026 Series H, with Claude Code's run-rate at $2.5B+ (more than doubled since the start of 2026) and enterprise over half its revenue. In 2025, $211B — roughly half of all global venture funding — flowed into AI.
The era is real. Not a trend word, but a fact that capital has finished betting on.
What AI Can Now Actually Do
Not in numbers — on the ground.
There's the story of an engineer who spent ten years in finance and payments. PCI rules, double-entry ledgers, payment flows — a decade of domain knowledge was his core asset. Then one day a manager said, "just ask the AI." What made that sting was that it wasn't wrong.
Debugging fell faster. Tracing distributed-systems bugs was the last redoubt people believed only humans could hold. That belief broke as the latest AI started solving one-to-two-day bugs in a single pass. Wasmer reported it used Codex to build an edge Node.js runtime 10–20× faster.
Design and architecture were thought to be the last human ground. That too is shrinking to "taste," the engineer says — companies have begun to prefer code that's easy for AI to process over code that's easy for humans to read. Domain → debugging → design: the places where humans were believed to hold a comparative advantage are narrowing in that order.
Now We Can't Even Measure It
That AI can do these things is a different question from whether it raises productivity. And that productivity — became unmeasurable before it was proven.
Start with the most-quoted number, "AI writes X% of the code." There's no agreement. Count by commit authorship and it's 5%; count tab-acceptance in files where Copilot is on and it's 46% (2023). Nadella said "20 to 30% of our repos" (Apr 2025), Pichai said "25%, now over 30% of new code," and DX's survey self-reports 22% (Q4 2025). The same question swings from 5 to 46 — an 8× gap. (The widely circulated "GitHub 41%" is a mutation of a 2023 interview with no primary source.) There is no measurement standard.
Productivity splits even harder. GitHub says Copilot makes tasks 55% faster — but that's 95 developers writing one HTTP server, a single toy task, in a vendor-run study (2022); a longitudinal study of the same tool found no statistically significant change in commit activity. On the other side is METR's RCT: 16 experienced developers in their own mature repos (avg 20k+ stars) were 19% slower with AI on. And after the trial they believed they'd been 20% faster.
That far is the "mirage." But in 2026 it went one step further — the measurement itself collapsed. METR tried to rerun the experiment and couldn't complete it. Developers refused to work without AI. There was no way to form a control group. One participant put it: "it felt like trying to get across the city walking when all of a sudden I was more used to taking an Uber." METR now believes AI modestly speeds up experienced devs in 2026 (raw −18% / −4%, uncertainty +9% to −38%) — but flags this as "only very weak evidence."
The move in one year is the point. From we measured it (19% slower) to we can't even measure it (no control group). I've lined up elsewhere how AI evaluation is collapsing — this is the developer-productivity edition of that same collapse. DORA 2025 calling AI an amplifier (throughput up, stability down) is the same grain — you can see what went up, but not whether it got better.
The Side Effects Show Up Only in the Records
The cost signals aren't one strand. Over the past year they landed in four.
- Vendor opacity — Anthropic admitted, in an official postmortem, to a Claude Code quality degradation from 3/4 to 4/20. A rare case of a company documenting its own silent downgrade.
- Attribution pollution — VS Code 1.117 auto-attributed even hand-written code from users who had Copilot off as Co-authored-by: Copilot (reverted in a 5/3 patch).
- Quality erosion — Per code-analytics firm GitClear, copy-pasted lines rose from 8.3% (2021) to 12.3% (2024), while refactored (moved) code fell from 25% to under 10% — copy-paste exceeded refactoring for the first time (vendor analysis, not causal).
- Burden-shifting — Daniel Stenberg shut down the curl bug bounty after a flood of AI slop reports drove the real-report rate below 5%.
And trust. In the Stack Overflow 2025 survey, 84% use AI, but only 3% highly trust its accuracy — 2.6% among experienced developers (20% highly distrust). The #1 frustration is code that's "almost right, but not quite" — 66% spend more time fixing it. The #1 reason to ask a human instead is "when I don't trust the AI's answer" (75%). They don't trust it, and accept it anyway.
Why It Stays On
Even after seeing all the side effects, I'm on the same side. Two reasons.
Asymmetry. Side effects are recorded only by the people using it from the inside. The felt-+20%, measured-−19% gap surfaces only inside an RCT; a silent downgrade is visible only once the company issues a postmortem. Those who don't use it don't even have the record. No record, weaker prescription and weaker critique.
Irreversibility. Side effects can be softened by how you use the tool — token cost drops 25–30% with six levers, and quality erosion is patched by the judgment that separates what to delegate from what not to. But time not spent isn't recovered. The gap between someone who's used it daily since last year and someone turning it on next month closes with no lever.
What Survives Is Judgment — and You Can't See It Erode
Three numbers point to the same place. Code share scattering across 5–46% means volume is no longer a meaningful metric; METR's felt-vs-measured inversion means Accept All without knowing what to accept actually makes you slower; DORA's stability drop means generation got easy and judgment didn't keep up. The bottleneck moved from writing code to deciding what to build, what to accept, what to throw away, what to verify. So the seat that survives is judgment, not production. AI writes the code for you, but it does not decide what's correct for you.
The cruel part is here — you can't tell from the inside that your judgment is dulling. METR's +20% illusion wasn't a measurement error; it was evidence that the ability to assess your own performance had broken. Judgment atrophy sits in the same blind spot. The two unmeasurable things — whether AI helps you, and whether your judgment is melting — both register as positive from the inside. The engineer's words are exact: "I'm increasingly becoming a reviewer of AI. And my judgment as a reviewer is dulling too." Fewer chances to be wrong means less learning. Domain knowledge doesn't accrue when there's nothing to ask; debugging skill doesn't hold without tracing it yourself. And the erosion, like the measurement, is invisible from the inside. That judgment matters more, not less, when you're alone — the solo developer with no code review gets it wrong most often.
If Even Judgment Isn't the Human's Seat
There are two honest objections to that claim.
First, judgment automates too. The logic that moved the bottleneck from writing to judging pushes it past judging with the same force. Verifier models, LLM-as-judge, and automated code review are already attaching themselves to "reading the diff you're handed and saying a wrong thing is wrong" (the design that referees the opposition sits exactly there). Pinning judgment as the human's permanent seat has the same shape as the comfort every technological turn's incumbents recite — taste is what's left in the end. As hard to falsify as it is, it's at equal risk of being hollow.
Second, the market rewards speed over judgment. YC's W25 cohort, with 95% AI-generated codebases, didn't win on careful judgment — it won by shipping first. In front of throwaway prototype code, "the seat that judges what to throw away" is a luxury, and there vibe wins. The claim that judgment gets more expensive quietly assumes code that's worth judging.
Neither is a weak card. But both treat "judgment" as one undivided thing — and that's where they close.
The One Who Picks What to Throw Away Stays
Both objections close once you narrow judgment. What automates is correctness judgment ("is this code right?") — that's as far as a verifier scores. The judgment that survives sits above it: what to build, what to accept, what to throw away. What METR's inversion exposed wasn't a failure of correctness-checking but a failure of accepting without knowing what to accept. Judgment of goals and accountability can't be delegated to the model — because when it's wrong, the human pays. The places where speed beats judgment are bounded to code you can throw away too. In front of code that has to survive, speed can't buy judgment, and the market only looks like it rewards speed up until the moment that code hasn't been thrown away yet.
So the answer isn't to turn AI off. It's intent — since you can't measure it, you have to enforce it as ritual. Before accepting AI's code, you should be able to explain why this code; when AI fixes a bug, tracing the logic of the fix is how the debugging muscle stays; even when you ask AI for domain knowledge, the habit of verifying the answer is what protects expertise. Accepting and judging are different acts.
The gap is a function of time. It isn't counting how much code AI writes — the one who judges, every day, what to accept and what to throw away is the one who holds that seat. That the era is real means the cost of sitting idle is real too.
So the gap is a function of time. The distance between someone counting how much code AI writes and someone deciding what to accept and what to throw away every day widens a month from now. A year not spent with AI isn't recovered by turning it on next month. That the era is real means the cost of sitting it out is real too.
References
The era / capital
- GitHub Octoverse 2025 — 180M developers, 43.2M PRs/month, GenAI projects +178%
- YC / Garry Tan — 25% of W25 with 95% AI-generated code (TechCrunch, Mar 2025)
- Anysphere (Cursor) — $50B-valuation round, $2B ARR (TechCrunch, Apr 2026)
- Anthropic Series H — $965B valuation, Claude Code $2.5B+ run-rate (May 2026)
- Crunchbase — $211B into AI in 2025, ~half of all global VC
The collapse of measurement
- METR — 16 experienced devs 19% slower with AI, felt +20% (RCT, Jul 2025)
- METR — 2026 follow-up: no control group (developers refuse to work without AI) (Feb 2026)
- GitHub — Copilot "55% faster" (95 developers, single toy task, vendor, 2022) · longitudinal — no significant change in commit activity (arXiv 2509.20353)
- DORA 2025 — AI as amplifier: throughput up, stability down
- Nadella — "20–30% of our repos are AI" (CNBC, Apr 2025) · DX — self-reported 22% (Q4 2025)
- The collapse of AI evaluation — another edition of measurement collapse
The side effects
- Anthropic — April 23 Postmortem: official admission of Claude Code degradation
- microsoft/vscode #314311 — Co-authored-by Copilot auto-attribution bug
- GitClear — AI code quality 2025: copy-paste up, refactoring down
- Daniel Stenberg — the end of the curl bug bounty (Jan 2026)
- Stack Overflow 2025 — 3% highly trust AI accuracy (2.6% among experienced), 66% "almost right but not quite"
Related