Developers will not give up AI coding tools, even temporarily for science. AI research lab METR tried to rerun its 2025 study measuring how much faster open source developers worked with AI versus by hand, and had to abandon the experiment in February 2026 because participants refused to code unplugged. The lab fell back on a self-report survey in May 2026, in which technical employees said AI made them roughly twice as valuable to their organizations.
That self-perception is now colliding with a growing pile of data suggesting the productivity gains are smaller, and the long-term costs higher, than the inside-the-IDE feeling implies. METR's original 2025 work already showed that while AI generated code faster, developers spent the saved time fixing errors, steering the model, and waiting on completions, with a net slowdown. The 2026 survey simply removed the stopwatch.
The dominant metric of the year has been tokenmaxxing — treating tokens consumed as a proxy for output. Amazon shut down Kirorank, its internal token-use leaderboard, after employees gamed it by deploying AI agents excessively and running up cloud bills, the Financial Times reported. Uber burned through its full 2026 AI budget in the first four months of the year, according to The Information, with COO Andrew Macdonald saying on a podcast that the spend did not translate into more shipped projects.
Key facts
- 01METR could not rerun its 2025 AI coding productivity study in February 2026 because developers refused to work without AI even briefly.
- 02Entelligence AI CEO Aiswarya Sankar says companies are spending 44% of their tokens on fixing bugs their AI generated.
- 03Code Rabbit's analysis of open source pull requests found AI-generated code produced 1.7x more problems than human code.
- 04Uber burned through its full 2026 AI budget in the first four months of the year with no measurable productivity gain, per COO Andrew Macdonald.
- 05Amazon shut down Kirorank, its internal AI token-use leaderboard, after employees gamed it by running up agent costs.
The maintenance bill is the part that doesn't show up on the leaderboard. Programmer James Shore argued in a Hacker News-trending post that faster code generation is only a win if the maintenance burden falls in step.
He put the trade-off in terms his readers couldn't miss, framing every saved hour of generation as a debt against future maintenance budgets that few engineering leaders have yet started to track.
Entelligence AI CEO Aiswarya Sankar, whose startup sells a reliability-engineering agent, posted that companies are now spending 44% of their tokens on bug fixes for code their AI generated. Code-review tool Code Rabbit says it analyzed open source pull requests and found AI-generated code produced 1.7x more problems than human-written code. Both numbers come from vendors with skin in the game, but independent work points the same direction: Singapore Management University researchers warned in April 2026 that AI-generated code can introduce long-term maintenance costs into real software projects.
Uber's experience is the cleanest enterprise data point so far.
“such spending hadn't led to a measurable increase in projects or productivity”— Andrew Macdonald, Uber COO
The vendors selling agents have a tidy answer: use more agents. Cognition CEO Scott Wu, whose company builds the autonomous coding agent Devin, has argued the fix for AI-generated bugs is more AI doing the cleanup. Wu himself rates Devin somewhere between a junior and mid-level engineer depending on the task, which is a long way from hand-it-off-and-forget. Cognition recently said Devin now writes 89% of its own internal code, a figure that lands differently after the 44% bug-fix stat.
The Singapore Management University team takes the opposite view. Developers, they argue, need to learn the failure modes of AI coding tools as deeply as they know their primary language, build QA systems designed for AI output, and review generated code with the scrutiny they'd apply to a junior hire. Architecture and security design should stay with humans. Wu agrees with the last part.
The awkward shape of the story is that all of these things can be true at once. AI coding tools feel indispensable to the developers using them. The same developers' employers are spending unprecedented sums on tokens without seeing the productivity bump in shipped output. And the code itself appears to carry a heavier long-term maintenance load than what it replaced. The METR refusal is the tell — when participants will not surrender a tool even for a controlled study, the tool has crossed from productivity aid into workflow dependency, and that's a different category.
For the AI coding market, the next twelve months are about whether the bug-fix tax is a transient artifact of model immaturity or a structural feature of probabilistic code generation. If it's the former, model releases like Claude Opus 4.8 will narrow the gap and the tokenmaxxing era will look like an early-adoption growing pain. If it's the latter, the winners will be the reliability and code-review layer — Entelligence AI, Code Rabbit, and whoever else can charge to clean up after the generators — and the buyers funding seven-figure agent budgets will start asking for output metrics the leaderboards were designed to obscure.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




