The AI coding slowdown: what METR's controlled trial measured
Series: The Measured View ยท 29 July 2026
The demo convinced everyone. The pilot goes well, the team adopts the coding assistant, and within a month everyone reports the same thing: this is saving us serious time. Then someone asks the uncomfortable question. Saving time compared to what? Nobody ran a control. Nobody timed anything. The entire business case rests on how the work felt.
Most organisations measure AI productivity exactly this way: self-reports, surveys, sentiment, anecdotes from power users. Feelings about speed are a specific kind of measurement. They capture effort, novelty and enjoyment, and they are weak at capturing elapsed time. In 2025, one research group decided to time it properly.
METR recruited 16 experienced open-source developers and gave them 246 real tasks in mature codebases they had worked in for five years on average. Each task was randomly assigned: AI tools allowed, or not. Before starting, the developers forecast that AI would make them 24 percent faster. Afterwards, they estimated it had made them 20 percent faster. The stopwatch said otherwise: allowing AI made them 19 percent slower.
Effect of allowing AI tools on task completion time
Why slower? The developers spent real time prompting, waiting, and cleaning up generated code. In codebases they knew intimately, with high quality bars, their own hands were often quicker. The setting matters: experts, mature million-line repositories, strict standards. AI assistance shines brightest where the developer knows least, and here the developers knew almost everything. The result comes with its boundaries stated: one study, 16 developers, using Cursor Pro and Claude 3.5 and 3.7 Sonnet in mid-2025. It is robust within its setting, and it is a data point about a specific frontier, a moment in time, and a type of work, with a year of tool progress between that setting and today.
The gap is the finding. Even after living through the slowdown, the developers still believed AI had sped them up by 20 percent. If skilled professionals misjudge their own productivity by 39 percentage points, every self-reported AI productivity number in your organisation deserves scrutiny. Adoption surveys tell you about satisfaction. Only timing tells you about time.
Two rules follow for any rollout. Measure time, not vibes: pick a handful of recurring tasks and time them with and without AI, on real work, with the people who normally do them; a modest sample beats an enthusiastic survey. And match the tool to the terrain: expect the largest gains where familiarity is low, new codebases, new domains, boilerplate, and the smallest gains, or losses, where your experts are already at home.
The feeling of speed is not speed. METR's trial says less about AI's ceiling and more about our habit of grading it by sentiment. Companies that measure will find the real gains. Companies that trust the feeling will fund an illusion.
The paper: Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity