Developers said AI made them faster. A controlled trial says they were slower.

Inside the METR study that measured the gap between how fast AI feels and how fast it is.

Malte Clausen · 15 August 2026 · The Measured View

The demo convinced everyone

The pilot goes well. The team adopts the coding assistant. Within a month, everyone reports the same thing: this is saving us serious time.

Then someone asks the uncomfortable question. Saving time compared to what? Nobody ran a control. Nobody timed anything. The entire business case rests on how the work felt.

We grade AI by feel

Most organisations measure AI productivity through self-reports: surveys, sentiment, anecdotes from power users.

Feelings about speed are a specific kind of measurement. They capture effort, novelty and enjoyment. They are weak at capturing elapsed time.

In 2025, one research group decided to time it properly.

A randomised trial, on real work

METR recruited 16 experienced open-source developers and gave them 246 real tasks in mature codebases they had worked in for five years on average. Each task was randomly assigned: AI tools allowed, or not.

Before starting, developers forecast AI would make them 24 percent faster. Afterwards, they estimated it had made them 20 percent faster. The measured result: allowing AI made them 19 percent slower.

Effect of allowing AI tools on task completion time

What developers forecast
24% faster
What developers felt afterwards
20% faster
What the stopwatch measured
19% slower

Why slower

The developers spent real time prompting, waiting, and cleaning up generated code. In codebases they knew intimately, with high quality bars, their own hands were often quicker.

The setting matters: experts, mature million-line repositories, strict standards. AI assistance shines brightest where the developer knows least. Here, the developers knew almost everything.

The gap is the finding

Read it again: even after living through the slowdown, developers still believed AI had sped them up by 20 percent.

If skilled professionals misjudge their own productivity by 39 percentage points, every self-reported AI productivity number in your organisation deserves scrutiny.

The lesson is about measurement discipline. Adoption surveys tell you about satisfaction. Only timing tells you about time.

Two rules for your rollout

Measure time, not impressions. Pick a handful of recurring tasks. Time them with and without AI, on real work, with the people who normally do them. A modest sample beats an enthusiastic survey.

Match the tool to the terrain. Expect the largest gains where familiarity is low: new codebases, new domains, boilerplate. Expect the smallest gains, or losses, where your experts are already at home.

The feeling of speed is not speed

METR's trial says less about AI's ceiling and more about our habit of grading it by sentiment. Perception ran 39 points ahead of the stopwatch.

Companies that measure will find the real gains. Companies that trust the feeling will fund an illusion.

One study, 16 developers, using Cursor Pro and Claude 3.5 and 3.7 Sonnet in mid-2025. The result is robust within its setting. It is a data point about a specific frontier, a moment in time, and a type of work. A year of tool progress separates that setting from today's.

METR. Becker, Rush, Barnes and Rein (2025). Measuring the impact of early-2025 AI on experienced open-source developer productivity. arXiv:2507.09089

Back to The Measured View