RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The archive · 100 retrospective records ↗
productivityatlas.

The archive / Evidence

Evidence / From the archive · 10 July 2025 event · prepared 16 September 2026

Experienced developers were slower with AI and did not notice

A randomised trial of 16 open-source developers found AI tools added time despite predictions of speed-up.

Visual for this record: Experienced developers were slower with AI and did not notice
Visual published by substackcdn.com, shown for identification of the record. Credit: substackcdn.com · source page ↗ Rights: owner-review-pending.

A randomised trial, not a benchmark

On 10 July 2025, the AI evaluations group METR published results from a randomised controlled trial measuring how AI coding tools affected experienced developers working on real, familiar codebases. Sixteen developers with substantial history contributing to large open-source repositories completed 246 real issues from their own projects, each issue randomly assigned to allow or disallow AI assistance. The full paper describes the tools used as frontier models of the time, mainly Cursor Pro with Claude 3.5 and 3.7 Sonnet, and the tasks as ordinary maintenance and feature work the developers would have done regardless of the study.

What was measured, and the gap between belief and result

Before starting, developers predicted AI assistance would cut their completion time by 24%. Economists and machine-learning researchers surveyed separately predicted even larger gains, of 38 to 39%. The measured result ran the other way: allowing AI use increased completion time by 19%. After finishing, developers on average still believed AI had sped them up, by about 20%, even when the clock showed otherwise. METR investigated 20 candidate explanations for the slowdown, including unfamiliarity with the tools and low code-quality standards, and reports the effect held up as a robust feature of this setting rather than an artefact of any single factor it tested.

What the trial does not claim

METR states directly that the result does not show AI systems fail to speed up developers generally. The sample was sixteen highly experienced maintainers working in codebases they already knew well, using specific tools available in early-to-mid 2025; the paper does not claim the finding extends to less experienced developers, unfamiliar codebases, greenfield projects, or later model versions. The mismatch between perceived and measured speed is the most transferable part of the finding, since it suggests that self-reports of AI-driven speed-up, on their own, may not be a reliable guide to whether a team is actually faster. It says nothing about code quality, maintainability, or outcomes beyond time to completion.

Questions to ask before you adopt it

  • Is the comparison a clock measurement or a self-reported impression of speed?
  • How similar are the codebase, developer experience level and tool version to the ones this trial actually tested?
  • What would change your mind if the tool felt faster but a time-tracked pilot showed it was not?

A randomised trial with real tasks and a real time measurement is unusually strong evidence for its own setting, and this one directly contradicts what its own participants believed about themselves. That combination is exactly why it should not be read as a verdict on AI coding tools everywhere; it is a caution about trusting felt speed over measured speed, in the specific setting the trial covered.

Sources & reading trail

Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ↗

METR's own account of the RCT design, the 16 developers, 246 tasks, the 19% slowdown, and METR's explicit statement of what the finding does not generalise to.

Source published: 10 July 2025 · Retrieved: 16 September 2026

Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ↗

Full paper text with the pre-registered design, the 24% predicted versus 19% actual slowdown, and the 20-factor robustness analysis.

Source published: 12 July 2025 · Retrieved: 16 September 2026

Announcements and papers establish the record; the friction reading and the adoption questions are Productivity Atlas editorial analysis. This retrospective draft does not imply the site published on the event date.