Measure the work that ships, including review and rework.
Keep the result attached to its setting
The original Atlas edition discusses METR's early-2025 study of experienced developers working in familiar open-source repositories. It is useful evidence precisely because the setting is specific. Turning that finding into “AI makes all developers slower” would discard the conditions that made it interpretable.
The 2026 update is a measurement warning
In February 2026, METR reported that a newer experiment had serious selection and measurement problems. Developers increasingly declined participation or withheld tasks they did not want to attempt without AI. Concurrent agent use also complicated recording time. METR considered greater acceleration plausible, but described the new data as weak evidence for its magnitude. It did not retract the earlier finding. Read the update and uncertainty discussion.
The lesson for a small team is methodological: a benchmark number is not a prediction for your backlog. Your projects, developers and definition of “done” may differ from a study's sample.
Compare whole tasks, not generated lines
Our suggested pilot records a task before work starts: acceptance criteria, relevant tests and the expected review standard. Track implementation, waiting, review, debugging and later corrections separately. Record abandoned attempts too. Otherwise the comparison can silently discard the difficult cases.
Use a varied sample of ordinary work rather than selecting only tasks that seem ideal for an assistant. Separate unfamiliar exploration from maintenance in familiar code. Small informal pilots will not establish a universal causal effect, but they can expose practical costs your team would otherwise miss.
Code review is another workflow to evaluate
GitHub documents Copilot code review as a system for feedback and suggested fixes. Its documentation also describes usage costs and configurable capabilities. A review response is not evidence that every relevant defect has been found. Inspect the feature and its limits.
Try the reviewer against changes with known issues and benign changes. Inspect missed problems and distracting comments. Keep established tests and human ownership of consequential decisions. A faster suggestion loop only matters if it improves the finished change at an acceptable cost.
Decide what would change your mind
Before a pilot, agree on adoption, restriction and rejection criteria. Perhaps documentation tasks benefit while risky migrations do not. That is a useful result, not a failed experiment. The Atlas recommendation is to buy a workflow you can evaluate, not to buy a universal productivity claim.
Source receipts
Primary pages checked for the stated product facts or evidence. The proposed pilots and buying questions are Atlas analysis, not reported study results.
METR · Published 2026-02-24 · Checked 2026-09-16
GitHub Docs · Checked 2026-09-16