
Two field experiments, not lab tasks
Late 2024 brought two field experiments measuring AI coding assistance on real work rather than a single scripted task. The first, described in a MIT Sloan account of a paper by Mert Demirer and colleagues at Microsoft, Princeton, Pennsylvania and MIT, tracked staggered GitHub Copilot rollouts across three employers: Microsoft, over seven months; Accenture, over four months; and an anonymous Fortune 100 electronics manufacturer, over two months with a six-week staggered access window. Because access rolled out to different developers at different times within each company, the researchers could compare otherwise similar developers before and after they gained the tool. The second, a randomised controlled trial posted to arXiv on 16 October 2024, took a stricter approach at a single employer: 96 full-time Google software engineers were randomly assigned, task by task, to use or not use internal AI coding features on one complex, enterprise-grade assignment.
What each measured
Across the three-company study, Copilot access raised completed weekly tasks by an average of 26%, concentrated heavily among newer hires (27 to 39% gains) against much smaller gains for senior developers (8 to 13%), with roughly 60% of eligible developers still using the tool after a year. In the Google trial, engineers assigned to use the AI features completed their task about 21% faster on average, with larger gains for those who spent more of their working day on code. The two studies differ in the strength of their design: the three-company study relies on a staggered natural rollout with before-and-after comparison, while the Google study randomly assigns the treatment task by task, a stronger basis for a causal claim within that one company.
What neither result generalises to
Both papers are explicit about scope. The Google trial's authors caution that their effect size need not apply more broadly, since it comes from one internal tool, tested in one lab-like arrangement, over a single summer. The three-company study covers only Copilot, only developers, and only weekly task counts as its productivity measure, not code quality or long-term skill development, and the effect varied enough by company and tenure that a single average obscures more than it reveals. Neither study says anything about non-coding knowledge work, and the consistent pattern across both, and across the customer-support and consulting studies elsewhere in this evidence set, is that junior or less-experienced staff gain the most.
Questions to ask before you adopt it
- Is the comparison a randomised assignment or a before-and-after look at a staggered rollout, and does that change how much weight the result deserves?
- Is the productivity measure task count, completion time, or something else, and does that map to what you actually want to improve?
- If gains concentrate in junior staff, does your rollout and training plan reflect that, rather than assuming uniform benefit?
Two employers' worth of real, measured coding work, using two different designs, point the same direction: meaningful average gains that are smaller for the most experienced people doing the work. That pattern is worth taking seriously without treating either single-company figure as a portable productivity rate.
Sources & reading trail
Full abstract of a randomised trial with 96 Google engineers testing internal AI coding features, finding a roughly 21% reduction in completion time with an explicit caveat against generalising the effect size.
Source published: 16 October 2024 · Retrieved: 16 September 2026
MIT Sloan's account of a separate staggered-rollout field study of GitHub Copilot across Microsoft, Accenture and an anonymous Fortune 100 manufacturer, naming the paper, its authors, and its per-company duration and results.
Source published: 4 November 2024 · Retrieved: 16 September 2026
Announcements and papers establish the record; the friction reading and the adoption questions are Productivity Atlas editorial analysis. This retrospective draft does not imply the site published on the event date.