
A preregistered trial, not a survey
On 13 July 2023, Science published a preregistered online experiment by MIT economists Shakked Noy and Whitney Zhang measuring what happens when college-educated professionals are given access to ChatGPT for real writing work. The working paper describes recruiting 453 experienced professionals through the survey platform Prolific between 27 January and 21 February 2023, drawn from six occupations: marketers, grant writers, consultants, data analysts, human resource professionals and managers. Each participant completed two 20-to-30-minute writing tasks resembling their actual work, such as press releases, short reports, analysis plans and delicate emails. Half were randomly assigned to register for ChatGPT before the second task; the control group registered for the LaTeX editor Overleaf instead, to hold signup friction constant. Output was graded blind, on a 1-to-7 scale, by three experienced evaluators per piece.
What moved, and by how much
The treatment group finished the second task 11 minutes faster than the control group's 27-minute average, a 0.75 standard-deviation shift, and their evaluator grades rose by 0.45 standard deviations. Both the time and quality distributions moved almost entirely, not just at the tails. Inequality between workers narrowed: in the control group, a worker's score on the first task predicted their second-task score at a correlation of 0.41; with ChatGPT access that correlation fell to 0.14, because participants who scored lowest on the first task gained the most. The paper's own tracking of participants' screens shows why: a third submitted ChatGPT's output essentially unedited, and most of the rest spent only a few minutes revising it.
What the design does not establish
The authors are explicit about the limits of their own evidence. The tasks demanded generic, self-contained writing without any requirement for company-specific knowledge or factual verification, exactly ChatGPT's strengths and exactly what many real jobs also require checking. Participants were paid per quality point, not managed through the promotion and reputation incentives that shape how professionals actually write. A two-week follow-up found 34% of the former treatment group using ChatGPT in their real job, against 18% of controls, but users rated it less useful there (3.66 out of 5) than in the experiment, citing a lack of company-specific context. The paper measures a one-off task under strong incentives, not sustained job performance or what happens once novelty wears off.
Questions to ask before you adopt it
- Does the task require company-specific knowledge or fact-checking the model cannot supply on its own?
- Are the people using it paid for speed and generic quality, or for judgement that a compressed draft cannot substitute for?
- Would you expect the same gain to persist after the first few weeks of novelty, or does the evidence only cover a first encounter?
Noy and Zhang's experiment is one of the cleanest randomised measurements of AI-assisted writing available, and its finding that lower performers gained the most is echoed elsewhere, including in the NBER paper on generative AI in customer support. That does not make it a general productivity multiplier for every writing job; it is evidence about a specific set of short, incentivised, fact-light tasks completed by strangers recruited online.
Sources & reading trail
Full working-paper text describing the 453-participant Prolific experiment, task types, treatment/control design, and the 40% time and 0.45 SD quality effects.
Source published: Not established · Retrieved: 16 September 2026
Independently cites Noy and Zhang's finding that ChatGPT compresses the productivity distribution, corroborating the inequality result from a different, real-workplace setting.
Source published: 1 April 2023 · Retrieved: 16 September 2026
Announcements and papers establish the record; the friction reading and the adoption questions are Productivity Atlas editorial analysis. This retrospective draft does not imply the site published on the event date.