The report’s observations, prices and forecasts are a dated snapshot. Examples of savings are estimates unless explicitly identified as study results. This is not a fresh verification of every claim.
Evidence table
| Study | Design and population | Finding | What it does not prove |
|---|---|---|---|
| Brynjolfsson, Li, Raymond, QJE / NBER | Staggered field rollout to 5,179 customer-support agents | About 14% more issues resolved per hour on average; roughly 34% improvement for novice and low-skilled workers; little effect for experienced/high-skilled workers | General knowledge work, long-term labor effects, or every chatbot product |
| Dell’Acqua et al., HBS/BCG, Organization Science | Preregistered experiment with 758 consultants and 18 realistic tasks | Inside the AI frontier: 12.2% more tasks, 25.1% faster, and higher quality; outside the frontier, AI could hurt performance | Safe delegation of arbitrary professional judgment |
| Dillon et al., Microsoft Research | Six-month randomized field experiment with about 6,000 knowledge workers using AI in email, documents, and meetings | Users spent around three fewer hours per week on email; documents appeared moderately faster; meeting time did not significantly change | Organization-wide productivity or whether saved time became leisure rather than more output |
| Cui et al., Management Science | Three company-run randomized field experiments at Microsoft, Accenture, and a Fortune 100 firm; 4,867 developers | Combined 26.08% increase in completed tasks; effects varied and less experienced developers adopted more | Every codebase, seniority group, or quality/security outcome |
| Becker et al., METR | Randomized trial with 16 experienced open-source developers and 246 tasks using early-2025 tools | Developers estimated a 20% speedup, but measured completion time increased 19% in this setting | Current models, novice developers, or simpler greenfield tasks |
Interpretation
The strongest evidence clusters around tasks with clear inputs, repeated patterns, fast feedback, and a measurable output: customer-support responses, code completion, email drafting, document transformation, structured research, and meeting capture. Gains are less certain when the task requires deep domain context, ambiguous goals, novel judgment, or long-horizon coordination.
Time saved can be redistributed into more output, higher quality, learning, or more meetings. The website should therefore report at least four separate metrics:
- Cycle time: time from request to acceptable completion.
- Throughput: number of acceptable outcomes per person or team.
- Quality: rubric score, error rate, rework, or customer outcome.
- Human cost: review time, context switching, cognitive load, and escalation burden.
An AI tool that cuts drafting time by 30% but adds 20 minutes of checking, or produces 10% more output that no one wants, has not delivered the claimed 30% productivity gain.
Pilot measurement protocol
- Define the baseline task, input quality, cycle time, acceptable-quality rubric, and error cost.
- Run a control period without the AI feature and a treatment period with it.
- Record user time, AI time, review time, rework, escalations, and exceptions.
- Segment by user experience, task type, language, and data complexity.
- Review a random sample for factuality, completeness, tone, security, and policy compliance.
- Measure whether the “saved” time changed output, focus, learning, service level, or workload.
- Stop or redesign the pilot if errors or review costs erase the benefit.