
An experiment run inside a real consultancy
On 21 September 2023, Boston Consulting Group published an account of a field experiment run with Harvard, Wharton and MIT researchers on 758 of its consultants worldwide, mostly junior staff with up to four years' experience. Participants were split across three conditions: no AI access, GPT-4 access, or GPT-4 access plus training, and given tasks in two categories. The first, a creative product-development exercise, sat inside what the researchers judged GPT-4 to handle well. The second, a business problem requiring interpretation of interview and financial data to make a recommendation, sat outside that competence. The MIT Sloan review of the same study, published a month later, calls this boundary the “jagged technological frontier”: a line, invisible to users in advance, separating tasks a model handles well from tasks it does not, that does not track how difficult a task looks to a human.
Gains on one side, losses on the other
On the task inside the frontier, about 90% of AI-assisted consultants improved, converging on output roughly 40% higher in quality than the unassisted group, with the largest gains among consultants who had scored lowest without AI (43%, against 17% for top performers). On the task outside the frontier, AI-assisted consultants performed roughly 19 percentage points worse than the unassisted group. BCG reports that participants often accepted GPT-4's confidently wrong output rather than checking it, and that training on how to use the model did not prevent this: trained participants did no better, and sometimes worse, on the out-of-frontier task than untrained ones.
What the boundary does not tell you in advance
The frontier is defined after the fact, by comparing task performance with and without the tool; it is not something a consultant, or a manager assigning work, can reliably see before starting. BCG also found that idea diversity among AI users fell by 41% on the creative task even as average quality rose, a genuine editorial trade-off between consistency and variety, not just a limitation of the study. The sample was junior BCG consultants using GPT-4 on two specific task types in 2023; a different model, a different seniority mix or a different task will sit differently against its own frontier, and the paper does not claim to have mapped that boundary for any task outside the two it tested.
Questions to ask before you adopt it
- Has anyone actually tested whether this specific task sits inside or outside the model's competence, or is that being assumed?
- Does your review process catch confidently wrong output, given that trained users in this study often did not?
- Are you measuring output diversity as well as average quality, since the two moved in opposite directions here?
“Jagged frontier” is a useful name for a real pattern: capability that is uneven in ways that do not match human intuition about difficulty. This is an editorial framework drawn from one field experiment, not a settled map of what any given AI tool can or cannot do.
Sources & reading trail
BCG's own account of the 758-consultant experiment, the task split between creative product innovation and business problem-solving, and the percentage gains, losses and diversity effects.
Source published: 21 September 2023 · Retrieved: 16 September 2026
Independent MIT Sloan summary of the same experiment describing the two task categories, the naming of the 'jagged technological frontier', and the skill-level breakdown of gains.
Source published: 19 October 2023 · Retrieved: 16 September 2026
Announcements and papers establish the record; the friction reading and the adoption questions are Productivity Atlas editorial analysis. This retrospective draft does not imply the site published on the event date.