When Coding Output Doubles, Review Becomes the Constraint

Case Study·Giovanni Leonardi·August 2026·7 min read

Researched by an agentic pipeline · reviewed and gated by the author

A programme has not doubled productivity when it has only doubled the work arriving at its next constraint.

The target was simple

In June 2025, the chief technology officer of a mid-sized B2B software company set a clear goal: double engineering productivity through AI, measured as merged pull requests per engineer per month. The firm was unusually well placed to try. It was young, already AI-forward, paid for commercial coding tools without seat or token caps, and treated AI fluency as an organisational priority. Researchers from Carnegie Mellon University and Stanford University obtained anonymised pull-request, review and tool-usage data; no company employee participated in the analysis or writing. [S1]

The study followed 802 developers and 196,212 non-bot pull requests across 364 repositories from January 2024 to April 2026. Its estimation panel contained 564 developers observed for at least three active months: 451 AI adopters and 113 never-adopters. The tool environment included Cursor, Claude Code and newer internal agents. [S1]

By April 2026, the chosen metric had moved from 21.2 to 44.3 merged pull requests per active developer per month: 2.09 times the pre-mandate baseline. The target was met.

The mandate met its chosen output target by moving the constraint downstream.

What changed across the delivery system

Measure Observed change
Merged PRs per active developer 21.2 to 44.3 per month
Raw PR volume 3.1× the early-2025 baseline
Developers acting as reviewers 1.5×
PR load per reviewer 2.0×
PRs with human review 89% to 68%
PRs with substantive human comments about 39% to about 21%
PRs with automated review about 19% to about 84%
AI-authored PR total cycle time 22% longer after the mandate

The output gain was not an immediate effect of a management instruction. The researchers’ preferred within-developer model associated adoption and accumulated use with roughly a 1.5× gain at observed usage, rising towards 2× after nine months on tool. They used developer and calendar-month controls, staggered difference-in-differences, alternative estimators and placebo tests. [S1]

That is meaningful evidence, but not a clean causal estimate. Adoption and intensity were chosen, not random. Early adopters were already higher-output developers. One alternative estimator reduced the adoption estimate to a non-significant result. The mandate was a common date, so its direct effect could not be separated from other organisation-wide changes. The study supports a strong association between sustained use and throughput; it does not show that issuing a mandate causes output to double.

The gain was uneven

The result was broadly similar from individual contributors through principal engineers, but codebase age mattered. In post-2022 repositories, adoption was associated with 44% higher throughput. In pre-2022 repositories, the estimated gain was 12% and not statistically significant. [S1]

This matters for transferability. The case company was close to a favourable upper bound: young, permissive and willing to fund uncapped use. A mature enterprise with legacy systems, stricter controls or constrained budgets should not use 2.09× as its planning assumption.

A separate Microsoft study of 16,223 engineers found a smaller but directionally consistent pattern. Engineers completed an estimated 40.5% more pull requests in their highest GitHub Copilot-use weeks than in zero-use weeks while holding measured coding effort constant. The design was observational and depended on a conditional-independence assumption. [S5] Independent reporting on another Microsoft rollout found a 24.0% increase in merged pull requests over four months, with regular users gaining more than those merely given access. It did not measure software quality, customer value, security or long-term maintainability. [S2]

The wider evidence therefore supports usage-dependent throughput gains, not a universal multiplier.

Review absorbed the overflow

Raw pull-request volume grew 3.1× while the reviewer pool grew only 1.5×. Reviewer load consequently doubled. Human-review coverage fell by 21 percentage points, automated review rose to 84%, and the human review that remained became thinner: substantive written review fell from about 39% to about 21% of pull requests while silent approvals stayed broadly flat. [S1]

The company continued to ship. Merge rates stayed flat and revert rates declined. Those are useful signals, but they are coarse and short-term. They do not measure latent defects, security failures, architecture quality, maintainability, ownership or what an automated reviewer failed to notice.

The queue did not disappear either. For AI-authored pull requests, time from first human review to merge was about 20% longer and total cycle time was 22% longer after the mandate. The researchers found that the organisation kept pace mainly by routing more work through automated review and auto-approval, not by making substantive human review faster. [S1]

Separate open-source evidence points in the same direction. One study found that after GitHub Copilot’s introduction, experienced core developers reviewed 6.5% more code while their own original-code productivity fell 19%; the authors interpret this as maintenance burden shifting towards scarce experts. [S6]

Contradictory evidence is part of the result

METR’s 2025 randomised trial assigned 246 real issues from 16 experienced open-source maintainers to AI-allowed or AI-disallowed conditions. Developers took 19% longer when allowed to use AI, despite believing afterwards that AI had made them 20% faster. METR explicitly warned against generalising that narrow result to most developers or future tools. [S3]

Its 2026 follow-up could not provide a reliable updated estimate. Developers increasingly refused work that might require them to operate without AI, task selection changed, the pay rate fell, and parallel-agent use made time measurement harder. METR judged the resulting signal too biased to quantify current uplift. [S4]

These findings do not invalidate the enterprise case. They show that productivity depends on task, codebase, experience, workflow, adoption intensity and measurement design. A pull-request count, a task-completion time and customer value are different outcomes.

What the case proves

The evidence supports three bounded conclusions:

  • Under favourable conditions, sustained use of coding agents can coincide with a large increase in pull-request throughput.
  • The benefit takes time and complementary organisational change; access or mandate alone is not the productive unit.
  • Faster authoring can relocate work into review, assurance and maintenance rather than remove it.

It does not establish net financial return. The study reports no licence cost, token spend in currency, reviewer labour cost, defect cost, incident rate or total cost of ownership. It does not establish long-term quality, and it has not been replicated across comparable enterprises.

The decision test

Before setting an AI engineering target, define a balanced operating measure that includes output, review capacity, cycle time and quality. Track at least:

  • merged output per engineer, segmented by repository age;
  • queue depth and time from first human review to merge;
  • substantive human-review coverage, not merely an approval event;
  • automated-review coverage and escalation rates;
  • reverts, escaped defects, security findings and incidents;
  • licence, token, review and maintenance cost per accepted change.

Treat a throughput gain as economically real only if downstream assurance remains adequate and total delivery cost improves. A programme has not doubled productivity when it has only doubled the work arriving at its next constraint.

Sources

  1. He et al. — AI Writes Faster Than Humans Can Review: A Longitudinal Study of an Enterprise “2×” Mandate — 2 July 2026 — https://arxiv.org/html/2607.01904v1
  2. TechRepublic — Microsoft Study Finds AI Coding Agents Lift Pull Requests by 24% — 13 July 2026 — https://www.techrepublic.com/article/news-ai-coding-agents-microsoft-pull-requests/
  3. METR — Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity — 10 July 2025 — https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
  4. METR — We are Changing our Developer Productivity Experiment Design — 24 February 2026 — https://metr.org/blog/2026-02-24-uplift-update/
  5. Heilman et al. — GitHub Copilot and Developer Productivity: An Observational Dose-Response Analysis — 30 May 2026 — https://arxiv.org/abs/2606.00438
  6. Xu et al. — AI-Assisted Programming Decreases the Productivity of Experienced Developers by Increasing the Technical Debt and Maintenance Burden — 28 January 2026 — https://arxiv.org/abs/2510.10165

Available on request — write me


More from Programme