The Acceleration Tax

Practice Brief·Giovanni Leonardi·October 2026·7 min read

Researched by an agentic pipeline · reviewed and gated by the author

Organisations investing in AI coding tools without proportional investment in review capacity — human, automated, or both — are accumulating unreviewed code in production.

The Trade Nobody Announced

Every AI coding throughput claim in this evidence base is accurate. So is the quality collapse that accompanies it.

Faros AI’s telemetry across 22,000 developers and 4,000 teams — the largest engineering-operations dataset published on AI tool adoption — documents both sides in the same data [S1]. Developers in the highest AI adoption quarters completed 66% more epics and 34% more tasks per head. They also produced 54% more bugs, triggered three times the production incidents per pull request, and generated code churn nearly ten times above baseline.

Three independent studies, using different methodologies and samples, corroborate the pattern. The throughput is real. The downstream cost is also real. And the downstream cost does not appear in the metrics most organisations use to judge whether AI coding tools are working.

What 22,000 Developers Show

Faros collected two years of engineering telemetry spanning task management, version control, CI/CD pipelines and incident management systems from over 4,000 teams. They compared each organisation’s lowest and highest AI adoption quarters — defined as the point at which more than 50% of developers used AI tools weekly — and reported only statistically significant correlations standardised per company, with a minimum of six companies per metric [S1][S2].

The throughput gains are unambiguous. Epics per developer rose 66.2%. Task throughput per developer rose 33.7%. Pull request merge rate rose 16.2%. Code-focused tasks surged 210%, confirming AI’s strongest effect on code generation specifically [S4].

The quality deterioration is equally clear. Bugs per developer rose 54%, up from 9% in 2025, meaning the gap is widening [S3]. The incidents-to-pull-request ratio jumped 242.7%. Monthly incidents rose 57.9%. Pull request size increased 51.3% and files edited per request rose 59.7%, producing changes harder to review, harder to roll back, and harder to contain when they fail.

Code churn — the proportion of code rewritten within two weeks of being committed — rose 861%. Faros acknowledges the ambiguity: this may reflect rework of insufficient AI-generated code, accelerated refactoring, or faster iteration cycles. The figure cannot be decomposed further in cross-customer telemetry [S3]. What it establishes is that nearly ten times more code failed to survive first contact with production.

The Review Bottleneck

The pipeline data describes a system generating more work than its downstream processes can absorb.

Median time to first pull request review rose 157%. Average code review time rose 200%. Median review duration rose 442% — roughly fivefold [S2][S3]. Senior engineers absorbed the burden disproportionately, spending escalating hours on code that reads plausibly but fails structurally.

The practical response has been to skip review. Pull requests merged without review increased 31.3% [S1]. Faros’s interpretation: reviewers cannot keep pace with the volume of AI-generated code [S3].

Commit-to-production lead time increased 480%. Weekly deployments fell 12% in the measured subset [S2]. More code was written. Less of it reached production faster. Developers reported 67% more daily pull request contexts, 14% more task restarts, and 26% more tasks stalled inactive for seven or more days [S3]. Starting work became easier. Finishing it became harder.

Three Corroborating Studies

The Faros findings do not stand alone.

CodeRabbit analysed 470 open-source GitHub pull requests — 320 AI-co-authored, 150 human-only — and found AI-authored code carried 1.7 times more issues per pull request [S5]. The multiplier varied by category: performance-related issues appeared at eight times the human rate; readability problems at three times; security issues at 2.7 times. Authorship attribution relied on signals rather than direct confirmation.

GitClear examined 623 million code changes across 2023–2026 and documented a structural shift in how code is written [S6]. Refactoring activity fell 70%. Cross-file function calls dropped 35%. Code block duplication rose 81%. Error-masking constructs rose 47%. Developers became five times more likely to copy and paste than to refactor — a pattern consistent with tools optimised for producing new code rather than maintaining existing systems.

New Relic and Hanover Research surveyed 200 US technology decision-makers and found 78% reporting measurable production incident spikes tied to AI-generated code [S7]. Eighty-two percent had suffered a major production failure from AI code within six months. The study is a perception survey, not telemetry, but the perception is near-universal among the respondent population.

What This Does Not Prove

Every source in this evidence base has a commercial interest in the finding that organisations need better measurement of AI-generated code. Faros sells engineering analytics. CodeRabbit sells AI-assisted code review. GitClear sells code quality measurement. New Relic sells observability. This does not invalidate the findings, but no source is disinterested [S1][S5][S6][S7].

All results are correlational. No study establishes that AI tools cause quality degradation. The relationship could reflect adoption patterns, team composition, project complexity, or measurement artefacts that correlate with both AI uptake and quality metrics.

No independent academic replication exists. The CodeRabbit sample is 470 open-source pull requests with uncertain authorship attribution. The New Relic study is 200 survey respondents. The Faros datasets are independent cross-sections, not longitudinal panels tracking the same teams over time [S4]. Whether structured mitigation practices — tighter pull request controls, earlier quality gates, mandatory review enforcement — eliminate the quality gap remains untested at scale.

The Maturity Question

One finding deserves separate attention. Faros reports that mature DevOps organisations showed the same downstream quality deterioration as less mature peers [S3]. This contradicts the DORA 2025 survey’s conclusion that AI amplifies existing organisational capabilities.

If the finding holds, the implication is significant: existing engineering maturity does not absorb the quality cost of AI-generated code. The investment required may be additional — new kinds of quality gates, new review tooling, new measurement — rather than a return on what organisations have already built.

What a Programme Leader Should Conclude

The evidence supports three bounded conclusions.

First, throughput gains from AI coding tools are real and measurable, but they are gross figures. The net gain — after accounting for review burden, rework, incident response, and code churn — is unknown. No study yet published measures it. Any headcount, capacity, or delivery-schedule decision based on gross throughput is built on an incomplete number.

Second, the review pipeline is the binding constraint. The data consistently shows code review as the stage where AI-generated volume overwhelms human capacity, producing either delays or skipped review. Organisations investing in AI coding tools without proportional investment in review capacity — human, automated, or both — are accumulating unreviewed code in production.

Third, the evidence base is commercially interested, correlational, and early. It is also the only enterprise-scale empirical evidence available. Ignoring it because it is imperfect is no more rational than treating it as definitive. The prudent response is to instrument: measure quality alongside velocity, track the review pipeline as a capacity constraint, and make adoption decisions on net figures rather than gross.

The 861% code churn figure is the asterisk on every throughput number in this evidence base. It measures what failed to survive — and no headline velocity metric accounts for it.

Sources

  1. Faros AI — AI Engineering Report 2026: The Acceleration Whiplash (PDF) — April 2026 — https://pages.faros.ai/hubfs/AI_Engineering_Report_2026_The_Acceleration_Whiplash_Faros.pdf
  2. Faros AI — AI Impact on Engineering Productivity: 2026 Report Data — 2026 — https://www.faros.ai/research/ai-acceleration-whiplash
  3. Faros AI — Ten Takeaways from the Acceleration Whiplash — 2026 — https://www.faros.ai/blog/ai-acceleration-whiplash-takeaways
  4. ADTmag — More Code, More Bugs — 22 April 2026 — https://adtmag.com/articles/2026/04/22/more-code-more-bugs.aspx
  5. CodeRabbit — State of AI vs Human Code Generation Report — 17 December 2025 — https://www.coderabbit.ai/blog/state-of-ai-vs-human-code-generation-report
  6. GitClear — The Maintainability Gap: 2026 AI Code Quality Research — 2026 — https://www.gitclear.com/the_ai_code_quality_maintainability_gap
  7. New Relic / Hanover Research — State of AI Coding 2026 — 10 June 2026 — https://newrelic.com/blog/ai/state-of-ai-coding-2026