The Early Warning Paradox

Analysis·Giovanni Leonardi·September 2026·11 min read

Researched by an agentic pipeline · reviewed and gated by the author

The humans were in the loop. They stopped looking.

The Quiet Erosion

In July 2026, the UK’s National Infrastructure and Service Transformation Authority published the annual health check for the government’s major projects portfolio: 189 projects, £924 billion in whole-life cost, 34 rated Red [S1]. Buried in the same reporting cycle was a less visible change. NISTA had deployed an AI-powered Early Warning System across the portfolio, ingesting the existing Government Major Projects Portfolio data to flag projects at risk of sliding to a Red rating [S1]. The system was live, embedded in review processes, and offered as evidence of a maturing assurance capability.

The difficulty is not the ambition. It is the data. The same GMPP data feeding the AI has been documented by the National Audit Office as persistently optimistic for over a decade. The government’s own analysis acknowledges that project data may be “incomplete, overly optimistic, differently defined between departments or updated too slowly” [S3]. The mandatory data standard that might address this does not take effect until 2027, with full compliance not required until 2030 [S1]. The AI arrived first.

This is the governance problem that matters: not whether AI can improve programme assurance — it plausibly can — but what happens in the interim when algorithmically confident assessments are built on structurally unreliable foundations.

A decade of documented optimism

The UK’s data quality problem is not a discovery. The NAO established in 2013 that optimism bias “persists, frequently undermining projects’ value for money as time and cost are under estimated and benefits over estimated” [S4]. A year earlier, the NAO had found the assurance system “not yet built to last,” with HM Treasury not routinely using portfolio data and departments failing to implement integrated assurance [S5]. By 2016, the NAO was reporting that one-third of projects due within five years were rated Red or amber-Red, with no “clear, consistent data with which to measure performance.”

The pattern continued into the AI era. The Public Accounts Committee’s September 2025 report found that NISTA itself described its data standard as a “minimum viable product” and acknowledged “more work to do in terms of standardisation of the data across projects and Departments” [S6]. The same committee, examining government AI use in March 2025, found that 62% of government bodies identified “access to good-quality data” as a barrier to AI adoption [S7].

The Government Major Projects Evaluation Review completed in 2025 found that only 34% of portfolio projects had robust evaluation plans — meaning 66%, representing £456 billion in total cost, lacked adequate evaluation frameworks [S8]. This is the evidential base on which the AI early warning system now operates.

Three simultaneous changes

The AI deployment did not happen in isolation. Three structural changes converged between late 2025 and mid-2026.

The Early Warning System went live, processing existing GMPP performance data — cost, schedule, milestones, risk, delivery confidence — to identify projects at risk of deterioration [S1]. Separately, the open-source Scout tool, built by i.AI for document analysis, entered beta with a claimed ability to halve assurance review preparation time and be “over 90% accurate in mimicking human judgement to establish lines of enquiry” [S2].

The portfolio was restructured from 189 projects to approximately 81 under central NISTA oversight, effective April 2026 [S1]. The remainder moved to a Departmental Major Projects Portfolio with strengthened but department-led assurance. Three programmes exceeding £10 billion — HS2, Sizewell C and Dreadnought — were classified as mega-projects.

The Government Reporting Integration Platform became mandatory from April 2026, and the Programme and Project Data Standard was published for trial in December 2025, with mandatory compliance from early 2027 and full adoption by 2030.

Each change has a defensible rationale. Together, they create a specific condition: AI producing confident outputs from data documented as unreliable, while the organisational structure simultaneously moves a majority of projects to assurance regimes led by the departments whose self-reporting produced that data.

How confident outputs displace scrutiny

The mechanism connecting AI assurance to governance risk is not dramatic failure. It is quiet displacement, operating through three reinforcing channels.

The first is data-confidence decoupling. AI systems processing programme data produce outputs with a computational precision that is unrelated to the reliability of the underlying data. When project teams report optimistically — as a decade of audit evidence confirms they do — the AI assessment inherits the bias but presents it with algorithmic authority. The bias becomes harder to detect than in a narrative human assessment precisely because it looks rigorous.

The second is attention displacement. When an AI system identifies which projects require attention, projects it does not flag receive less. If the system’s training data reflects the same optimistic reporting, it may systematically miss projects where reported data conceals deterioration — the projects where human review would add most value.

The third is institutional incentive erosion. Maintaining human assurance processes alongside AI is expensive and politically difficult to justify when the AI appears to work. The erosion is rarely an explicit decision. It is gradual resource reallocation: fewer reviewers per project, shorter review cycles, delegated oversight — each individually reasonable, collectively corrosive.

These channels operate together. Biased data produces confident assessments. Confident assessments redirect attention. Redirected attention reduces the human scrutiny that would catch what the data misses. Reduced scrutiny weakens the case for maintaining expensive human assurance. The cycle is self-reinforcing and largely invisible until a project fails in a way that retrospective review traces back to the gap.

The strongest case for the defence

The sceptical reading of this deployment deserves serious engagement, because it is strong.

The government has designed the AI system with explicit institutional humility. It triggers human review rather than automated decisions [S2]. Data limitations are openly acknowledged in official publications [S3]. The Scout tool is open-source, MIT-licensed and subject to external technical scrutiny. Its accuracy claim — 90% in mimicking human lines of enquiry — is honestly scoped; it is a productivity metric for document preparation, not a claim about predicting project outcomes [S2].

The portfolio restructuring concentrates central expertise on 81 high-priority projects rather than spreading it across 189. Twenty-six projects successfully exited the portfolio in 2025-26, up from 14, suggesting the broader reform programme is improving delivery [S1]. The data standard, GRIP and evaluation reforms represent a coherent programme to fix inputs — and the AI deployment provides immediate operational value while the slower institutional reforms take effect.

This is not naive automation. It is a considered deployment that openly acknowledges its constraints.

Where the defence is fragile

The fragility is not in the current design. It is in the assumption that the design will hold under institutional pressure over time.

Automation bias research from Georgetown University’s Center for Security and Emerging Technology documents a consistent pattern across aviation, military systems and civilian technology: humans reduce vigilance and defer to automated outputs even when human-in-the-loop protocols are formally maintained [S9]. The USS Vincennes incident, Boeing 737 Max crashes, and Tesla Autopilot fatalities all occurred in systems where human oversight was structurally present and procedurally required. The humans were in the loop. They stopped looking.

Programme assurance operates on slower timescales than aviation or driving, which makes the bias harder to detect but no less consequential. The gap between an AI assessment and a project outcome may be measured in years, not seconds. A gradually eroding review process does not produce a visible incident — it produces a portfolio where problems surface later, cost more, and are harder to attribute to any single oversight failure.

The sequencing gap compounds the risk. The AI is live now. The mandatory data standard does not take full effect until 2030. During this four-year interval, the system operates in precisely the conditions that historical audit evidence has documented as unreliable. The government is aware of this, which is why the data reforms exist. But the reforms follow the deployment rather than preceding it, and the evidence for optimism about the reforms themselves is limited: a standard described by its own creator as a “minimum viable product” [S6], arriving into a landscape where 62% of government bodies identify data quality as a barrier to AI adoption [S7].

The portfolio restructuring adds a structural dimension. The approximately 100 projects delegated to departmental portfolios move to assurance regimes led by the same departments whose self-reporting has been documented as optimistic. The central AI system provides a layer of analytical visibility — but the data it analyses for those departmental projects is generated by the departments themselves, with the same incentive structures that produced the documented bias.

The governance conditions

Any organisation contemplating AI-powered programme assurance faces the same structural question the UK deployment exposes. The question is not whether to deploy — the productivity case is real — but what governance architecture prevents the quiet erosion that well-designed automation enables.

Four conditions emerge from the evidence.

First, independent data validation. Programme data feeding AI assurance tools must be validated by a function independent of the project teams whose performance is being assessed. Self-reported data processed by AI does not become reliable because an algorithm processed it.

Second, maintained human capacity. Human assurance resources must be ring-fenced against the efficiency case that AI inevitably creates. This means explicit organisational commitment: staffing levels, review depth and independent challenge capacity that cannot be reduced by reference to algorithmic coverage.

Third, outcome-tested accuracy. AI assurance systems must be evaluated against outcome data — did the AI predict what actually happened? — not process metrics. The Scout tool’s 90% accuracy in mimicking human lines of enquiry tells us the AI can replicate what humans do; it does not tell us whether that replication improves what humans catch.

Fourth, sequenced reform. Data quality reform must precede or run alongside analytical deployment, not follow it. The four-year gap in the UK case is an institutional experiment in whether AI can add value while its inputs are being fixed. That experiment may succeed, but any organisation replicating it should understand that it is an experiment, not a proven architecture.

The distinction that matters

The UK government’s deployment is not a cautionary tale of reckless automation. It is something more instructive: a transparent, well-documented case where a considered institutional design confronts the structural conditions that automation bias research predicts will erode it. The AI early warning system has not produced wrong assessments. Human review has not visibly declined. The data reforms are serious.

The question is whether institutional humility can be made permanent in the face of cost pressure, political urgency, and the ordinary organisational tendency to trust systems that appear to work. The automation bias evidence from other domains suggests it cannot — not because the people are careless, but because maintaining vigilance against a system that produces confident outputs is cognitively expensive and institutionally unrewarded.

The governance implication for programme and portfolio leaders is precise: before trusting algorithmic programme assurance, verify that the data it ingests has been independently validated, that the human oversight it nominally augments has been structurally protected, and that the reforms intended to fix its inputs are not scheduled to arrive after the organisation has already built operational dependence on its outputs. The UK has provided the first evidence-rich test of what happens when these conditions are not fully met. The results will arrive over the next four years.

Sources

  1. NISTA — NISTA Major Projects Annual Report 2025-26 — July 2026 — https://www.gov.uk/government/publications/nista-major-projects-annual-report-2025-26/nista-major-projects-annual-report-2025-26
  2. Government Project Delivery — Progress on embedding data analytics and AI in project delivery — 1 August 2025 — https://projectdelivery.gov.uk/blogs/progress-on-embedding-data-analytics-and-ai-in-project-delivery/
  3. Construction Magazine UK — £924bn Government Projects Portfolio Puts AI and Delivery Risk Under New Scrutiny — July 2026 — https://www.constructionmagazine.uk/2026/07/nista-government-projects-ai-delivery-risk.html
  4. NAO — Over-optimism in government projects — 19 December 2013 — https://www.nao.org.uk/insights/optimism-bias-paper/
  5. NAO — Assurance for major projects — 2 May 2012 — https://www.nao.org.uk/reports/assurance-for-major-projects/
  6. UK Parliament Public Accounts Committee — Governance and decision-making on major projects — 10 September 2025 — https://publications.parliament.uk/pa/cm5901/cmselect/cmpubacc/642/report.html
  7. UK Parliament Public Accounts Committee — Use of AI in Government — 26 March 2025 — https://publications.parliament.uk/pa/cm5901/cmselect/cmpubacc/356/report.html
  8. UK Government — Government Major Projects Evaluation Review — 2025 — https://www.gov.uk/government/publications/government-major-projects-evaluation-review/government-major-projects-evaluation-review-html
  9. CSET Georgetown University — AI Safety and Automation Bias — November 2024 — https://cset.georgetown.edu/wp-content/uploads/CSET-AI-Safety-and-Automation-Bias.pdf

More from Programme