The Risk the Model Did Not Ask About

Analysis·Giovanni Leonardi·August 2026·11 min read

When AI proposes the questions, it does more than accelerate preparation. It allocates institutional attention.

The question list arrives before the reviewers

A major-programme assurance team receives thousands of pages: schedules, risk registers, cost forecasts, minutes, dependencies, benefits cases and previous findings. Before the first interview, an AI tool proposes the lines of enquiry. It flags supplier concentration, an unresolved planning consent and a benefits baseline that has not been updated. The reviewers save hours and begin with sharper questions.

One material issue is absent. A policy dependency sits outside the documents the tool was given. The experienced reviewer notices and adds it. The less experienced reviewer assumes the generated list represents broad coverage and moves on.

This is an illustrative workflow, not a report of a particular review. It exposes the central control problem. When AI proposes the questions, it does more than accelerate preparation. It allocates institutional attention.

That distinction matters as the UK Government Project Delivery Function prepares to use Assure Assist in major-programme assurance. Its official July 2026 update says the tool is secure, stable and ready to support Government Major Projects Portfolio reviews. It also reports a 92.5% accuracy rate in identifying appropriate lines of enquiry during beta testing [S1]. A separate official post said rollout across projects and programmes was expected soon [S2].

Those statements establish readiness as an official claim. They do not establish that assurance quality has improved, that programme risks are found earlier or that delivery outcomes change. The denominator, sample, benchmark, false-negative rate and independent validation behind the 92.5% figure have not been published.

The consequential question is therefore not whether the tool can agree with reviewers. It is whether the combined human–AI review system finds the material risk that neither party would have found alone.

Assurance is the allocation of scarce attention

Programme assurance is sometimes described as an independent judgement applied to a body of evidence. In practice, it begins earlier. Reviewers decide what to inspect, whom to interview, which inconsistencies to pursue and how much time to give each risk. No team can test everything. The first act of assurance is the allocation of attention.

An AI system that retrieves a document is helpful. A system that proposes or ranks lines of enquiry is more consequential. Its output shapes the sequence of the review:

  1. Programme evidence enters an approved data environment.
  1. The tool identifies patterns and proposes questions.
  1. Reviewers accept, modify, reject or add enquiries.
  1. Those enquiries shape interviews and evidence requests.
  1. Findings become recommendations and assurance judgements.
  1. Sponsors and oversight bodies decide whether to continue, intervene or change delivery.

Formal accountability remains human, but influence enters at step two. Risks surfaced early gain cognitive priority. Risks omitted require a reviewer to recognise an absence, which is harder than challenging a visible suggestion. The model becomes a salience engine.

That is why “human in the loop” is an incomplete control. A named person can sign the final judgement while relying heavily on the questions placed in front of them. The meaningful evidence of human control is behaviour: what reviewers challenged, supplemented, rejected and discovered beyond the tool’s frame.

What 92.5% can and cannot mean

An accuracy figure looks precise because it compresses evaluation into one number. Its governance value depends on what was counted.

If the benchmark is agreement with historical reviewers, 92.5% may show that the tool reproduces conventional lines of enquiry. That could be useful for consistency and training. It would not show that 92.5% of material risks were found. Historical reviewer choices are not a complete inventory of truth, and major programmes often fail through interactions, emerging dependencies or contextual changes that previous reviews did not anticipate.

The missing error profile is especially important. False positives consume reviewer time. False negatives conceal relevant questions. The two costs are not symmetrical. A harmless extra enquiry is inconvenient; a missed dependency can distort the assurance conclusion. Aggregate accuracy can hide poor performance on rare, consequential conditions.

The appropriate validation questions are concrete:

  • What was the unit counted: suggested enquiries, reviewer agreement, risks, documents or programmes?
  • Which programme types and delivery phases were represented?
  • Who judged an enquiry “appropriate”, and were those judges independent of development?
  • What did the tool miss, particularly where reviewers found novel or contextual risks?
  • How did performance change with incomplete, optimistic or inconsistent programme data?
  • Did reviewers discover more consequential issues, or merely prepare faster?

Until those answers are available, the figure is a development signal rather than an assurance result.

The strongest case for augmentation

The case for Assure Assist is substantial. Major programmes generate more evidence than review teams can absorb. A well-designed tool can search consistently, connect patterns across documents and ensure that familiar lines of enquiry are not forgotten. It can give reviewers a common starting point while leaving more time for interviews, professional scepticism and contextual judgement.

The operating environment is also becoming more favourable. Government Project Delivery reports a common Project Data Standard and investment in data and AI capability [S1]. More consistent inputs can improve retrieval and comparison. The National Audit Office notes that major and mega-project governance must be clear about authority, accountability, outcomes and the role of the investing organisation [S3]. AI could help reviewers surface evidence relevant to those questions across a large portfolio.

The strongest augmentation model is not replacement. It is cognitive division of labour. The tool handles repeatable scanning and pattern retrieval; reviewers concentrate on ambiguity, contradiction, political context, emergent dependencies and judgement. Strong reviewers could become more effective because they spend less time assembling the obvious and more time testing the consequential.

This view also avoids a simplistic assumption that people always defer to algorithms. Peer-reviewed public-sector experiments found no universal automation bias; responses varied with task, context and the relationship between advice and prior beliefs [S5]. Actual reviewer behaviour must be measured, not inferred from a generic risk label.

The strongest case against it

The sceptical interpretation is equally serious. “Accuracy” against historical practice can reward conformity. If past assurance concentrated on schedule, cost and standard risk categories, the tool may learn to make those questions more visible while underweighting novel delivery models, institutional incentives or cross-programme interactions.

Common data create another trade-off. Standardisation improves comparability, but it can produce common-mode blind spots. A risk absent from the standard, poorly recorded across departments or systematically described in optimistic language may be absent everywhere at once. The model can then produce consistent analysis from consistently incomplete evidence.

There is also a capability paradox. Experienced reviewers may use suggestions as hypotheses and recognise omissions. Inexperienced reviewers may treat them as coverage. The same tool can therefore augment one team and narrow another. Formal human sign-off does not resolve the difference.

The sceptical case sharpens the conclusion: the control model must evaluate the review system, not the model in isolation. Technical readiness, security and stability are necessary. They are not operational proof that institutional attention is being allocated well.

Traceable question framing

A practical control model should make the origin, limits and human treatment of every material enquiry visible.

Control Evidence required
Provenance Source documents, data fields and reasoning basis linked to each material suggestion
Coverage Explicit statement of scope, unavailable evidence and known data gaps
Omission test Independent reviewer pass asking what material risk is missing
Human action Accept, modify, reject and add decisions recorded with concise reasons
Performance False negatives, unique material issues, evidence strength and downstream interventions
Change control Revalidation after material model, data, portfolio or workflow changes

Provenance is more than a clickable citation. The reviewer needs enough context to assess whether the source is current, complete and relevant. A risk-register entry can support an enquiry while still understating exposure. Traceability should invite challenge, not merely explain where text came from.

The omission test is the most important addition. Reviewers should complete an initial risk-framing exercise without seeing the generated list, or a separate reviewer should conduct a blind challenge on a sample. This preserves a route for non-obvious risks to enter the review.

Human actions should also be treated as performance data. Frequent additions can reveal gaps in scope. Rare overrides can indicate either excellent output or weak challenge. The pattern needs interpretation by programme type and reviewer experience.

Finally, evaluation should follow consequences. A suggested enquiry has value when it leads to stronger evidence, a material finding, an earlier mitigation or a better decision. Adoption, usage and agreement are intermediate measures.

A better test design

The next phase should compare assisted and unassisted review teams on representative programmes. A credible design would stratify by programme type, maturity and risk, then measure:

  • preparation time;
  • overlap and difference in proposed enquiries;
  • unique material issues identified;
  • false positives and false negatives;
  • quality of evidence supporting findings;
  • reviewer additions and overrides;
  • downstream decisions, mitigations and escalations.

Some cases should deliberately contain incomplete or contradictory data. Others should include unusual dependencies not prominent in historical reviews. This tests whether the combined system handles novelty rather than only reproducing known patterns.

Independent re-review is essential. Development teams have an understandable interest in demonstrating progress, while historical reviewers may prefer outputs that resemble their own practice. The National Institute of Standards and Technology recommends separating model development and use from verification and validation where possible, and frames accountability, transparency, validity and context as connected properties [S4].

The Financial Reporting Council’s 2026 guidance offers a useful professional analogy: confidence and mitigation should vary with intended use, while accountability remains with the human auditor [S6]. Programme assurance is not financial audit, but the design principle transfers. The closer AI output sits to a consequential judgement, the stronger the evidence and challenge required.

What leaders should govern

Senior responsible owners and assurance leaders need to distinguish three accountabilities.

The tool owner is accountable for scope, data access, validation, change control, incidents and retirement.

The assurance lead is accountable for how the tool is used in a review, whether omissions are tested and whether evidence supports the final judgement.

The sponsor and oversight body remain accountable for acting on assurance findings and for deciding whether programme exposure is acceptable.

These roles should not collapse into a generic statement that “a human remains responsible”. Responsibility must attach to observable decisions. The National Audit Office’s 2026 good-practice guide similarly emphasises clear ownership, measurable objectives, realistic resources and governance that supports responsible scaling [S7].

The leader’s dashboard should therefore show more than licences, prompts or time saved. It should reveal which material enquiries originated with the tool, which were added by reviewers, how misses were found, what actions followed, and where performance varies across programme types.

The risk beyond automated review

Assure Assist has crossed a meaningful institutional threshold: the government says it is ready to support major-programme assurance. That makes disciplined evaluation timely. It does not justify the claim that assurance is automated, that rollout is complete or that programme delivery will improve.

The opportunity is real. AI can widen the evidence a reviewer can inspect and release time for deeper challenge. The risk is equally real: a persuasive question list can turn historical attention patterns into a default frame that reviewers do not see.

The governing standard should be simple. Do not ask only whether the model produced the questions reviewers expected. Ask what material risk the combined system found that would otherwise have remained invisible—and what it still failed to ask.

Sources

  1. Government Project Delivery Function — From strategy to scale: Building data and AI capability for the project delivery profession — 30 July 2026 — https://projectdelivery.gov.uk/blogs/from-strategy-to-scale-building-data-and-ai-capability-for-the-project-delivery-profession/
  2. Government Project Delivery Function — Strengthening our foundations — 24 July 2026 — https://projectdelivery.gov.uk/blogs/strengthening-our-foundations/
  3. UK National Audit Office — Governance and decision-making on mega-projects — 14 March 2025 — https://www.nao.org.uk/insights/governance-and-decision-making-on-mega-projects/
  4. US National Institute of Standards and Technology — Artificial Intelligence Risk Management Framework 1.0 — 26 January 2023 — https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf
  5. Journal of Public Administration Research and Theory — Human–AI Interactions in Public Sector Decision Making: Automation Bias and Selective Adherence to Algorithmic Advice — 8 February 2022 — https://academic.oup.com/jpart/article/33/1/153/6524536
  6. UK Financial Reporting Council — Innovative new guidance supports audit firm adoption of emerging AI technologies — 30 March 2026 — https://www.frc.org.uk/news-and-events/news/2026/03/innovative-new-guidance-supports-audit-firm-adoption-of-emerging-ai-technologies/
  7. UK National Audit Office — Good practice guide for organisations using AI — 15 May 2026 — https://www.nao.org.uk/wp-content/uploads/2026/05/good-practice-guide-for-organisations-using-ai.pdf

Giovanni Leonardi  ·  About  ·  LinkedIn

Leave a Reply

Your email address will not be published. Required fields are marked *