The Calibrated Assurance Test
Researched by an agentic pipeline · reviewed and gated by the author
Delegation without a common evidence layer is merely the relocation of risk.
The Calibrated Assurance Test
How to reduce approval friction without moving complex programme risk out of sight
The United Kingdom has begun a consequential experiment in major-programme governance. From 1 April 2026, the Government Major Projects Portfolio was reduced from more than 200 programmes and projects to a more concentrated group of roughly 80 to 100. Work that still requires Treasury approval but no longer sits in that central portfolio is expected to be managed through departmental major-project portfolios, supported by common reporting through the Government Reporting Integration Platform.
It is tempting to read this as a familiar argument between central control and local freedom. That is too crude. The more useful interpretation is a capability-contingent assurance architecture: delegate routine decisions, retain common information and feasibility disciplines, and concentrate independent challenge where consequence, complexity or weak local capability justify it.
This distinction matters because approval volume and assurance quality are not the same thing. A programme can pass through many approvals and still lack a clear owner, an honest delivery forecast or an effective escalation route. Conversely, fewer approval points can improve accountability if decisions sit with capable sponsors who work to common standards and face credible independent challenge at the moments that matter.
The reform therefore creates a test, not yet a proven result. There is no transparent post-reform evidence showing that programmes transferred from central oversight decide faster, preserve delivery confidence or realise benefits more reliably. The design should be judged by whether it removes duplicate control while preserving the capacity to detect, challenge and escalate material risk.
What has changed — and what has not
The April reform delegates decisions to the lowest appropriate level, consolidates several central controls and reserves more central attention for strategically important or high-risk work. Programmes seeking entry to the central portfolio must first undertake an initial feasibility study. A separate mega-project category, typically for work above £10 billion with national transformational impact, receives enhanced attention. Version 3 of the Teal Book, published in July 2026, aligns government delivery guidance with the new arrangements.
The latest public performance picture is still the March 2026 baseline. The pre-reset portfolio contained 189 projects with £924.2 billion of whole-life cost. At that point, 58 per cent were rated Amber and 18 per cent Red. Those ratings are snapshots of delivery risk, not predictions of ultimate success, and they pre-date the new operating model. They cannot establish whether the reform works.
Nor has departmental accountability disappeared. Business cases, funding authority, benefits ownership, gates and assurance remain necessary. A smaller central portfolio does not make excluded digital, service or transformation programmes less complex. It changes where challenge is organised and who must possess the capability to act.
The Office for Value for Money diagnosed a system in which overlapping controls diluted accountability and consumed delivery capacity. Its answer was to place decisions closer to delivery, set common standards centrally and focus scarce specialist attention on the largest risks. That logic is sound only if the receiving department can perform the work that central scrutiny previously supplied.
The three-tier framework
Leaders should classify assurance through three tests: consequence, complexity and local capability. Cost is relevant, but it is not a sufficient proxy for any of them.
| Tier — Appropriate use — Assurance pattern — Movement trigger |
|---|
| 1. Local delivery — Repeatable work; limited exposure; mature team; reversible commitments — Departmental gates, common data, benefits tracking and proportionate peer review — Rising interdependence, irreversible commitment or weaker delivery confidence |
| 2. Federated oversight — Material programmes with cross-functional complexity but credible local capability — Departmental ownership plus scheduled independent challenge and specialist access — Enterprise-wide consequence, capability failure or persistent adverse indicators |
| 3. Enterprise assurance — Nationally significant, highly novel, systemically exposed or capability-constrained work — Intensive independent assurance, central specialist support and explicit executive escalation — Sustained evidence that consequence or uncertainty has reduced and local capability is proven |
This is not a permanent classification. A programme should move between tiers as evidence changes. A locally managed programme may require enterprise assurance when a critical supplier fails, benefits assumptions deteriorate or cross-system dependencies emerge. A centrally assured programme may step down when uncertainty narrows and the department demonstrates reliable control.
The mechanism is straightforward.
- Remove overlapping approvals that ask different bodies to reconsider the same decision.
- Name the accountable sponsor and place routine decisions within that authority.
- Preserve a common evidence layer so performance, cost, risk and benefits can be compared.
- Test feasibility before detailed design and sunk cost make challenge politically difficult.
- Direct independent assurance and specialist expertise towards high-consequence decisions and weak capability.
- Escalate quickly when thresholds are crossed, rather than waiting for formal failure.
Delegation without a common evidence layer is merely the relocation of risk.
The capability assessment
Before a programme leaves intensive oversight, its sponsor should demonstrate five capabilities.
- Decision authority: the sponsor can make or secure material scope, funding and sequencing decisions without creating a shadow approval chain.
- Delivery intelligence: cost, schedule, risk, dependencies and benefits are defined consistently, updated promptly and open to challenge.
- Independent judgement: reviewers are sufficiently separate from the delivery incentives they assess, and their findings lead to explicit decisions.
- Specialist access: commercial, digital, infrastructure, security and benefits expertise can be brought in before commitments become difficult to reverse.
- Escalation discipline: thresholds, recipients and response times are agreed in advance, including a route back to enterprise assurance.
This capability assessment should accompany an escalation contract. The contract specifies the indicators that move the programme between tiers, who can call for additional assurance, how quickly the decision must be taken and what evidence cannot be waived. It turns delegation from an expression of trust into an observable operating agreement.
Consider a composite public-service programme replacing a legacy case-management platform across several agencies. Its cost may sit below a central threshold, yet its migration risk, supplier concentration, data dependencies and operational irreversibility are substantial. A cost-only rule could place it in Tier 1. The calibrated test would place it in Tier 2 because complexity and consequence remain high. If the department lacks migration expertise or reporting quality deteriorates, the escalation contract moves it temporarily to Tier 3. Once rehearsals, independent testing and local capability provide credible evidence, it can step down again.
The strongest objection
The sceptical case deserves more than a footnote. Reducing the central portfolio may remove independent challenge and specialist support from programmes that remain difficult. Classification can be shaped by cost or political priority while operational complexity sits elsewhere. Departments may inherit accountability faster than they build capability. A common reporting platform can standardise fields without improving the candour of forecasts or the quality of decisions.
Historical evidence reinforces the concern. The National Audit Office found varied outcomes among projects leaving the major-project portfolio and warned against treating departure as evidence of successful completion. Parliamentary scrutiny has repeatedly highlighted weaknesses in governance, decision-making and accountability. Research on central assurance finds associations between assurance activity and improved delivery confidence, but observational relationships do not prove that assurance caused better outcomes. Stronger programmes may be more able to act on recommendations, while weaker ones may receive more scrutiny precisely because they are already in difficulty.
There are also alternative explanations for any early improvement. Faster decisions might reflect weaker challenge rather than better governance. Departments granted more autonomy may already be stronger, producing selection effects. Better data or earlier feasibility work might create the benefit independently of portfolio tiering. These possibilities are not reasons to reject delegation; they are reasons to make the causal claims modest and the evaluation explicit.
How executives should apply the test
The operating rhythm can be built into existing portfolio governance.
- At initiation, assess consequence, complexity and local capability separately. Record the evidence and do not let a financial threshold substitute for judgement.
- Assign the lowest tier that preserves credible challenge. State which decisions are genuinely delegated and remove duplicate approvals.
- Establish a common minimum dataset covering forecast cost, schedule, benefits, dependencies, delivery confidence and decision latency.
- Require an initial feasibility challenge before the detailed business case. Ask whether the objective, options and delivery vehicle are plausible, not merely whether the preferred solution is well described.
- Write the escalation contract. Include adverse trends, unresolved dependencies, sponsor turnover, supplier distress, benefits erosion and poor data quality as possible triggers.
- Review the tier at material gates and after shocks. A programme’s assurance level is a current judgement, not an inherited status.
- Publish the consequences of assurance. Track recommendations accepted, decisions changed, proposals stopped and support deployed — not simply reviews completed.
What would prove the model wrong
The reform should be treated as a testable design, with a baseline taken before its effects are claimed. It would be failing if:
- approval-cycle time does not fall, or duplicate controls reappear inside departments;
- data completeness, timeliness or comparability deteriorate outside the central tier;
- transferred programmes show worsening delivery confidence after adjustment for prior risk and complexity;
- escalation occurs only after cost, schedule or service failure becomes unavoidable;
- feasibility studies rarely change, pause or stop weak proposals;
- departments with limited sponsorship or specialist capability perform no differently because assurance tiers ignore that capability;
- faster approvals are not accompanied by stable or improving benefits confidence.
Evaluation should compare programmes by initial risk, complexity and departmental capability rather than contrasting central and departmental portfolios at face value. It should also examine near misses and decisions changed, because effective assurance may prevent harm without producing a visible delivery event.
Beyond Central Control
The central question is not how much control the centre retains. It is whether the system directs challenge to the decisions where challenge can still alter the outcome.
For executive sponsors and enterprise portfolio leaders, the lesson travels beyond government. Large organisations often respond to slow governance by abolishing forums, raising thresholds or pushing approvals into business units. That can remove friction, but it can also make risk less visible. The better move is calibrated assurance: local authority where capability is demonstrated, a common evidence layer across the enterprise, early feasibility work before commitment, and escalation tied to consequence and uncertainty.
This model will not suit every context. Small, repeatable projects with mature teams and reversible commitments may need little beyond local control. Organisations without reliable data, clear decision rights or independent assurance should not delegate high-consequence work merely to appear agile. They must first build the conditions that make delegation safe.
The United Kingdom’s reform is therefore best viewed as an operating hypothesis. If departments gain real authority, common information remains trustworthy and independent challenge follows consequence rather than bureaucracy, decisions may become faster without becoming weaker. If capability and escalation lag behind delegation, the system will have displaced risk rather than reduced it. The difference will be visible in evidence, not in the size of the central portfolio.
Sources
- Government Project Delivery, “Government strengthens project delivery and accountability”, 1 April 2026. https://projectdelivery.gov.uk/2026/04/01/government-strengthens-project-delivery-and-accountability/
- Government Project Delivery, “The Teal Book has been updated”, 1 July 2026. https://projectdelivery.gov.uk/2026/07/01/the-teal-book-has-been-updated/
- National Infrastructure and Service Transformation Authority, “Government major projects demonstrate strong foundations for delivery”, 13 July 2026. https://www.gov.uk/government/news/government-major-projects-demonstrate-strong-foundations-for-delivery
- Office for Value for Money and HM Treasury, “Reforming the spending control and accountability framework”, November 2025. https://assets.publishing.service.gov.uk/media/6925f3a122424e25e6bc31b3/Controls_report.pdf
- House of Commons Committee of Public Accounts, “Governance and decision-making on major projects”, 10 September 2025. https://publications.parliament.uk/pa/cm5901/cmselect/cmpubacc/642/report.html
- National Audit Office, “Projects leaving the Government Major Projects Portfolio”, 19 October 2018. https://www.nao.org.uk/wp-content/uploads/2018/10/Projects-leaving-the-Govenment-Major-Projects-Portfolio-Summary.pdf
- Kirkham et al., “An empirical study of assurance in the UK government major projects portfolio: from data to recommendations, to action or inaction”, International Journal of Managing Projects in Business, 2021. https://research.manchester.ac.uk/en/publications/an-empirical-study-of-assurance-in-the-uk-government-major-projec/
- Construction Management, “UK just dropped scores of major projects from central control. Here’s why that’s not good”, 18 May 2026. https://constructionmanagement.co.uk/uk-just-dropped-scores-of-major-projects-from-central-control-heres-why-thats-not-good/