The Executive Judgment Boundary
The governance object is not the model; it is the decision role the model is permitted to perform.
Executive Summary
The executive AI debate is being framed too loosely. “Copilot”, “augmentation” and “human in the loop” describe an aspiration, not a control system. They do not tell a board which parts of a strategic decision may be delegated, which parts should be challenged by AI, and which parts must remain under accountable human judgement.
The evidence now supports a more exact approach. In a preregistered experiment involving 758 consultants, AI improved speed, task completion and quality on work inside the model’s capability frontier, yet reduced correctness on a deliberately selected complex task outside it. [S1] A 2026 laboratory experiment with 273 business professionals found that intermediate, critically interrogated use was associated with lower anchoring, confirmation and overconfidence in one market-entry exercise, while both low and high use produced weaker outcomes. [S2] A qualitative study of 33 executives in three AI-mature multinationals found a practical distinction between delegating structured analytical preparation and using AI iteratively when decisions were less predictable and more consequential. [S3]
None of these studies proves that a particular human-AI configuration improves live board decisions over a long horizon. Together, however, they reveal the wrongness of a single adoption rule. AI performance is jagged. Human oversight is behaviourally fragile. Strategic work combines tasks with very different levels of structure, context, verifiability and consequence. The useful governance unit is therefore not “AI use”. It is the decision-role assignment.
This paper introduces the Executive Judgment Boundary: a framework for assigning one of seven AI roles—retrieve, summarise, generate, compare, challenge, recommend or execute—according to eight characteristics of the decision. As uncertainty, irreversibility, context density and stakeholder consequence rise, AI should move away from recommender and executor roles towards evidence processing, option generation and structured dissent. Humans should retain problem framing, contextual synthesis, accountable commitment and final choice.
The boundary is not a permanent defence of human status. It is a dynamic allocation of authority. It should move when models become more capable, when firm-specific context becomes more reliable, when outputs become independently verifiable and when outcome evidence demonstrates that a wider role improves decisions. Until then, executive teams need a protocol that makes AI’s practical decision rights visible before fluent outputs quietly acquire them.
The Decision Changed Before the Governance Did
Consider a familiar strategy cycle. A management team is examining entry into a neighbouring market. The strategy function asks an internal AI system to summarise competitor moves, identify regulatory barriers, estimate likely responses, propose entry options and produce the first investment-committee paper. The output arrives overnight. It is comprehensive, internally consistent and formatted in the organisation’s preferred language.
By the time executives meet, several decisions have already been made without appearing on the agenda. The market has been framed as attractive or unattractive. Certain competitors have been treated as relevant. Particular risks have been elevated. The option set has narrowed around the patterns the model could retrieve and articulate. The base case has become an anchor. The human committee still holds formal authority, but practical discretion has begun to move upstream into the generation of the evidence pack.
This is why “a human makes the final decision” is an inadequate governance statement. Final approval is only one point in a decision chain. Strategic judgement also sits in problem definition, evidence selection, causal interpretation, option construction, disagreement, trade-off and commitment. If AI shapes those stages, human sign-off may preserve accountability on paper while diluting it in practice.
The governance object is not the model; it is the decision role the model is permitted to perform.
That distinction matters because executive AI is no longer confined to drafting emails or compressing reading. It is entering market analysis, acquisition screening, portfolio reviews, scenario design, investment cases and board preparation. These activities sit close enough to the decision that a recommendation can become influential before anyone decides whether the system was authorised to recommend.
The Evidence Supports Conditional Authority
The strongest current evidence does not justify either enthusiasm or prohibition. It supports conditional authority.
The jagged-frontier experiment is instructive because the same technology helped and harmed within a recognisably professional workflow. Participants using AI completed 12.2 per cent more inside-frontier tasks, finished 25.1 per cent faster and produced higher-quality work. On the complex outside-frontier task, AI users were 19 percentage points less likely to reach the correct solution. [S1] The mechanism is not that AI is broadly competent or broadly unreliable. It is that apparent task similarity does not reveal where the capability boundary lies, and users may not know when they have crossed it.
The 2026 cognitive-bias experiment sharpens the human side of the problem. In one high-uncertainty market-entry task, the relationship between use and measured bias was curvilinear. Intermediate use combined with critical interrogation was associated with stronger cognitive processing and lower bias; low use and high use produced weaker outcomes. [S2] This is meaningful, but it is not a universal dosage curve. It is one laboratory task under time pressure. The durable finding is narrower: the quality of engagement changes how AI affects judgement.
The qualitative executive study adds organisational texture. Its 33 participants across three AI-mature multinationals described greater confidence in delegating structured analytical work and more iterative human-AI co-deliberation when decisions were less predictable. [S3] The authors are explicit that demonstrations were not broad observations of unstaged board decisions and that their propositions remain conditional. That limitation is important. Practice can reveal an emerging architecture without proving that the architecture is optimal.
Taken together, these studies point to three conclusions.
- AI can create substantial value on bounded tasks.
- Human involvement does not automatically prevent error, anchoring or over-reliance.
- The appropriate role depends on the characteristics of the decision and the design of the process.
The available evidence does not show that moderate AI use is always best, that AI cannot do strategy, or that executives should reserve every consequential activity for themselves. Nor does it demonstrate improved long-horizon board outcomes. The honest position is more demanding: organisations must govern the boundary while the evidence remains incomplete.
Why Human Sign-off Comes Too Late
A formal human decision owner is necessary. It is not sufficient.
The first failure occurs through framing. A strategic question rarely arrives fully formed. Is declining growth a product problem, a channel problem, a pricing problem or a portfolio problem? The frame determines what evidence is sought and which options appear reasonable. An AI system that drafts the first issue tree can influence the decision before it makes any explicit recommendation.
The second failure occurs through anchoring. Fluent, numerical and well-structured output creates a reference point. Subsequent debate tends to adjust around it, even when the underlying assumptions are weak. Requiring a human to approve the final paper does not remove the anchor.
The third occurs through convergence. Executive teams already face pressure to reach coherence. A single AI answer can compress disagreement further by giving every participant the same synthesis, vocabulary and option set. Multiple agents can help only when they introduce genuinely distinct assumptions or evidence. Repeated outputs from closely related systems can manufacture the appearance of pluralism while reproducing correlated error.
The fourth occurs through context substitution. Models work from represented context. Strategic judgement often depends on what has not been codified: the credibility of a partner, the readiness of a leadership team, an informal regulatory signal, a history of failed integration, or a cultural constraint that changes the feasible option set. General pattern recognition can displace local knowledge precisely because it is easier to articulate.
The fifth occurs through authority transfer. When an organisation repeatedly accepts AI-generated frames, comparisons and recommendations, practical decision rights move even if the governance chart does not. Responsibility remains human, but the path of influence becomes opaque.
NIST’s AI Risk Management Framework provides a useful baseline: targeted application scope should be specified, human oversight defined, roles differentiated and executive responsibility made clear. [S4] For strategy, those principles need to be translated from system governance into decision governance.
The Executive Judgment Boundary
The Executive Judgment Boundary determines how much authority AI may exercise in a particular decision. It combines two elements: the character of the decision and the role assigned to AI.
Eight Decision Dimensions
Each strategic decision should be classified across eight dimensions.
| Dimension | Lower-bound condition | Higher-bound condition |
|---|---|---|
| Problem structuredness | Variables, criteria and outputs are clear | The problem itself is contested or ambiguous |
| Situational predictability | Patterns repeat and relevant relationships are stable | Discontinuity, novelty or genuine uncertainty dominates |
| Verifiability | Output can be checked independently and promptly | No reliable near-term ground truth exists |
| Feedback latency | Consequences become visible quickly | Effects emerge over years or through noisy signals |
| Reversibility | Errors can be corrected at low cost | Commitment creates path dependence or lock-in |
| Context density | Relevant context is codified and available | Tacit, political or proprietary context is decisive |
| Stakeholder consequence | Impact is limited and containable | Legal, ethical, employment or societal effects are material |
| Accountability clarity | A named owner can explain and defend the choice | Responsibility is diffuse, symbolic or contested |
These dimensions are not a scoring exercise designed to produce false precision. Their purpose is to make the sources of judgement difficulty visible. A decision can be highly structured but irreversible. It can be verifiable but politically consequential. It can contain excellent data but depend on a causal break from the past. The profile matters more than a single total.
Seven AI Roles
The framework distinguishes seven roles of increasing decision authority.
- Retrieve — locate relevant internal or external information.
- Summarise — compress evidence without selecting the preferred course.
- Generate — expand the option set, scenarios, hypotheses or questions.
- Compare — apply stated criteria, model sensitivities and expose differences.
- Challenge — construct counterarguments, surface assumptions and seek disconfirming evidence.
- Recommend — propose a preferred course and explain the rationale.
- Execute — initiate or complete action within defined limits.
The critical move is to assign a role rather than grant generic access. “The executive team may use AI” says almost nothing. “AI may retrieve, summarise and challenge for this acquisition decision, but may not define the acquisition thesis, recommend the target or initiate contact” creates a governable boundary.
Three Operating Zones
The decision profile and role assignment produce three operating zones.
Delegation Zone
AI can perform analytical work with limited iterative supervision when the task is structured, predictable, independently verifiable and reversible, with sufficient context and clear accountability.
Typical roles include retrieve, summarise, generate, compare and, in narrow cases, recommend or execute. Examples include assembling comparable-company data, checking arithmetic consistency, identifying duplicated assumptions, drafting sensitivity tables or monitoring agreed thresholds.
Delegation is not absence of control. The organisation still needs provenance, access controls, performance measures and exception handling. The difference is that the output can be tested against a sufficiently reliable standard.
Augmentation and Challenge Zone
AI should work as an evidence processor, option generator and dissent resource when the task contains uncertainty or competing interpretations but still benefits from wider search and structured comparison.
Here, executives should expect iteration. The model may produce scenarios, compare strategic logics, identify missing evidence, test consistency and generate the strongest case against the emerging view. Recommendation can be permitted only when its assumptions and sources are explicit and when an independent human frame already exists.
This zone is where much serious strategy work belongs. The aim is not consensus with the system. It is better human deliberation through wider search and disciplined challenge.
Human-owned Commitment Zone
As decisions become difficult to verify, slow to reveal outcomes, costly to reverse, dense in tacit context or consequential to stakeholders, AI’s role should narrow. It may still retrieve, summarise, compare or challenge. It should not acquire practical authority over problem framing, contextual synthesis, accountable commitment or final choice.
Market exits, major acquisitions, restructurings, senior succession, ethically sensitive product decisions and irreversible capital bets frequently sit here. The point is not that AI has nothing to contribute. It is that its output should enter as evidence or argument, not instruction.
The boundary should move from “What can the model produce?” to “What decision right has the organisation authorised it to exercise?”
A Worked Decision Chain
Return to the market-entry example.
At first glance, the decision appears data rich. Market size, pricing, competitor share and regulatory requirements can be researched. That makes several tasks suitable for delegation. AI can retrieve source material, summarise filings, compare market structures and generate an initial range of entry modes.
The boundary changes when the team moves from analysis to commitment.
Suppose the entry requires a local partner, a five-year capacity agreement and a public promise to retain employment at an acquired site. Feedback on the strategy will arrive slowly. Exit will be expensive. Partner quality depends on relationships that are incompletely documented. The employment commitment creates stakeholder consequence. The organisation has moved towards the right-hand side of several dimensions.
A disciplined role architecture would work as follows.
- Before AI exposure, the accountable executive writes an independent problem frame: what decision is being made, what success means and which commitments are unacceptable.
- AI retrieves and summarises evidence, but each material claim carries provenance and a confidence statement.
- AI generates at least three structurally different entry options, including a no-entry or delay case.
- A separate challenge process attacks the assumptions behind demand, partner reliability and exit cost.
- Humans add tacit context that the system cannot reliably infer and record where it changes the analysis.
- AI may compare options under agreed criteria, but it does not select the weights or recommend the final commitment.
- The investment committee records its rationale, including where it rejected or modified AI output.
This is not a ceremonial human loop. It is a division of cognitive and accountable labour.
The Strongest Case for Moving Faster
The strongest objection is that this framework may freeze today’s limitations into tomorrow’s bureaucracy.
Models are improving rapidly. Firm-specific retrieval, evaluation systems and specialised agents can reduce context loss. AI may already outperform executives on some analytical and judgement tasks. Human reservation can preserve executive bias, political compromise and status protection. A cautious framework may become an excuse for leaders to keep authority while delegating only administrative work.
That objection is valid. The judgment boundary should not be used to protect human identity or to assume human superiority. Human teams are inconsistent, selective and prone to groupthink. There will be cases where an AI recommendation is more rigorous than the executive preference it challenges.
The answer is not to abandon the boundary. It is to make the boundary dynamic and evidence-led.
A task should move towards recommendation or execution when the organisation can show that the relevant context is represented, performance is stable across realistic cases, outputs are independently verifiable, failure modes are understood, escalation works and measured outcomes improve. Conversely, a task should move back towards augmentation when model changes, data drift, new regulation or unusual conditions weaken those assurances.
The framework therefore imposes an obligation on cautious leaders as well as enthusiastic ones: they must provide evidence for keeping authority human. “We have always decided this ourselves” is no more adequate than “the model is usually right”.
Failure Modes the Framework Must Catch
A boundary is useful only if it exposes predictable failure.
Polished error occurs when confident form substitutes for evidential strength. The control is provenance, uncertainty disclosure and independent verification.
Generic advice occurs when broad patterns displace firm-specific advantage. The control is explicit identification of tacit context, proprietary causal beliefs and assumptions the model cannot access.
Correlated consensus occurs when several outputs appear to agree because they share data, architecture or prompting logic. The control is independent evidence and materially different challenge methods, not a count of agreeing agents.
Automation bias occurs when users accept a recommendation because it is available, fast or apparently authoritative. The control is an independent human frame, forced consideration of alternatives and recorded reasons for acceptance or rejection.
Algorithm aversion occurs when leaders dismiss useful output after visible error or because it threatens their authority. The control is task-level performance evidence and comparison with human baselines.
Symbolic oversight occurs when the human approver lacks time, context or permission to challenge the AI-shaped case. The control is to place human judgement before and during framing, not only at the signature stage.
The Judgment Boundary Protocol
For high-consequence decisions, organisations should use an eight-step protocol.
- Classify the decision. Describe its position across structuredness, predictability, verifiability, feedback latency, reversibility, context density, stakeholder consequence and accountability clarity.
- Assign the AI role. State whether AI may retrieve, summarise, generate, compare, challenge, recommend or execute. Do not grant an undefined “copilot” role.
- Frame independently. Above a defined consequence threshold, require an accountable human to state the problem, decision criteria and non-negotiable constraints before seeing AI output.
- Expose the evidence. Require source provenance, uncertainty, known limits and the distinction between fact, inference and generated illustration.
- Create independent challenge. Use a human red team, a genuinely separate model or both. Treat agreement as a signal to investigate, not as verification.
- Name the owner. Record who holds authority, what advice was accepted or rejected and why the commitment remains defensible.
- Set stop conditions. Halt recommendation or execution when critical evidence cannot be verified, disagreement remains unresolved, confidentiality is at risk or the system operates outside the approved role.
- Review behaviour and outcomes. Track acceptance, modification, rejection, overrides, incidents and outcome quality. Reassess the same task after material model or context changes.
The protocol should be proportionate. Low-stakes retrieval does not need an investment-committee process. High-consequence strategy should not inherit controls designed for drafting productivity.
What Would Move the Boundary
The main uncertainty is empirical. There is still no strong causal evidence showing which configurations improve long-horizon outcomes in live board and top-management decisions. Most available evidence comes from laboratory tasks, professional experiments, simulations, retrospective counterfactuals, interviews or practitioner reports.
The boundary should move when evidence improves in four areas.
- Capability: repeated testing shows stable performance on the relevant task, including unusual and adversarial cases.
- Context: firm-specific data, definitions and causal assumptions are represented accurately enough to reduce generic substitution.
- Verification: critical outputs can be checked independently before commitment, and feedback arrives soon enough to support learning.
- Outcomes: prospective evidence shows that the human-AI configuration improves decision quality, not merely speed, confidence or document quality.
The key challenge is that underuse can be costly. If models become better at identifying weak assumptions or comparing complex evidence, excessive human reservation may preserve bias and slow decisions. The framework must therefore measure both sides: errors caused by inappropriate delegation and value lost through unjustified restriction.
The Question for the Strategy Table
Executive AI governance will remain weak while it is organised around adoption levels, tool access or generic human oversight. Those controls describe the technology estate. They do not reveal how judgement is being allocated.
The more consequential question is specific: which parts of this decision are we willing to let AI shape, and on what evidence?
Once that question is asked, “copilot” becomes too vague. The organisation must distinguish retrieval from recommendation, challenge from authority, and analytical preparation from accountable commitment. It must also accept that the correct line will change.
The mature position is neither human primacy nor machine autonomy. It is a visible, reviewable allocation of decision rights that uses AI where its contribution can be tested, turns it into a dissent resource where uncertainty remains, and keeps responsibility attached to the people who can explain and carry the consequences of the choice.
Sources
- INFORMS / Organization Science — Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality — 11 March 2026 — https://pubsonline.informs.org/doi/10.1287/orsc.2025.21838
- SAGE / Strategic Organization — Generative AI and Cognitive Biases in Strategic Decision-Making — 26 May 2026 — https://journals.sagepub.com/doi/10.1177/14761270261457350
- Springer Nature / Group Decision and Negotiation — Integrating Artificial Intelligence in Strategic Decision-Making: Contexts for Delegation and Augmentation — 17 July 2026 — https://link.springer.com/article/10.1007/s10726-026-10016-x
- National Institute of Standards and Technology — AI RMF Core — 26 January 2023 — https://airc.nist.gov/airmf-resources/airmf/5-sec-core/