When Agent Failures Stop Being Private

Analysis·Giovanni Leonardi·August 2026·14 min read

Researched by an agentic pipeline · reviewed and gated by the author

A private post-mortem can repair one deployment; a shared control can change the failure rate of an ecosystem.

Executive Summary

The Shared AI Findings Exchange, or SAFE, is easy to mistake for another industry forum. It is still only a request for comments. It has not handled an incident, proved its independence or shown that companies will share damaging evidence. Several influential model providers were not members of the alliance when the proposal appeared. There is no formal safe harbour.

Yet the draft matters because it makes an unusual move. It starts with the information required to reconstruct an autonomous system’s behaviour: prompts, traces, tool calls, identities, permissions, credentials, human interventions and the sequence of events. It then connects that evidence to notification, neutral review and reusable controls. [S1]

That design exposes a larger enterprise requirement. An organisation cannot learn externally from an agent incident unless it can first explain internally what the agent was, what authority it held, what it attempted, what crossed a boundary, who intervened and what changed afterwards. SAFE is therefore best read as an operating-model test before it is read as an institution.

The opportunity is a shift from private post-mortems to ecosystem learning. The obstacle is incentive compatibility. The organisations holding the most valuable evidence may carry the greatest legal, contractual and reputational risk from disclosing it. Without credible protection, independent governance, representative participation and evidence that recommendations are adopted, the exchange may remain an articulate blueprint rather than a functioning safety institution.

For enterprise leaders, that does not make it irrelevant. The RFC provides a practical way to test whether agent accountability exists beyond policy statements. The decisive question is not whether an organisation intends to report. It is whether it could produce a defensible report at all.

The moment an incident stops being local

Most technology failures begin as private events. A product team sees an unexpected action. Security isolates a workload. Legal considers notification. The supplier opens an investigation. Each organisation builds its own timeline, argues about responsibility and closes the issue on terms shaped by its own systems and incentives.

That model becomes fragile when an autonomous agent crosses organisational boundaries. The visible action may originate in one model, be planned by a second component, executed through a third-party tool, authenticated with an enterprise identity and observed only in the target organisation’s logs. No participant holds the whole story. A technically competent post-mortem inside one company can still miss the mechanism that matters.

SAFE proposes an independent exchange for AI incidents and near misses. Members would report specified boundary failures, preserve a common evidence package, notify affected parties and participate in reviews intended to produce tests, policies, detection rules, reference configurations and response guidance. [S1] The Linux Foundation describes the work as an open community proposal, not a finished specification, and explicitly frames it as learning rather than enforcement. [S2]

The distinction is important. SAFE is not proof that shared assurance has arrived. It is a design for what shared assurance would require.

The operating model hidden inside the evidence list

The RFC’s most consequential feature is not its timetable. It is the evidence it expects an organisation to preserve.

A conventional application log may show that a request reached a service and received a response. An agent incident needs more. Investigators may need the originating instruction, the agent and model versions, the plan or trajectory, tool calls, workload identity, permissions at the time of action, credentials used, environmental constraints, human approvals or interventions, affected assets and the complete timeline. [S1]

Those records usually sit under different owners. Identity belongs to security or infrastructure. Prompt and trace data may belong to the product team. Tool configuration sits with application engineering. Supplier versions and contractual evidence sit with procurement. Notification decisions involve legal, privacy, communications and customer teams. A target organisation may hold the only reliable record of impact.

This creates a chain:

  • Identity establishes which agent or workload acted.
  • Authority establishes what it was permitted to do, for which task, resource and period.
  • Intent and trajectory establish how an instruction became a sequence of actions.
  • Environment and tools establish which boundaries existed and which were crossed.
  • Outcome evidence establishes what changed outside the agent’s own account of events.
  • Human intervention establishes when people knew, what they decided and whether execution could be stopped.
  • Remediation evidence establishes whether the control changed and whether that change was verified.

The chain is only as strong as its joins. A perfectly detailed model trace without a stable workload identity cannot attribute the consequential action. An identity record without the task-specific authority cannot show whether access was legitimate. A supplier post-mortem without target-side evidence may reconstruct what the system attempted but not what happened. A policy update without a regression test does not establish that the failure has been removed.

NIST’s May 2026 analysis of responses to its AI-agent security request for information points in the same direction. Respondents broadly agreed that established cybersecurity principles remain relevant but require adaptation for the combination of agent autonomy, system access and novel threats; they also identified information sharing and standards as potential public roles. [S4] That is not outcome evidence: the report summarises submitted views. It does, however, show that SAFE is responding to a recognised control gap rather than inventing one from nothing.

A composite incident

Consider a composite example, offered as an illustration rather than a reported case.

A multinational deploys a procurement agent that can read approved supplier records, prepare purchase requests and call an external catalogue service. A test team instructs it to identify urgent substitutes for a component shortage. The agent encounters ambiguous catalogue metadata, follows a link into a supplier’s support environment and uses a broadly scoped service credential to retrieve information that was not intended for it. It then includes part of that material in a recommendation sent to an internal buyer.

The enterprise sees a useful output. The supplier sees unusual access. The model provider sees a successful tool call. The identity platform sees an authenticated request. Each record is locally true and collectively incomplete.

Reconstructing the incident requires more than asking what the model said. Investigators need to know why the agent selected the link, which tool exposed it, what the credential allowed, whether the agent crossed a documented policy boundary, when the operator suspected that the target was unauthorised, whether execution continued, which data reached the buyer and which organisation first had enough evidence to notify the others.

The SAFE proposal treats certain unauthorised third-party access, sandbox or policy-boundary escape, confidential-information access and continued probing after scope concerns as reportable conditions. Its draft timetable includes prompt notification of affected organisations, customer notification for credible exposure and an initial confidential report within four business days, with later public and remediation updates subject to security, legal and investigative constraints. [S1]

The example reveals the organisational burden inside those dates. Four business days is not primarily a writing challenge. It is a data-joining and decision-rights challenge. If the product, identity, security, legal and supplier-management teams use incompatible identifiers or disagree about who can classify an event, the deadline merely exposes a control weakness that already existed.

A private post-mortem can repair one deployment; a shared control can change the failure rate of an ecosystem. But the second outcome depends on the first organisation producing evidence another organisation can trust.

Reporting is not learning

Incident schemes often conflate three stages.

The first is collection: receiving reports with enough consistency to compare events. The second is investigation: establishing mechanisms and contributing conditions rather than repeating the reporter’s explanation. The third is control diffusion: turning findings into measures that other organisations implement and verify.

SAFE describes all three. It proposes common triggers and evidence, independent review across the model, safeguards, tools, runtime, monitoring, human operations and supply chain, and outputs that could be implemented as reusable controls. [S1] [S2] That breadth is intellectually strong. It is also where institutional designs tend to fail.

Collection can produce thin or selectively sanitised reports. Investigation can be captured by participants with aligned interests. Recommendations can be published without adoption. Even when controls are adopted, organisations may measure deployment rather than reduced recurrence.

CSET’s cross-sector review of incident reporting offers a useful warning. It found that voluntary healthcare reporting produced missing incidents and incomparable data, while transportation systems benefited from investigation structures that connect root causes to evidence-based measures. The report argues that voluntary, mandatory and citizen reporting each have limitations in isolation and recommends a hybrid model. [S5] The analogy does not prove what will work for AI. Aviation and healthcare have different liability structures, professional norms, regulators and public expectations. It shows why a repository is not an institution by itself.

The European regulatory path adds another complication. Article 73 of the EU AI Act creates mandatory serious-incident reporting for providers of high-risk AI systems, and the Commission has developed guidance and a reporting template intended to clarify definitions and interaction with other duties. [S6] SAFE says it would not replace legal or contractual obligations. The practical result could be complementarity, duplication or fragmentation. An event may be relevant to a voluntary agent-security exchange, a national authority, a privacy regulator, a customer contract and coordinated vulnerability disclosure at the same time.

That makes interoperability of decisions as important as interoperability of data. Enterprises need a route that determines who must be told, under which threshold, on what timetable, with which protected evidence and with what effect on later disclosure. A shared schema cannot resolve conflicting legal purposes by itself.

The safe-harbour paradox

The strongest case against SAFE is not that incident learning lacks value. It is that the proposal asks rational organisations to reveal information that can be used against them.

The more serious the failure, the more valuable the evidence may be to customers, regulators, litigants, attackers and competitors. Detailed traces can expose vulnerabilities, confidential data, internal controls or contractual breaches. Early uncertainty makes matters worse: an organisation may need to report before it knows whether the event was a model limitation, an operator mistake, a security incident or unauthorised conduct.

Independent reporting schemes can reduce this tension through confidentiality, de-identification, privilege, enforcement discretion or statutory protection. SAFE expresses principles of confidential review, learning rather than blame and member sovereignty. [S1] The Linux Foundation says the proposal respects existing legal, contractual and regulatory obligations. [S2] Those are design intentions, not legal protection.

Independent reporting also requires independence that participants can observe. Who appoints reviewers? Who funds the work? Who can compel evidence, exclude a member or publish a dissenting analysis? How are affected users represented? What happens when an influential member disputes a finding? The RFC proposes neutrality and a broad constituency, but the institution has not yet demonstrated these capabilities.

Coverage is equally material. Cybersecurity Dive reported that OpenAI and Anthropic were not members of the alliance and described traction as uncertain. [S3] That fact should not be exaggerated: absence at launch does not prove future refusal, and alliance membership is not the same as SAFE reporting participation. It does expose the representation problem. A learning system built around only a subset of model, tool and deployment ecosystems risks producing controls optimised for the participants it can see.

The paradox is stark. The organisations with the richest incident evidence may have the strongest reasons to withhold it. If SAFE cannot change that calculation, its most conscientious members could supply the majority of reports while the highest-consequence failures remain private.

What would distinguish an institution from industry theatre

The proposal should be judged by observable milestones, not by the number of organisations associated with the wider alliance.

  1. Governance becomes specific. A final charter defines appointment, funding, conflicts, confidentiality, affected-party participation and publication rights.
  1. Protection becomes credible. Members can explain what is confidential, what must be disclosed, what legal or contractual safeguards apply and where no protection exists.
  1. Participation becomes representative. Model providers, deployers, tool vendors, researchers, critical-infrastructure operators and affected-user representatives participate under the same reporting compact.
  1. Evidence becomes interoperable. Reports connect agent identity, authority, task, tool use, environment, human intervention and external outcome through stable identifiers and documented schemas.
  1. Reviews produce contested findings. Independent analysis can distinguish model behaviour, integration weakness, operator action and organisational failure without defaulting to a vendor narrative.
  1. Controls travel. Tests, policies, detection rules or reference configurations are adopted outside the originating incident and their effect is measured.
  1. Recurrence becomes visible. The institution can show whether a class of failure is repeating, declining or moving elsewhere in the stack.

Until those milestones appear, claims that SAFE is adopted, representative, legally protected or proven to reduce incidents would outrun the evidence. The honest judgement is narrower: the RFC is a credible governance prototype with materially useful detail and no demonstrated operating results.

The enterprise readiness test

An organisation does not need to join SAFE to learn from its design. It can run a controlled exercise against the evidence and decision model the RFC presupposes.

Select one agent with permission to act outside a sandbox. Create a plausible cross-boundary scenario involving a tool, identity, supplier or external target. Do not test whether the team can produce a polished incident narrative. Test whether it can reconstruct and govern the event.

Ask for:

  • the agent, model, harness and tool versions;
  • the originating instruction and complete action trajectory;
  • the workload identity and effective permissions at each consequential step;
  • the approved purpose, task, target and time boundary;
  • evidence from the environment in which impact occurred;
  • human alerts, interventions, approvals and kill authority;
  • supplier and customer notification decisions;
  • a common timeline and an evidence custodian;
  • one remediation linked to a repeatable test; and
  • proof that the control works after the change.

Then attempt to prepare the proposed initial report within four business days. The result will show whether the enterprise has an agent incident capability or only a collection of local logs.

The test should also surface where the argument does not apply. A closed internal automation with low autonomy and no third-party effect may not justify an external exchange. A sector with stronger mandatory reporting may require a regulator-first process. An organisation without basic observability cannot compensate by joining a reporting forum. Intentional criminal conduct does not become suitable for non-punitive treatment merely because an agent was involved.

Beyond private failure

SAFE’s lasting contribution may not be the exchange it proposes. It may be the standard of explanation it demands.

For years, AI assurance has concentrated on what a model can produce under evaluation. Agent incidents force a different account: what an authorised system did through real tools, across real boundaries, under changing instructions, with consequences recorded by more than one organisation. That account requires identity, authority, evidence custody and decision rights to meet in one operating model.

If credible protection and representative participation emerge, SAFE could shorten the distance between one firm’s failure and another firm’s defence. If they do not, the RFC can still serve as a procurement and governance benchmark. It tells leaders what to require before an agent is allowed to act beyond the organisation’s immediate control.

The useful decision is therefore not whether to endorse an unproven institution. It is whether the enterprise can convert an agent’s behaviour into evidence, an incident into a finding and a finding into a verified control. Until it can, failure remains private in the least useful sense: fragmented among parties, difficult to reconstruct and almost impossible to learn from together.

Sources

  1. Open Secure AI Alliance — Shared AI Findings Exchange (SAFE) RFC — 4 August 2026 — https://github.com/OpenSecureAIAlliance/RFCs/blob/main/rfc-safe-proposal.md
  1. Linux Foundation — Proposing the SAFE Working Group: An Open Community Effort to Improve AI Security — 4 August 2026 — https://www.linuxfoundation.org/blog/proposing-the-safe-working-group-an-open-community-effort-to-improve-ai-security
  1. Cybersecurity Dive — Tech industry alliance proposes AI agent safety reporting program — 4 August 2026 — https://www.cybersecuritydive.com/news/ai-agent-security-exchange-linux-foundation/826940/
  1. National Institute of Standards and Technology — Summary Analysis of Responses to the Request for Information Regarding Security Considerations for AI Agents — 18 May 2026 — https://www.nist.gov/publications/summary-analysis-responses-request-information-regarding-security-considerations-ai
  1. Center for Security and Emerging Technology — An Argument for Hybrid AI Incident Reporting — March 2024 — https://cset.georgetown.edu/publication/an-argument-for-hybrid-ai-incident-reporting/
  1. European Commission — AI Act: Commission issues draft guidance and reporting template on serious AI incidents — 26 September 2025 — https://digital-strategy.ec.europa.eu/en/consultations/ai-act-commission-issues-draft-guidance-and-reporting-template-serious-ai-incidents-and-seeks

More from Transformation