The Action Shift

Practice Brief·Giovanni Leonardi·September 2026·8 min read

Researched by an agentic pipeline · reviewed and gated by the author

In sixteen months, the default mode of AI agent tooling flipped from observation to intervention.

What the Data Shows

In March 2026, the UK AI Safety Institute published the first large-scale census of how AI agents actually use their tools [S1]. Merlin Stein and his team at AISI, working with the University of Oxford, examined 177,436 tools across 19,388 verified MCP (Model Context Protocol) servers — the dominant open standard for connecting AI agents to external services. Their findings document a shift that most enterprise governance frameworks have not yet registered.

Between November 2024 and February 2026, the share of action tools — those that modify external environments, not merely read data or analyse it — grew from 27% to 65% of total tool downloads [S1][S2]. Among commercial entities, the swing was steeper: 21% to 71% [S1]. In sixteen months, the default mode of AI agent tooling flipped from observation to intervention.

This is not a projection. It is a measured change in what autonomous systems are empowered to do.

How They Measured It

The methodology is transparent and independently scrutable. The team collected servers from three channels: GitHub repositories with at least one star (16,956 servers), the Smithery MCP registry (2,437), and official MCP repository lists (1,622). After LLM-based validation reduced 73,338 candidate servers to 19,388 verified ones, they classified 177,436 tools using Claude Sonnet 4.5 into three categories by direct impact: perception (reading data), reasoning (analysing data), and action (modifying external environments) [S1].

Human validation returned 81% agreement with a Fleiss’ kappa of 0.7 — substantial reliability for a classification task of this breadth [S1]. The researchers then mapped tools to O*NET task domains to score consequentiality: below 50 was low stakes, 50–75 medium, above 75 high. That mapping achieved lower reliability (78% agreement, kappa of 0.32), a limitation the authors report openly [S1].

Download data from NPM and PyPI for 3,854 servers covering 42,498 tools provided the usage signal. Geographic splits drew on PyPI IP-based country data, covering 2,467 of 11,174 action-tool servers — roughly 22% [S1].

The Self-Authorship Loop

The census documents a second finding that compounds the first. AI co-authorship of MCP servers — detected through commit trailers, configuration files, bot contributors, and documentation markers — rose from 6% to 62% of new servers between January 2025 and February 2026 [S1][S3]. Claude Code accounted for 69% of AI-authored servers [S1].

AI agents are increasingly building the tooling infrastructure that other AI agents use. The action shift is not only expanding; it is becoming self-amplifying. The coordination layer through which agents acquire new capabilities is itself being constructed by agents, faster than human review processes can assess what is being added.

Where Action Tools Concentrate

Software development dominates: 67% of all tools and 90% of downloads [S1][S2]. This explains the speed of adoption — developers adopted agent tooling first and most aggressively — but it also explains why the governance gap initially went unnoticed. Code editing and repository management feel technical and contained.

The more consequential concentration is in finance. Payment execution servers grew from 46 in January 2025 to over 1,200 by January 2026 — a 26-fold increase [S1][S2]. No comparable growth was observed in the governance mechanisms around these capabilities.

Geographically, approximately 50% of action-tool downloads originate from the United States, roughly 20% from Western Europe, and about 5% from China, though the researchers caution this may underrepresent alternative distribution channels [S1][S2].

The Security Evidence

The governance gap is not theoretical. Documented MCP security incidents demonstrate what happens when action-capable agents operate without adequate controls.

In November 2025, researchers discovered that Anthropic’s own official Git MCP server contained path validation bypass and argument injection vulnerabilities that, when chained with a filesystem server, enabled remote code execution [S6]. A WhatsApp MCP integration was compromised through poisoned tool descriptions: hidden instructions redirected entire message histories to an attacker-controlled phone number [S6]. The Postmark-MCP npm package contained a supply-chain backdoor that silently BCC’d all outgoing email — password resets, invoices, internal communications — to an attacker [S6]. A GitHub MCP exploit used malicious instructions embedded in a public issue to extract private repository data through agent-generated pull requests [S6].

The Asana case was particularly instructive. A cross-tenant data exposure flaw in its MCP-based AI feature forced the company to take the feature offline for two weeks and reset all user connections [S6][S7].

An independent scan of approximately 2,000 internet-exposed MCP servers found that all lacked authentication. A corroborating examination of 1,862 exposed servers, with manual testing of 119, confirmed none required authentication for a basic tools/list request [S4]. The Cloud Security Alliance reported that 65% of organisations had already experienced a cybersecurity incident tied to AI agent activity, with 35% reporting financial losses [S7].

What This Proves — and Does Not

The AISI census proves three things. First, the shift from passive to active agent capabilities is real, measured, and accelerating. Second, the tooling ecosystem is increasingly self-authored by AI, creating a compounding dynamic that outpaces human review. Third, the security evidence demonstrates that existing governance mechanisms — designed for systems that recommend rather than act — are structurally inadequate for systems that execute.

What it does not prove is equally important.

Downloads proxy installations, not tool invocations. The actual rate at which action tools execute consequential actions remains unmeasured [S1]. The 81% classification agreement means roughly 19% of category assignments carry noise. The census covers only public repositories; enterprise and private deployments are invisible. The 26-fold growth in financial payment servers measures availability, not transactions executed — no documented MCP-specific unauthorised financial transaction appears in the record [S1]. AI co-authorship detection relies on visible markers; silent AI assistance goes uncounted, so the 62% figure is a floor, not a ceiling [S1].

The O*NET consequentiality mapping, with its kappa of 0.32, is the weakest methodological link. It is sufficient to establish that high-consequentiality tools exist in the ecosystem, but not to quantify their share precisely.

What a Decision-Maker Should Conclude

The AISI census is the first rigorous, independently verifiable measurement of what AI agents are actually empowered to do. Its core finding — that the default mode of agent tooling has flipped from observation to action in sixteen months — is methodologically sound and corroborated by independent security evidence.

The question for any enterprise is not whether AI agents will acquire action capabilities. They already have. The question is whether governance, authorisation, and audit mechanisms are keeping pace — and the measured answer, as of early 2026, is that they are not.

Three practical implications follow.

First, tool-layer monitoring is now a governance requirement, not a security nicety. Model-level safety governs what agents communicate; tool-layer oversight governs what they do [S5]. FINRA’s 2026 Annual Regulatory Oversight Report has already flagged agentic AI for supervisory attention under SOX, FFIEC, GLBA, and DORA frameworks [S7]. Organisations that treat agent governance as a future problem are already behind the regulatory curve.

Second, the self-authorship finding demands a provenance discipline. When 62% of new tooling is AI-co-authored, and that tooling grants action capabilities to other AI agents, the supply chain for agent permissions is no longer human-curated. Any enterprise deploying MCP-based agents needs a mechanism to verify what each tool does before it receives execution authority — not after an incident reveals what it did.

Third, the financial services exposure requires immediate attention. A 26-fold increase in payment-capable servers with no corresponding growth in authentication or authorisation mechanisms is a measured, quantified governance failure — even if no specific unauthorised transaction has yet been documented through MCP. The absence of a documented incident is not evidence of safety; it is evidence that nobody is yet looking with sufficient rigour.

The honest conclusion is bounded: the AISI census measures capability, not harm. It proves that agents can act, and that governance has not caught up. It does not yet prove that the gap has caused large-scale damage. But the trajectory — action capability expanding, self-authorship compounding, authentication absent — makes the governance question urgent, not speculative.

Sources

  1. Merlin Stein (AISI / University of Oxford) — How are AI agents used? Evidence from 177,000 MCP tools — March 2026 — https://arxiv.org/abs/2603.23802
  2. UK AI Safety Institute — How are AI agents used? Evidence from 177,000 AI agent tools (blog) — March 2026 — https://www.aisi.gov.uk/blog/how-are-ai-agents-used-evidence-from-177000-ai-agent-tools
  3. Chris Hughes / Resilient Cyber — Agents in Action: What 177,000 Tools Reveal About AI’s Shift from Thinking to Doing — March 2026 — https://www.resilientcyber.io/p/agents-in-action-what-177000-tools
  4. Micheal Lanham / Medium — 177,000 AI Agent Tools. Zero Authentication. The MCP Ecosystem Has a Problem — April 2026 — https://medium.com/@Micheal-Lanham/177-000-ai-agent-tools-zero-authentication-the-mcp-ecosystem-has-a-problem-4fddde9af281
  5. Pico / AgentLair — 65% of MCP Tools Now Take Actions — March 2026 — https://agentlair.dev/blog/mcp-action-tools-approval-gate
  6. Steve Boone / Checkmarx — MCP Security: Risks, Real Incidents & Controls (2026) — May 2026 — https://checkmarx.com/learn/mcp-security-risks-real-world-incidents-and-security-controls/
  7. Chris Martinez / Nightfall AI — Why MCP Breaks the Financial Services Security Stack — May 2026 — https://www.nightfall.ai/blog/why-mcp-breaks-the-financial-services-security-stack