The Tool Description Is Part of the System
Researched by an agentic pipeline · reviewed and gated by the author
A tool description is part of the executable control surface: changing it can change both task success and execution cost.
The interface the model actually sees
An MCP tool definition gives a model a name, a natural-language description and an input schema. The protocol standardises how tools are discovered and called, but the model still has to infer which tool fits the task and how its arguments should be populated. That makes the description part of the operating interface, not supporting documentation. [S6]
This distinction matters because a description can be syntactically valid and still steer an agent badly. An unclear purpose can send the model to the wrong tool. Missing usage boundaries can make it call a tool in the wrong situation. Opaque parameter guidance can produce a technically valid but operationally wasteful request.
A tool description is part of the executable control surface: changing it can change both task success and execution cost.
What two studies found
A Queen’s University research team examined 856 tools across 103 MCP servers. Its automated rubric assessed purpose, usage guidance, limitations, parameter explanations, completeness and examples. The authors reported that 97.1% of descriptions contained at least one identified defect, while 56% did not state the tool’s purpose clearly. Missing limitations, missing usage guidance and opaque parameters were even more common. Official servers did not show a statistically significant quality advantage over community servers on the measured components. [S1]
The team then changed tool descriptions and ran agents on MCP-Universe, a benchmark built around multi-step tasks using real MCP servers and execution-based evaluators. Fully augmented descriptions increased task success by a median 5.85 percentage points across the tested domain-model combinations and improved partial goal completion by 15.12%. Those are meaningful changes from an artefact that many engineering teams would treat as prose. [S1][S3]
A separate 2026 preprint tested description defects across 10,831 MCP servers. In controlled mutations, correcting functionality and accuracy defects changed tool selection by 11.6% and 8.8% respectively, with reported p-values below 0.001. When five functionally equivalent servers competed for selection, standards-compliant descriptions reached a 72% selection probability against a 20% baseline. The methods differ from the Queen’s study, but the direction is consistent: description quality affects agent behaviour. [S2]
The improvement was not free
The Queen’s study also recorded a 67.46% median increase in execution steps after full augmentation, and performance regressed in 16.67% of tested cases. More explanation could improve guidance while making trajectories longer. No single combination of description components performed best across every model and domain. Removing examples did not produce a statistically significant degradation across the tested combinations, and shorter targeted variants often matched the fully augmented descriptions. [S1]
The practical tension is therefore not good documentation versus bad documentation. It is semantic guidance versus context and execution cost. A description can be too thin to guide the model or too broad to justify the tokens and steps it consumes.
Anthropic has reported the same capacity problem from a builder’s perspective. In one five-server example, 58 tool definitions consumed about 55,000 tokens before work began; the company says it has seen definitions consume 134,000 tokens before optimisation. Its deferred Tool Search pattern loaded only relevant definitions on demand. Anthropic reported an 85% reduction in tool-token use and evaluation gains from 49% to 74% for Opus 4 and from 79.5% to 88.1% for Opus 4.5. These are interested-party, Claude-specific internal results, not independent production evidence, but they show why description design and tool discovery have to be engineered together. [S4]
What this proves — and what it does not
The evidence supports three bounded conclusions.
- Descriptions are behavioural inputs. They influence selection, parameterisation and multi-step execution, so changes deserve tests and release control.
- Completeness is not the objective. The tested optimum varied by domain and model. The correct target is the shortest description that preserves reliable behaviour for the relevant workload.
- Tool-set size changes the design problem. As connected tools multiply, loading every definition can consume context and make selection harder. MCP-Universe found performance declines when unrelated tools were added; for example, Claude 4 Sonnet’s location-navigation success fell from 22.22% to 11.11%, while GPT-4.1 browser-automation success fell from 23.08% to 15.38%. [S3]
The limits are equally important. Both description-quality studies are preprints. Their results come from controlled benchmarks, not longitudinal enterprise deployments. The available evidence does not establish a universal description template, a monetary return, a maintenance baseline or durability across model updates. Anthropic’s results are builder-reported. No independent production replication in the source set shows that the measured benchmark gains translate directly into lower incident rates or operating cost.
A production discipline for tool descriptions
Teams should manage tool descriptions as versioned interface artefacts with an evaluation contract.
- Start with explicit purpose, usage conditions, limitations and unambiguous parameter semantics. Enforce expected inputs and outputs with strict schemas where possible. Anthropic’s engineering guidance reaches the same conclusion from practice: clear tool boundaries, precise parameter names and evaluation-driven refinement improve tool use. [S5]
- Build representative tasks around actual user intents, including near-duplicate tools, ambiguous requests, invalid parameters and multi-step workflows.
- Record at least task completion, correct tool selection, argument validity, execution steps, token use and latency. A higher success rate that requires materially more calls may still be the wrong operating point.
- Test compact variants by model and domain. Do not assume that adding examples or every possible caveat helps.
- For large catalogues, evaluate deferred discovery or another retrieval layer so irrelevant definitions do not occupy the model’s working context.
- Release description changes with the tool implementation, model configuration and evaluation results tied together. When the model or tool catalogue changes, rerun the tests.
The decision test is simple: if a team cannot show how a description change affects task success, bad calls and execution cost on its own workload, it is not yet managing the agent’s tool interface as a production system.
Sources
- Queen’s University / arXiv — Model Context Protocol (MCP) Tool Descriptions Are Smelly! — 31 May 2026 — https://arxiv.org/html/2602.14878v3
- Wang et al. / arXiv — From Docs to Descriptions: Smell-Aware Evaluation of MCP Server Descriptions — 21 February 2026 — https://arxiv.org/abs/2602.18914
- Salesforce AI Research / arXiv — MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers — 20 August 2025 — https://arxiv.org/html/2508.14704v1
- Anthropic — Introducing advanced tool use on the Claude Developer Platform — 24 November 2025 — https://www.anthropic.com/engineering/advanced-tool-use
- Anthropic — Writing effective tools for AI agents — 11 September 2025 — https://www.anthropic.com/engineering/writing-tools-for-agents
- Model Context Protocol — Tools specification — 28 July 2026 — https://modelcontextprotocol.io/specification/2026-07-28/server/tools