Skip to main content

Your Deploy Agent Doesn't Need 100 Tools in Context: Progressive Discovery for Infrastructure MCP Servers

10 min readDora NodaDora Noda
Share
On this page

Three MCP servers. 143,000 tokens of tool definitions. That's 72% of a 200,000-token context window consumed before the agent has read a single word from the user. That measurement, from Stephanie Goodman's 2026 AgentPMT analysis of real MCP setups, is the failure mode every infrastructure control plane is walking into: connect a deploy agent to a 100-tool server and it must ingest the entire catalog — logs, rollbacks, DNS, env vars, all of it — before it knows which one it needs.

The fix the ecosystem converged on in 2026 is progressive discovery: expose a small read-first entry point, reveal deployment and destructive tools only after resource selection, and let the agent search the catalog instead of swallowing it. This post gives you the concrete design — a three-tier tool catalog for a Render-compatible infrastructure server — plus the numbers that justify it and a method to measure your own payoff.

The 100-tool cliff, in numbers​

Tool overload doesn't degrade agents gradually. It hits a cliff, and multiple independent benchmarks agree roughly where:

  • The RAG-MCP project measured tool-selection accuracy falling from 43% to under 14% as the tool catalog grew — a threefold silent degradation before any bad answer ships.
  • Nebula's benchmarks put a focused 4-tool setup at ~95% correct tool selection versus ~71% with a full 46-tool MCP toolset — a 24-point gap caused purely by context bloat.
  • Anthropic's internal guidance says selection accuracy "degrades significantly" past 30–50 tools, and documents ~55K-token context bloat from basic MCP setups.
  • TaskBench (NeurIPS 2024) showed graph accuracy dropping from 96% with one tool to 25% on 8-tool chains; MCPVerse (2025) found most models degraded with larger tool sets.

The token math explains why. Roughly 50 tools cost 10,000–20,000 tokens in definitions alone. A 100-tool infrastructure server therefore opens every session somewhere around 20,000–40,000 tokens in the hole — and the damage isn't just cost. Researchers call the reasoning-quality decay context-rot: a hundred tool schemas in context means the agent wades through all of them on every decision, and its focus on what actually matters weakens with each one.

Infrastructure servers hit this first because the mapping is so tempting: one REST endpoint, one MCP tool. Render's public API covers services, datastores, deploys, projects, and environments — model each operation naively and a Render-compatible surface sails past 100 tools without trying. The agent that only needs to tail logs pays the same upfront tax as the agent about to rewire DNS.


The design in one table​

Progressive discovery inverts the default. Instead of "everything visible, agent figures it out," the server exposes tiers and reveals downward:

TierNameWhat's visibleExample toolsRevealed when
0Read-first entry~5–8 tools + catalog searchlist_services, get_service, get_status, search_toolsAlways (session start)
1Resource-scopedTools bound to a selected resourcetail_logs, list_deploys, get_env_vars, describe_dnsAfter the agent names a service/environment
2Gated actionsDeploy + destructive toolstrigger_deploy, rollback, update_dns, rotate_secretAfter explicit intent + confirmation surface

Two rules make the tiers work:

  1. Nothing in Tier 1 or 2 enters context until Tier 0 has identified the target. The agent can always list and inspect; it cannot touch what it hasn't selected.
  2. Tier 2 tools arrive with their confirmation story attached. The reveal isn't just a schema — it's the schema plus the guardrail (dry-run output, blast radius, required confirmation), so the most dangerous tools are also the best-documented at the moment of use.

The search_tools meta-tool is the escape hatch that keeps the tiers from becoming a straitjacket: any tool in the catalog is one search away, but its full schema loads only when the agent actually reaches for it.


Three tasks, three reveals​

The spec's trio — logs, rollback, DNS — exercises all three tiers. Here's what each session looks like against a flat 100-tool server versus the tiered design.

Task 1: "Why is api-web erroring?" (read-only). Flat: 100 schemas load (~20–40K tokens) so the agent can call one log tool. Tiered: Tier 0 only — list_services → get_status → Tier 1 reveals tail_logs for that service. Full schemas in context: under 10. The read-only path, which dominates real agent-ops traffic, never pays for deployment tools at all.

Task 2: "Roll back api-web to the last good deploy." (destructive). Flat: same 100-schema tax, plus the rollback tool sits beside 99 distractors at selection time. Tiered: Tier 0 identifies the service, Tier 1 reveals list_deploys to find the last good one, and only then does Tier 2 surface rollback — with the deploy history it needs as arguments already in context. The dangerous tool arrives last, when the agent has maximum information and minimum distraction.

Task 3: "Point staging DNS at the new deploy." (config change). Flat: full catalog again. Tiered: Tier 0 → Tier 1 describe_dns shows current records → Tier 2 reveals update_dns scoped to that zone. Note what didn't load: database tools, secret rotation, autoscaling — none of them relevant, none of them in context competing for the agent's attention.

The pattern across all three: context grows with demonstrated intent, not with catalog size. A 200-tool server costs the log-tailing session exactly the same as a 50-tool server does.


What the spec actually says​

The MCP 2026-07-28 protocol revision made this official — with one placement decision that surprises server builders: progressive discovery lives primarily at the host/client layer, not the server layer. The client best-practices guidance recommends that hosts fetch tools/list as usual but defer injecting definitions into context, expose a lightweight search_tools meta-tool, and load full definitions only as needed — switching to this mode once definitions cross roughly 1–5% of the context window.

That threshold is already productized. Anthropic's Claude Agent SDK exposes it as ENABLE_TOOL_SEARCH=auto:5: tool search activates automatically when tool definitions exceed 5% of context. OpenAI's GPT-5.4+ generation added deferred loading and tool search at the API level too, and Anthropic reports >85% reduction in the ~55K-token bloat their basic MCP setups were carrying.

Two adjacent pieces from the same spec revision matter for infrastructure servers:

  • The Tasks extension (io.modelcontextprotocol/tasks, SEP-2663) moved long-running work — exactly what a deploy or a rollback is — into a first-class protocol extension. A tiered server should return a task handle from Tier 2 actions, not block the request path.
  • Cache hints (ttlMs/cacheScope) travel on tools/list results, so the catalog metadata the client does fetch can be cached across sessions instead of refetched.

Server builders sometimes read "client-layer" as "not my problem." It's the opposite: the client can only defer what the server makes deferrable. A server that dumps 100 undifferentiated tools into tools/list with no search surface forces every host to improvise. A server designed for progressive discovery gives hosts tiers, search, and hints to work with.


Server-side patterns that work today​

You don't need to wait for universal client support. These patterns ship now, on both sides of the 2026-07-28 transition:

  1. Hide bulk tools from tools/list; serve them through search_tools. The Equinix docs MCP server runs this in production: API tools stay out of the default listing and are discovered on demand, with cache hints on the listing. This is the single highest-leverage change — it converts the flat tax into pay-as-you-search on every host, including older ones.
  2. Mark deferrable tools and set an auto-threshold. Claude Code's MCP connector supports mcp_toolset defer_loading; the Agent SDK's auto:5 flips on tool search at 5% of context. If your server ships client config or setup docs, prescribe these instead of leaving users on flat loading.
  3. Serve both protocol eras from one deployment. FastMCP 4.0 — built for the 2026-07-28 stateless, sessionless protocol — negotiates per-client, dual-serving handshake-era (2025-11-25) and new-era clients. Key migration notes: background tasks moved to the optional fastmcp-tasks package (@mcp.tool(task=True) raises unless the Tasks extension is registered), and SEP-2577 deprecates Roots, Sampling, and Logging. Pin accordingly and don't strand older hosts during the transition.
  4. Keep Tier 2 schemas strict and Tier 0 schemas tiny. Deferred loading only helps if the always-loaded surface is genuinely small. Push enums, field details, and examples into searchable descriptions or resources — the cesteral/mcp-open-advertising pattern of small inputSchemas plus nextAction hints in errors lets the model self-recover without preloading everything.
  5. Learn from the big-catalog operators. GitHub's 101-tool and Financial Modeling Prep's 253-tool MCP setups both recovered accuracy through deferred-loading architecture after hitting the same cliff. A 64-tool setup consumes roughly 30% of Claude's context window — past the ~20-tool reliability threshold — so if you're anywhere near that size, you're already paying.

A minimal server config sketch captures the posture:

text
tools/list (default surface):
  - list_services, get_service, get_status   # Tier 0: always loaded
  - search_tools                             # catalog search meta-tool
  - cache hints: ttlMs=3600000, scope=server
 
hidden from default listing, via search_tools:
  - tail_logs, list_deploys, get_env_vars…   # Tier 1: on resource select
  - trigger_deploy, rollback, update_dns…    # Tier 2: on explicit intent

Measuring the payoff​

Don't take the published numbers on faith — they bound your expectations, but your catalog's schemas and your agents' tasks set the real figures. Run the comparison yourself: same task suite, flat loading versus progressive, and record both cost and correctness.

Use the published points as sanity bounds. Token math first, extrapolated from the ~200–400 tokens per tool definition the field reports:

Catalog sizeFlat: definitions in context (est.)Progressive entry surface (est.)
20 tools~4–8K tokens~2–3K tokens
50 tools~10–20K tokens~2–3K tokens
100 tools~20–40K tokens~2–3K tokens

The progressive column is flat by construction — that's the whole point — while the flat column grows linearly into the 1–5% auto-threshold and past it. On accuracy, expect the published shape: near-focused-toolset selection (~95% at 4 tools, Nebula) degrading toward ~71% at 46 tools and 43%→under-14% at RAG-MCP scale, with the 30–50-tool cliff (Anthropic) marking where flat loading starts failing silently.

Your benchmark needs three ingredients:

  1. A task suite drawn from real agent-ops traffic — weighted toward read-only (status, logs) since that's the majority, with a destructive minority (rollback, DNS, env changes). Score tool-selection correctness per task, not just final-answer quality: wrong-tool-right-answer hides the rot.
  2. Both loadings on identical tasks. Flat run first to establish your baseline error rate at your catalog size, then progressive. The delta is the deliverable.
  3. A threshold policy from the results. If your Tier 0 surface stays under ~1% of context and Tier 2 reveals stay scoped to one resource, you've matched what the spec authors had in mind — and what Anthropic's >85% bloat-reduction figure says is achievable.

The uncomfortable finding this benchmark usually surfaces: most teams discover their agent was already past the cliff. The errors were just silent — slightly wrong tool, plausible-looking output, nobody diffing the choice. Progressive discovery doesn't only save tokens; it makes the remaining tool choices auditable, because a five-tool decision is reviewable in a way a hundred-tool decision never is.


A 100-tool infrastructure server isn't a failure of scope — a Render-compatible API genuinely has that many operations. It's a failure of presentation if all 100 arrive in context at once. Tier the catalog, search the rest, gate the destructive tools behind demonstrated intent, and the deploy agent stays as sharp at 100 tools as it was at 10.

Bex.co is the open-source, AI-native Render alternative — push a git repo, get a running HTTPS service on machines you own. Star the repo on GitHub or deploy your first app today.

Related articles

Give your agents a chain backend

Autonomous agents hit RPC endpoints very differently than people do. See what bex router handles on their behalf.

Read the agents guide