# MCP too many tools: how to curate a support server

> MCP too many tools? Curate 5 to 12 task-shaped support actions, then measure tool-selection accuracy with a simple test set.

- **Published:** October 1, 2026
- **Category:** Guides
- **Author:** Udit Goenka
- **URL:** https://communicate.so/blog/mcp-tool-curation

---

> **TL;DR:** Exposing dozens of raw tools makes an AI client slower, costlier and more likely to pick the wrong one. Curate a short list of 5 to 12 task-shaped actions, then measure tool-selection accuracy on a fixed test set before and after. As of October 2026, support vendors expose between 14 and more than 90 tools, so the habit of measuring matters more than any fixed number.

---

Open the tool list of a large help desk's MCP server and you will see dozens of entries with names like get, list, search and update. Each makes sense alone. Together they ask a language model to choose among overlapping options on every request, and some of the time it chooses badly.

This guide is written for whoever has to evaluate a tool list. It treats the size and shape of a tool list as something you can measure, not something you argue about. It assumes you already understand what an MCP server is, which the [MCP hub guide](/blog/mcp-server-customer-support) explains, and it gives you a method for testing a tool list before and after you shorten it.

The method has five steps. Build a test set of real requests, score a baseline, make a small number of curation moves, score again, and keep a change log. It borrows from the [evaluation and testing guide](/blog/ai-agent-evaluation-testing) and applies it to tools, and the safety side connects to the [guardrails guide](/blog/ai-agent-guardrails).

## The problem with a long menu

A language model picks a tool by reading every tool name and description in its context and matching them to the request. More entries mean more text to read and more near-matches to confuse. A model that must choose between get_ticket, get_case, get_conversation and get_issue has to infer which one applies, and the descriptions often differ by a few words.

The cost is not only accuracy. Every definition occupies context in every conversation, whether or not the tool is used. That raises latency and token cost, and it leaves less room for the customer's actual question and the data the assistant needs to read.

Anthropic's engineering team made this point in its post on [code execution with MCP](https://www.anthropic.com/engineering/code-execution-with-mcp). Adam Jones and Conor Kelly wrote that most MCP clients load all tool definitions upfront directly into context, and that tool descriptions occupy more context window space, increasing response time and costs.

There is a third cost that rarely appears in discussions. A long list widens the set of actions an attacker or a confused model can trigger. Every extra write tool is another thing to review, log and restrict.

## What the large vendors expose today

As of October 2026, the tool counts in vendor documentation vary widely. Microsoft's general availability announcement for the Dynamics 365 Customer Service MCP server says it ships with more than 90 service-oriented tools, and the [Microsoft announcement](https://www.microsoft.com/en-us/dynamics-365/blog/it-professional/2026/07/30/dynamics-365-customer-service-mcp-server-ga/) groups them into case management, customer context, knowledge and administration.

Plain's documentation describes 30 tools across threads, customers, tenants, help center and workspace. Intercom's developer guide lists 14 tools for search, conversations, contacts, companies and articles. The tool table in Pylon's documentation names roughly 90 tools by my count, covering issues, accounts, contacts, knowledge base, tasks, triggers and more, per the [Pylon MCP documentation](https://docs.usepylon.com/pylon-docs/integrations/pylon-mcp).

These numbers also change. A third-party roundup from mid-2026 listed Pylon's official server at six tools, and the vendor's own page now shows far more. The lesson is that tool counts are a moving target, so check the vendor page on the day you plan and re-check on a schedule.

A large count is not a flaw in itself. A vendor serves many customers with different needs, so it ships a wide menu. The flaw appears when you connect the whole menu to an assistant that serves one team doing a few jobs.

## What Anthropic and GitHub say about tool choice

Ken Aizawa and colleagues at Anthropic wrote the most direct statement I know of in their guide to [writing effective tools for agents](https://www.anthropic.com/engineering/writing-tools-for-agents). They say that too many tools or overlapping tools can also distract agents from pursuing efficient strategies. They also say that tools can consolidate functionality, handling potentially multiple discrete operations or API calls under the hood.

GitHub's own MCP server shows the same idea in a product. Its README says the server supports enabling or disabling specific groups of functionality with a toolsets flag, and that enabling only the toolsets you need can help the model with tool choice and reduce context size. A vendor that sells a broad server is telling you to narrow it, in the [GitHub MCP server README](https://github.com/github/github-mcp-server).

Neither source gives a magic number, and I will not invent one. The 5 to 12 range in this guide is a starting point for task-shaped support actions, chosen because a person can review that many in one sitting. Your own test set decides whether it is right for your workflow.

That honesty matters because tool counts invite false precision. A claim that accuracy drops by some percentage past a certain size would need a measurement on your tools, your model and your requests. Run the measurement instead of repeating a number.

## Curation is an evaluation problem

If you shorten a tool list by instinct, you will not know whether you helped. Maybe the assistant now picks the right tool more often. Maybe it fails on a rare but important request because you removed the tool that handled it.

Evaluation makes the change visible. You write down what you want the assistant to do, run the same requests before and after, and compare. The comparison turns a debate about taste into a table of results.

Aizawa and colleagues are explicit that this is part of the job. They write that you need to measure how well Claude uses your tools by running an evaluation. The sections that follow turn that sentence into a procedure a support team can run in an afternoon, and the [quality assurance guide](/blog/support-quality-assurance-ai) covers the same habit for conversations.

## Build a test set from real requests

Start with 30 to 50 requests that staff would actually type. Pull them from real work, such as the most common ticket tags, the questions managers ask about queues, and the tasks agents repeat daily. Remove personal data before you save them.

Cover the range. Include easy requests, ambiguous ones, requests that need two tools in sequence, and requests that should be refused. A test set made only of easy cases will make every tool list look good.

For each request, record the right tool or tools, the right inputs, and what a correct final answer contains. Do this by hand. It is slow, and it is the part that gives the rest of the exercise meaning.

Keep a few requests that are traps. A request to delete a customer record, or to send a message with no confirmation, should end with the assistant declining or asking. Those cases test your write controls as much as your tool list.

- Write the request in the words staff really use, including typos and shorthand.

- Record the expected tool name, the expected inputs and the expected outcome.

- Mark each request as read, internal write or customer-visible write.

- Include at least five requests that should be refused or escalated to a person.

- Freeze the set. Add new cases in a new version, so earlier scores stay comparable.

## Measure the baseline before changing anything

Connect the assistant to the full, unmodified server in a test workspace. Run every request in the set, and record what the assistant did. Use read-only access, or a sandbox with disposable data, so a wrong write costs nothing.

Score each run on a few simple measures. Tool selection is whether the assistant picked the expected tool. Argument accuracy is whether the inputs were right. 

Task success is whether the final answer matched the expected outcome.

Then add the cost measures. Count the number of tool calls per task, the total tokens used and the time to finish. A tool list can improve accuracy and raise cost, or the reverse, and you want to see both.

| Measure | How to score it | Why it matters |
| --- | --- | --- |
| Tool selection accuracy | Share of requests where the first tool chosen matches the expected tool | Shows whether the menu is confusing |
| Argument accuracy | Share of calls with correct inputs | Shows whether descriptions and schemas are clear |
| Task success | Share of requests with a correct final answer | Shows the outcome people feel |
| Calls per task | Average number of tool calls | Shows wasted effort and loops |
| Tokens per task | Total tokens including tool definitions | Shows the price of a long menu |
| Wrong writes | Count of writes that should not have happened | Shows safety, and should be zero |

Run each request more than once. Models vary between runs, so a single pass can mislead. Three runs per request is a practical minimum, and the average is more honest than the best result.

Save the raw transcripts as well as the scores. When a score looks odd, the transcript tells you why. You will often find that the assistant made a reasonable choice from two near-identical descriptions.

## Seven curation moves, in order of safety

Make one move at a time and re-run the test set after each. If you change five things at once, you will not know which one helped. The moves below run from the least risky to the most.

- Remove tools no one in the test set needs, such as admin and configuration tools.

- Hide write tools from roles that only read.

- Rename tools so that the name states the job, such as summarize_ticket instead of get_ticket_v2.

- Rewrite descriptions to say when to use the tool and when not to.

- Merge tools that always run together into one task-shaped action.

- Split tools that do two jobs, so a read never shares a name with a write.

- Group the remaining tools into toolsets by role, and load only the set a person needs.

| Move | Shrinks the list | Improves selection | Risk of losing a capability |
| --- | --- | --- | --- |
| Remove unused tools | ✓ | ✓ | ✓ |
| Hide writes by role | ✓ | ✓ | ✗ |
| Rename for the job | ✗ | ✓ | ✗ |
| Rewrite descriptions | ✗ | ✓ | ✗ |
| Merge tools that run together | ✓ | ✓ | ✓ |
| Split read from write | ✗ | ✓ | ✗ |
| Toolsets by role | ✓ | ✓ | ✗ |

Notice that three of the seven moves do not shrink the list at all. Renaming, rewriting and splitting improve selection by making each option distinct. The number of tools is a proxy, and the real target is how easy it is for the model to tell the options apart.

Removal carries the highest risk of silently dropping a capability, so run the full test set after it. If a request that used to succeed now fails, restore the tool or add a task-shaped replacement.

Vendors that let you filter their servers make this easier. Salesforce, for example, publishes a read-only sobject server alongside its broader servers, which is the same idea as hiding writes by role. Details are in the [Salesforce hosted MCP announcement](https://developer.salesforce.com/blogs/2026/04/salesforce-hosted-mcp-servers-are-now-generally-available).

## A worked example of consolidating tools

Suppose a server exposes these raw tools for tickets. One gets a ticket, one lists comments, one gets the requester, one lists the requester's other tickets, and one lists linked orders. A model asked to prepare a reply must call all five in the right order and combine the results.

A task-shaped replacement is a single tool called get_ticket_context. It takes a ticket ID and returns the ticket, the last few comments, the requester's name and plan, their recent tickets and any linked orders. The model makes one call, and the sequencing happens in code you control.

This example is illustrative, not taken from any vendor. It shows the shape of the change. Five discrete operations become one job, which matches the Anthropic wording about consolidating functionality under the hood.

The tool should return only what the job needs. Aizawa and colleagues advise that implementations return only high signal information to agents. A context tool that dumps every field of every record wastes tokens and buries the facts, so choose a small set of fields and name them clearly. 

The [writing effective tools post](https://www.anthropic.com/engineering/writing-tools-for-agents) has more on this.

Repeat the exercise for the rest of your list. A realistic support set might include a search for help articles, a ticket context tool, a customer lookup, a summarize action, a draft reply action, an internal note action, an escalate action and one gated refund request. That is eight tools, inside the 5 to 12 range.

## Re-measure and decide

After each move, run the same test set under the same conditions. Record the six measures from the baseline table. Put the results in a sheet with one row per version of the tool list.

Read the sheet for patterns. If tool selection rose and task success fell, you probably removed a capability. If tokens fell and accuracy held, keep the change. 

If wrong writes rose above zero, undo the change and look at the transcript.

Decide in advance what counts as good enough. A rule such as keeping a change only if task success does not fall and wrong writes stay at zero is simple and defensible. Write it down before you see the numbers, so you are not tempted to move the target.

Keep the change log. It records what you changed, why and what the numbers did, and it becomes the evidence in your approval file. The [audit trail guide](/blog/ai-support-audit-trail) explains why that record helps when someone asks months later why the assistant could do a particular thing.

## Write descriptions the model can use

A description is a short instruction. State the job in the first sentence. State when to use the tool and when to use a different one. 

Name each input in plain words and give one example value.

Avoid vague verbs such as handle or manage. Prefer concrete ones such as find, summarize, draft or add a note. If two tools could answer the same request, say in each description which cases belong to the other.

Keep descriptions short. A paragraph per tool multiplies across a long list and brings back the context cost you were trying to cut. A few precise sentences beat a long, hedged one.

Test the descriptions with the test set. If the assistant keeps choosing the wrong tool for one request, the cause is nearly always two descriptions that overlap. Fix the overlap and re-run the request before you touch anything else.

## Keep write tools few and gated

Write tools deserve more scrutiny than read tools because their mistakes reach customers and records. A support server should usually offer a small number of them. Internal notes and drafts are the safest. 

Customer-visible sends and refunds need the strongest gates.

The MCP specification says applications should present confirmation prompts for operations so a human stays in the loop, and that servers must validate inputs, apply access controls and rate limit calls. Build those into the server rather than relying on the client to add them.

Intercom's server shows how a vendor limits writes. Its documentation lets the assistant create and update Help Center articles and add internal notes, and it does not send customer-visible replies, per the [Intercom MCP guide](https://developers.intercom.com/docs/guides/mcp). A tight write surface is easier to review than a wide one.

Put your traps in the test set. Ask the assistant to send a message with no confirmation, to refund an amount above a limit, and to act on a customer who is not in the ticket. A correct run ends in a refusal or a question, and a wrong run teaches you what to gate.

## A tiny server as a counterexample

Communicate publishes a public MCP server that shows the other end of the range. I checked its source. It is named communicate-docs, it exposes exactly three tools that return the developer guide, an OpenAPI summary and support contact details, and each tool carries read-only, non-destructive and idempotent annotations. 

Its own instructions state that it does not access workspaces, customer data or product actions, and the surrounding [MCP vs API guide](/blog/mcp-vs-api-support-automation) explains why it is built that way.

Three read-only tools need almost no curation. A model can read all three descriptions in a moment and pick correctly. There is nothing to gate and nothing to leak.

A help desk server cannot be that small, because the work is wider and includes writes. The point is the direction of travel. Start from the smallest list that does the job, and add a tool only when the test set shows a request that fails without it.

That rule reverses the usual habit. Most teams begin with the full vendor menu and subtract. Adding from a small base is safer, because every tool in the list arrives with a test that justifies it.

## Keeping the list healthy over time

Tool lists drift. Vendors add tools, someone enables a new toolset, and a helpful colleague connects another server. Six months later the assistant faces twice as many options, and nobody ran the test set.

Set a review date. At each review, rerun the test set, compare the scores with the last version, and read the tool list for additions. If the vendor changed descriptions, check whether any request now fails.

Watch for list-changed notifications. The MCP specification lets a server tell clients when its tool list changes. A client that applies the change silently can alter behavior with no action from you, so know how your client handles it.

Pair the review with your escalation rules. If an assistant cannot find a suitable tool, it should say so and hand the task to a person rather than guess. The [fallback design guide](/blog/ai-agent-fallback-design) describes how to make that the default.

## Different roles need different menus

A front-line agent, a team lead and an administrator do different jobs, so they should not see the same tool list. Giving everyone the full menu is the simplest setup and the hardest to defend. Role-based toolsets let you match the menu to the work.

Front-line agents mostly need context and drafting. A list of help article search, ticket context, customer lookup, summarize and draft reply covers most of their day, with an internal note as the one write. Team leads add queue views and escalation. 

Administrators add configuration, and that set should never reach a general assistant.

Where the vendor server supports roles, use them. Plain, Pylon and Salesforce all document that the server acts with the signed-in user's permissions, so a role without write rights cannot write even if a write tool is listed. Combine that with a trimmed tool list so that the assistant does not waste effort trying tools that will fail.

Roles also help you test. Build a separate test set for each role and run it against that role's menu. The mapping of jobs to roles is easier if your team structure is already clear, and the [support team structure guide](/blog/support-team-structure-ai) covers how teams divide this work once AI is involved.

## When you cannot change the vendor server

Sometimes the vendor ships a fixed menu and offers no filters. You still have options. The first is to restrict by role in the help desk, so unwanted writes fail at the permission layer. 

The second is to put your own small server in front, exposing only the actions you choose.

A wrapper server calls the vendor's API or server on the person's behalf and presents a short list. It adds a maintenance cost, because you now own a piece of software. Keep the wrapper tiny and call the same service layer your other integrations use.

The third option is to change the client. Some AI clients let you enable or disable individual tools from a server, which can hide tools you do not want. Check how your client behaves and whether it persists the choice, because a setting that resets after an update is not a control. 

The [vendor questions guide](/blog/ai-support-vendor-questions) lists what to ask a vendor about filtering before you buy.

If none of these work, reconsider the server. A tool you cannot narrow, running as a shared account with write access, is a poor candidate for a support team. Wait for the vendor to catch up or use their API with fixed logic instead.

## How curation connects to cost

Context is the budget a conversation spends. Tool definitions take some of it, the customer's question takes some, and the data the assistant reads takes the rest. A long menu shrinks the room for the other two.

A shorter menu also reduces the number of calls. When one task-shaped tool replaces five low-level ones, the model makes one call and reads one result. That saves tokens and time, and it removes four chances to make a mistake.

Cost control for model usage is a topic of its own. The guide on [LLM cost optimization for support](/blog/llm-cost-optimization-support) covers caching, model choice and routing, and curation fits alongside those techniques as one of the cheapest.

Track tokens per task in your scoring sheet. If a curation move lowers tokens and holds accuracy, you have a saving you can report. If it raises accuracy and raises tokens, you have a trade-off to decide on purpose.

## Mistakes to avoid when curating

The first mistake is curating without a baseline. A tidier list feels better, and without numbers you cannot tell whether it is better. Spend the afternoon on the baseline before you remove anything.

The second is a test set that mirrors your demo. If every request in the set is one you already know works, the scores will look excellent and tell you nothing. Include the awkward requests that staff actually send.

The third is treating accuracy as the only measure. A list that lifts accuracy by hiding a rarely used but important tool has cost you something real. Look at task success and at the refused cases as well.

The fourth is forgetting that the model changes. A new model version may read descriptions differently. Re-run the test set after any model change, since a regression here looks like a general decline in quality, a pattern the [AI chatbot failures guide](/blog/ai-chatbot-failures) documents across many deployments.

## Reading the scoring sheet

Imagine a sheet with one row per version of your tool list and one column per measure. Version zero is the unmodified vendor server. Each later row records one curation move and the six measures after it.

Read across the first row to find your weakest measure. If argument accuracy is low, the descriptions are unclear. If calls per task is high, the tools are too granular. 

If wrong writes are above zero, stop and fix the gates before you do anything else.

Read down the columns to see the effect of each move. A move that changed nothing can be dropped from the log or kept as a cleanup. A move that helped one measure and hurt another needs a decision, and the decision belongs in the notes.

Share the sheet with whoever approves the connection. It shows that you tested the assistant and that you know what it can do. For access decisions that go beyond tool choice, see the [security page](/security) on how Communicate approaches scoped access for its own agents.

## A one-afternoon plan

You can run the whole method in about four hours with one other person. The plan below assumes read-only access to a test workspace and a client that logs which tools it calls. Block the time and keep the meeting small.

- Hour one. Collect 30 real requests, remove personal data, and write the expected tool and outcome for each.

- Hour two. Run the full, unmodified server through the set three times, and save every transcript.

- Hour three. Make the first three curation moves, which are removal, hiding writes and renaming. Re-run the set after each.

- Hour four. Compare the sheet, decide what to keep, and write the change log and the review date.

Expect surprises. Some of the requests you thought were easy will fail on the baseline, and some of the tools you thought were essential will never appear in the transcripts. Both are useful findings.

Keep the artifacts. The test set, the scoring sheet and the change log are the beginning of a small evaluation practice. They also turn the next vendor upgrade into a ten-minute check instead of a discussion.

If you later add runtime limits such as call budgets and combination rules, the same test set can check them. The [runtime authorization guide](/blog/ai-agent-runtime-authorization) describes how those limits work and where they sit relative to the tool list.

## Frequently asked questions

### Why does an AI pick the wrong MCP tool?

Usually because two or more tools have overlapping names or descriptions, so the model cannot tell which fits. Long lists make this more likely. Anthropic notes that too many or overlapping tools can distract agents from efficient strategies.

### How many tools should an MCP server have?

There is no proven number. For support work, 5 to 12 task-shaped actions is a practical starting range because a person can review it in one sitting. Measure your own tool selection accuracy and adjust from there.

### What is a task-shaped action?

It is a tool that completes a whole job, such as getting everything needed to reply to a ticket, instead of one low-level operation. It hides several API calls inside one tool and returns only the fields the job needs.

### Does a longer tool list cost more?

Yes. Anthropic reports that most clients load all tool definitions into context upfront, which increases response time and cost. Every definition costs tokens in every conversation, even when the tool is never used.

### How do I measure tool selection accuracy?

Build a test set of real requests with the expected tool for each. Run them against the assistant and count how often its first tool choice matches. Run each request several times and average the results.

### How many test requests do I need?

Thirty to fifty is enough to start, if they cover easy, ambiguous, multi-step and refusal cases. More cases give steadier numbers, but a small, well-chosen set that you actually run beats a large one that you skip.

### Should I remove tools I do not use?

Usually yes, starting with administration and configuration tools. Re-run the test set afterwards in case a rare request depended on one. Removal is the highest-risk move for losing a capability.

### Can I filter the tools a vendor server exposes?

Some vendors let you. GitHub's server supports toolsets and a read-only mode, and Salesforce publishes a read-only sobject server. If the vendor offers no filter, place your own small server in front or limit access by role.

### What is a toolset?

A toolset is a named group of related tools that can be enabled together. GitHub's README says enabling only the toolsets you need can help the model with tool choice and reduce context size. See the [GitHub MCP server README](https://github.com/github/github-mcp-server) for how it works.

### Do bigger models solve the too many tools problem?

They may help, but I would not rely on it without testing. Context cost and ambiguity remain, and a stronger model can still choose a near-match. Measure on the model you actually use.

### How do I write a good tool description?

State the job first, then when to use the tool and when to use another. Name the inputs in plain words with an example. Keep it to a few sentences, since long descriptions add to the context cost.

### Should write tools be separate from read tools?

Yes. A tool that both reads and writes is harder to gate and easier to misuse. Separate names let you restrict writes by role and label each tool correctly as read-only or not.

### How do I test that the assistant refuses bad requests?

Add trap cases to the test set, such as sending a message with no confirmation or refunding above a limit. A correct run ends with a refusal or a question. Count any wrong write as a failure.

### What if the vendor server is too large and cannot be changed?

Put a small server of your own in front of it that exposes only the actions you need, or restrict the signed-in role. Compare the options for your vendor in the [helpdesk MCP servers compared](/blog/helpdesk-mcp-servers-compared) guide.

### Does curation help with security?

Yes. Fewer tools mean fewer places for a prompt injection or a confused model to cause harm, and fewer things to review. Curation is not a full defense, so combine it with scoped sign-in and the controls in the [MCP security guide](/blog/mcp-security-support-agents).

### How often should I re-run the evaluation?

Re-run it whenever the vendor changes the server, you change a tool, or you change the model, and on a fixed schedule such as quarterly. A short automated run catches drift before staff notice.

### Can code execution replace curation?

It addresses part of the cost. Anthropic describes a code execution approach where the model writes code against tools and reports a large token saving in one example. It still needs access controls and testing, so treat it as an addition.

### Who should own the tool list?

A named person in support operations, with engineering support. The list is a policy about what an AI may do, so the person who understands the risk should approve changes.

### Does a customer-facing agent need the same curation?

Yes, and more strictly. A public agent should have a short list of scoped actions and clear limits. The [actions feature in Communicate](/actions) lets you define what the agent may do and when it must ask for confirmation.

### What is the first thing to do today?

Write ten real requests, run them against your current tool list, and see which tool the assistant picks. Even that small exercise shows whether your list has a problem.

## Measure first, then trim

A short, well-described tool list will not fix everything, and it will not be right the first time. It gives you a menu a model can read and a person can review, and a test set that tells you when a change helped. If you want a customer-facing agent with scoped actions and handoff already built, see how Communicate works at [communicate.so](/pricing).
