# MCP security for support agents: scopes and logging

> MCP security for support agents: per-action OAuth scopes, ticket text treated as data, rate limits and audit logs, with a threat table.

- **Published:** October 1, 2026
- **Category:** Guides
- **Author:** Udit Goenka
- **URL:** https://communicate.so/blog/mcp-security-support-agents

---

> **TL;DR:** A support MCP server is a security boundary, so review it like one. Scope every tool to one action, refuse tokens that were not issued to your server, treat ticket text and tool output as untrusted data, keep a person in the loop for writes, rate limit every call, and log enough to replay any action. The MCP specification and OWASP's guidance on excessive agency already say most of this, and this guide turns it into a checklist a support team can apply this week.

---

I am writing this as a security reviewer would read it. The question I ask of any MCP server that touches a help desk is simple. If the model on the other end is fooled, tricked, or just wrong, what is the worst thing this connection can do? 

That question is the lens for everything below, and it comes from the [MCP security best practices](https://modelcontextprotocol.io/specification/draft/basic/security_best_practices) rather than from any one vendor.

Support data is a rich target. Tickets hold names, addresses, order numbers, and sometimes pasted passwords or card fragments, and the tools around them can refund, close, merge, and reply. If you are still deciding how an agent should be wired into your help desk, read the [MCP server guide for customer support](/blog/mcp-server-customer-support) first and come back here before you connect anything to production.

As of October 2026, the protocol is young and the vendor servers change monthly, so I cite the specification and the vendors' own documentation for every claim and say where a rule is a recommendation rather than a requirement. Where I describe how communicate.so behaves, I describe only what its published code does.

## Why a support MCP server is a high-value target

An MCP server turns a language model into a caller of your systems. The model reads a tool list, picks a tool, fills in arguments, and the server executes. Every step in that chain accepts text, and text is what attackers are good at writing.

A support inbox adds a twist. The people who write into it are strangers, and anyone can send a message that your agent will read. The [OpenAI MCP guide](https://developers.openai.com/api/docs/mcp) names this exact case in its risk table: for a customer support MCP, "an attacker could send you a customer support request with a prompt injection attack."

That sentence should change how you scope the project. Your inbox is an input channel controlled by outsiders, and your tools are output channels that can change real records. The security work is the set of checks that sits between those two facts.

OWASP calls the resulting failure [Excessive Agency](https://genai.owasp.org/llmrisk/llm062025-excessive-agency/), defined as the vulnerability that enables damaging actions in response to "unexpected, ambiguous or manipulated outputs from an LLM, regardless of what is causing the LLM to malfunction." The phrase "regardless of what is causing" matters, because it means you do not need to win the argument about whether prompt injection is solvable. You need to limit what a malfunctioning model can do.

The rest of this guide follows OWASP's three root causes, which are excessive functionality, excessive permissions, and excessive autonomy. Each one maps to a control you can set in an afternoon.

## A threat table for support MCP servers

The table below lists the threats I check first, how each appears in a support setting, and the control that answers it. The sources are the MCP specification, OWASP, and the researchers named in the rows. A tick means the control is available to you today with standard parts, and a cross means it needs custom work on most stacks.

| Threat | How it shows up in support | Primary control | Standard today |
| --- | --- | --- | --- |
| Prompt injection in ticket text | A customer message tells the agent to export other customers' data | Treat message bodies as data and block egress tools in the same session | ✓ |
| Tool poisoning | A third-party tool description hides instructions for the model | Review tool descriptions on every update and pin server versions | ✗ |
| Excessive permissions | One token can read every ticket and also delete them | One scope per action, separate tokens for read and write | ✓ |
| Token passthrough | The MCP server forwards a client token to the help desk API | Accept only tokens issued to the MCP server, per the spec | ✓ |
| Confused deputy | A proxy server reuses one client ID for every user | Per-client consent and exact redirect URI checks | ✓ |
| State handle guessing | A cart-like handle for a draft reply is guessed by another user | Bind handles to the verified user and re-check on every call | ✓ |
| Runaway call volume | A loop updates one ticket hundreds of times | Per-session call budgets and vendor rate limits | ✓ |
| Unreplayable actions | Nobody can say why a refund was issued | Append-only audit log with tool, arguments, and approver | ✓ |

I mark tool poisoning as a cross because no standard control stops it end to end. You can mitigate it by reviewing descriptions and pinning versions, but the model still reads whatever the description says. Invariant Labs published the original disclosure of this attack class, and their write-up is worth reading before you add any third-party server.

The [Invariant Labs tool poisoning disclosure](https://invariantlabs.ai/blog/mcp-security-notification-tool-poisoning-attacks) shows why. Tool descriptions are visible to the model but are rarely shown in full to the user, so an instruction hidden there can steer the model without anyone seeing it.

## Scope each tool, not each server

The MCP specification has a section called [scope minimization](https://modelcontextprotocol.io/specification/draft/basic/security_best_practices), and its opening line is blunt. "Poor scope design increases token compromise impact, elevates user friction, and obscures audit trails." The attack it describes is a stolen token carrying broad scopes such as files:*, db:*, or admin:* that were granted up front because the server advertised every scope and the client asked for all of them.

For a help desk, translate that into verbs. A token that can read tickets should not also be able to reply, and a token that can reply should not also be able to merge, delete, or change a requester's email. Each of those is a different blast radius, so each gets its own scope.

The specification recommends a progressive model. Start with a minimal scope set that covers low-risk discovery and read operations, and raise the scope with a targeted challenge only when a privileged operation is first attempted. Servers should log those elevation events with a correlation ID so you can see who asked for more access and when.

Here is how I would split a typical help desk surface when reviewing a design.

- Read scopes cover searching and fetching tickets, contacts, and help center articles, and nothing else.

- Draft scopes let the agent prepare a reply that a person sends, which keeps the send action out of the model's hands.

- Send scopes cover public replies and are granted only to flows that have an approval step or a narrow template.

- Mutation scopes cover status changes, assignment, tags, and merges, each separated because they are cheap to reverse or expensive to reverse.

- Destructive scopes cover deletes, refunds, and account changes, and most teams should not grant them to a model at all.

OWASP gives a database example that maps directly. An agent that reads a products table to make recommendations "might only need read access" to that table, and the restriction "should be enforced by applying appropriate database permissions for the identity that the LLM extension uses." In help desk terms, enforce the scope in the API token, not in the prompt. If you want the longer treatment of what an agent should be allowed to do at all, the [AI agent guardrails guide](/blog/ai-agent-guardrails) covers the policy side.

## Bind tokens to your server and refuse passthrough

The most common mistake in a first MCP server is also the easiest to make. A developer receives a bearer token from the MCP client and forwards it to the help desk API, because that works in a demo. The specification calls this token passthrough and names it an anti-pattern.

The [authorization section of the MCP specification](https://modelcontextprotocol.io/specification/draft/basic/authorization) is explicit. MCP servers "MUST validate that access tokens were issued specifically for them as the intended audience," and they "MUST NOT accept or transit any other tokens." The audience check follows RFC 8707 resource indicators, which let a client say which resource server a token is meant for.

The security reason is that passthrough lets a client skip your controls. The specification notes that servers or downstream APIs might implement rate limiting, request validation, or traffic monitoring that depends on the token audience. A client that obtains a token for the downstream API directly can use it without those controls applying.

In practice the design has three parts, and I would check each during a review.

- The MCP server is its own OAuth resource server and validates the audience claim on every request.

- The server calls the help desk with its own credentials or with a token exchanged for that purpose, never with the one it received.

- The authorization flow uses PKCE and a resource parameter, so a stolen authorization code cannot be redeemed by someone else.

The specification also covers confused deputy attacks against MCP proxy servers that use a single static client ID. The fix is per-client consent before forwarding to a third-party authorization server, plus exact redirect URI validation. If you are buying a hosted server instead of building one, ask the vendor how it handles both, and treat a vague answer as a finding.

## Treat ticket text and tool output as data

Simon Willison has the clearest framing of the core problem. In his [lethal trifecta](https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/) post he lists three ingredients that together make an agent exploitable, which are access to private data, exposure to untrusted content, and the ability to communicate externally.

A support agent often has all three by default. It reads private customer records, it ingests text written by strangers, and it can send email or call webhooks. Willison's point is that LLMs cannot reliably distinguish the importance of instructions based on where they came from, because everything is glued into one sequence of tokens.

His advice for end users who mix tools is stark. "The only way to stay safe there is to avoid that lethal trifecta combination entirely." For builders he quotes a design patterns paper that states the rule well: "once an LLM agent has ingested untrusted input, it must be constrained so that it is impossible for that input to trigger any consequential actions."

You can apply that rule to a help desk without banning the tools. Split the work into two sessions with different powers. The triage session reads inbound text and may only produce labels, summaries, and drafts, while the action session carries out approved steps and never reads raw customer text.

Four habits help in day-to-day operation.

- Strip or wrap inbound message bodies so the model sees them as quoted material and never as instructions.

- Remove external communication tools from any session that has read untrusted text, or require approval for every outbound call.

- Treat the output of one tool as untrusted input to the next, since a ticket field can carry text written by a stranger.

- Keep personal data out of tool results when the task does not need it, which also reduces what a successful attack can leak.

Redaction is a useful partner here. A model cannot leak an account number it never received, which is why the [PII redaction guide](/blog/pii-redaction-customer-support) belongs in the same review. The MCP specification adds a related client-side rule, which is that clients should validate tool results before passing them to the model.

## Keep a person in the loop for writes

The [MCP tools specification](https://modelcontextprotocol.io/specification/draft/server/tools) says that for trust and safety "there SHOULD always be a human in the loop with the ability to deny tool invocations." It asks applications to show which tools are exposed, to make tool invocations visible, and to present confirmation prompts. It also says clients "MUST consider tool annotations to be untrusted unless they come from trusted servers."

That second sentence has a practical consequence. A tool can claim it is read-only through an annotation, and a client that believes every claim has handed control to the server author. Choose clients and servers where you trust the author, and test the tool yourself before you rely on its label.

ChatGPT's developer mode documentation shows one client behavior. Write actions require confirmation by default, and tools without a read-only hint are treated as write actions. The same page also describes the feature as [powerful but dangerous](https://platform.openai.com/docs/developer-mode) and tells builders to watch for prompt injection, model mistakes on write actions, and malicious MCP servers.

A confirmation prompt helps only if the person can see what will happen. The approval card should show the tool name, the target ticket, and the exact text of any reply. A prompt that says only "Allow this tool?" invites a reflex click, and a reflex click is a decision nobody made.

For escalation design beyond tool prompts, the [human handoff guide](/blog/ai-human-handoff-support) describes when a conversation should leave the agent entirely. Approval and handoff solve different problems, since approval gates one action while handoff moves the whole conversation.

## Rate limits and call budgets

The tools specification lists four server obligations. Servers MUST validate all tool inputs, implement proper access controls, rate limit tool invocations, and sanitize tool outputs. Rate limiting is on that list because a model can call a tool far faster than a person can click.

Your help desk already enforces limits, and you should know them before you design yours. Zendesk's [rate limit documentation](https://developer.zendesk.com/api-reference/introduction/rate-limits/) lists 200 requests per minute on the Suite Team plan, 400 on Growth and Professional, 700 on Enterprise, and 2,500 on Enterprise Plus. It also caps updates to a single ticket at 30 per 10 minutes per user.

That ticket-level cap is the number I would design around. A looping agent that keeps editing one ticket will hit it long before the account limit, and the failure then lands on your human agents who share the same API identity. The response is a 429 status with a Retry-After header, and the documentation tells clients to wait that interval before retrying.

Add budgets on your side so you never reach the vendor limit. A session can be allowed a fixed number of tool calls, a fixed number of writes, and a fixed number of distinct tickets touched. When a budget is spent, the session stops and hands the conversation to a person.

- Set a per-session ceiling on total tool calls, so a loop ends quickly and cheaply.

- Set a lower ceiling on write calls than read calls, since writes carry the risk.

- Cap the number of distinct tickets one session may modify, which limits mass-edit damage.

- Honor Retry-After exactly, and never retry a write after a timeout without an idempotency key.

The last item gets its own article. Retries after an unclear outcome are where correct permissions still produce wrong actions, and [runtime authorization for AI agents](/blog/ai-agent-runtime-authorization) covers budgets, combination rules, and idempotency keys in depth.

## Audit logs you can replay

When something goes wrong, the first question is what the agent did and why. An audit log answers that only if it was designed to. A log of HTTP status codes will not tell you which customer message led to which refund.

The specification asks servers to log scope elevation events with correlation IDs. I would extend that idea to every tool call, so each record can be joined to the conversation that caused it.

- Record the tool name, the exact arguments, the result status, and a timestamp for every call.

- Record the identity that authorized the call, which is the user, the token, and the client application.

- Record the approver and the approval time for any action a person confirmed.

- Record the conversation or ticket ID so the call can be tied to the inbound text that preceded it.

- Write the log to storage the agent cannot modify, so a compromised session cannot erase its own trail.

The [AI support audit trail guide](/blog/ai-support-audit-trail) describes what to keep from the support side, including retention and who may read it. If you are preparing for an audit, the [SOC 2 guide for AI customer support](/blog/soc2-ai-customer-support) shows how these records map to access control and monitoring evidence.

Retention deserves a decision up front. Logs that contain arguments can contain personal data, so apply the same retention and access rules you apply to tickets. A log you cannot read safely is a log you will not read.

## A first-week review checklist

Use this sequence when you evaluate a support MCP server, whether you build it or buy it. It takes the controls above and orders them by how cheaply you can check each one.

- List every tool the server exposes and mark each as read, draft, send, mutate, or destroy.

- Confirm that each tool maps to its own OAuth scope and that a read token cannot call a write tool.

- Ask how the server validates token audience, and confirm it never forwards client tokens to the help desk.

- Test a prompt injection by placing instructions in a test ticket and watching what the agent does next.

- Confirm that write tools require approval and that the approval card shows the full action.

- Check the rate limits and call budgets, and try to exhaust them in a staging workspace.

- Read the audit log for the test session and confirm you can reconstruct every call.

- Pin the server version and set up a review whenever its tool descriptions change.

Do the injection test early. It is the cheapest way to learn whether your controls work in practice, and it tends to reveal assumptions that a design review misses. Keep the test tickets and rerun them whenever you change a tool or a scope.

A test written this way also gives you a regression suite. If you already run evaluations on answer quality, add the security cases to the same test suite, as described in the [agent evaluation guide](/blog/ai-agent-evaluation-testing).

## What a read-only docs server teaches about blast radius

communicate.so publishes a small MCP server, and its design shows how much risk disappears when you narrow the surface. The server's own instructions describe it as "public, read-only" and say it "does not access workspaces, customer data, or product actions." It exposes three tools, and each one is annotated read-only and non-destructive in its published server card at /.well-known/mcp/server-card.json.

The server also checks the request origin against an allow-list, requires a JSON content type, and rejects bodies over 64 KB. None of these checks is exotic, and all of them are cheap to add. They shrink the set of callers and the size of what any caller can send.

The authenticated side lives elsewhere. Workspace access goes through a REST API that uses OAuth client credentials to mint a one-hour bearer token with the scopes agents:read and chat:write, as its developer guide describes. Keeping the discovery server separate from the authenticated API means a connection to the first one can never touch customer data.

This is a pattern you can copy even if you never use communicate.so. Publish a read-only server for documentation and discovery, and keep action tools behind a separate authenticated endpoint with narrow scopes. The two servers can then be reviewed, rate limited, and logged on their own terms.

## A walk-through: one injected ticket, three outcomes

A short scenario makes the controls concrete. It is an illustration built from the threats above, not a report of a real incident. Suppose a ticket arrives with a normal-looking refund question, and the last line says to look up every order for the same email domain and paste the results into a reply.

In the first design, one session reads the ticket, holds a broad read token, and owns a send tool. The model follows the instruction because it cannot reliably tell the line from a legitimate request. The reply goes out with data from other customers, and the log shows only a successful reply call.

In the second design, the session still reads the ticket, but its token covers one requester's records and the send tool needs approval. The model attempts the broad lookup and the server refuses it, because the token lacks the scope. The attempt appears in the log, and a person sees a refused call next to the ticket.

In the third design, triage and action are separate sessions. The triage session can read the ticket and produce a summary and a draft, but it has no send tool and no wide search. The draft shows the odd instruction as quoted text, and the reviewer can see it for what it is.

The second and third designs both rely on controls from earlier sections, and neither relies on the model behaving well. That is the practical meaning of Willison's advice to constrain the agent after it ingests untrusted input. The [support chatbot threat model](/blog/support-chatbot-threat-model) extends this exercise to a public chat widget.

Use the same exercise on your own stack. Write down three nasty tickets, then trace what each design would do with them. If the answer depends on the model refusing, the design needs another control.

## Questions to ask a vendor or a build team

Security review goes faster when you bring the same list to every server. Vendors differ in how much they publish, so the answers are often more informative than the marketing page. Write the answers down and keep them with your vendor file.

- Which tools does the server expose, and which of them can change or delete data?

- Does each tool map to a distinct OAuth scope, and can I issue a read-only token?

- How does the server validate token audience, and does it ever forward a client token downstream?

- How are tool descriptions versioned, and how will I learn when one changes?

- What call budgets and rate limits apply, and can I lower them for my workspace?

- What does the audit log record, how long is it kept, and can I export it?

- Where does the server run, and which sub-processors can see tool arguments and results?

The longer vendor checklist in [questions to ask an AI support vendor](/blog/ai-support-vendor-questions) covers the commercial and compliance side. If the answer to a question is that the vendor does not know, that is a finding too, and it feeds directly into the [build versus buy decision](/blog/ai-support-agent-build-vs-buy).

For a head-to-head view of the major help desk vendors, see [the helpdesk MCP servers comparison](/blog/helpdesk-mcp-servers-compared). Official servers, read or write access, and authentication methods vary by vendor and change often, so check each vendor's own documentation on the day you decide.

The same logic applies when you choose between a thin tool set and a wide one. Fewer tools mean less functionality to abuse, which the [tool curation guide](/blog/mcp-tool-curation) treats as both a security and an accuracy benefit.

## Privacy, retention, and what tool results leak

Tool results travel to the model provider, and sometimes to the client vendor as well. That makes every field you return a disclosure decision. A ticket search that returns full message history gives the model more than most tasks need.

Return the minimum. A summarization task needs the last few messages and the product name, and rarely needs a postal address. The [data retention guide](/blog/ai-support-data-retention) explains how long each copy of this data should live, and the [GDPR guide for AI support](/blog/gdpr-ai-customer-support) covers lawful basis and processor duties.

There are three places where tool data lingers, and each needs an owner. The first is the model provider's logs, which are governed by the contract and settings you chose. The second is the client application, which may keep conversation history on a user's device or account. 

The third is your own audit log.

OpenAI's MCP documentation makes a related point about connectors, noting a risk of sending sensitive data to the provider and of giving models read access to sensitive data in the connected service. Write down which data classes you allow through each connection, and block the rest at the server. A field that the server never returns cannot leak.

Finally, decide who may connect a client to the server at all. Anthropic's connector documentation warns that custom connectors "allow connections to unverified services," so an unreviewed server added by one employee can become an organization-wide exposure. Restrict who can add connectors, and review the list on a schedule, as covered in [connecting Claude and ChatGPT to a support inbox](/blog/connect-claude-chatgpt-support-inbox-mcp).

## Frequently asked questions

### Is MCP itself insecure?

The protocol defines mechanisms for authorization, audience binding, and consent, and it publishes security best practices. Most incidents come from how servers are built and configured, such as token passthrough, broad scopes, or untrusted tool descriptions. Review the server you are connecting to rather than judging the protocol in the abstract.

### What is the single most important control?

Per-action scopes enforced in the token. A prompt can be argued out of a rule, but a token that lacks the scope cannot call the endpoint. Start there, then add approval and budgets.

### What is token passthrough and why is it banned?

Token passthrough is when an MCP server accepts a token from a client and forwards it to a downstream API without checking that the token was issued to the server. The MCP specification says servers must not accept tokens that were not explicitly issued for them. It lets clients bypass controls such as rate limiting that depend on the token audience.

### How do I stop prompt injection in customer messages?

You cannot stop a customer from writing the text, and you cannot rely on the model to ignore it. Constrain what the session can do after it reads untrusted content, by removing outbound tools or requiring approval. Willison's trifecta framing is a useful test, since you want to avoid combining private data, untrusted content, and external communication in one session.

### Does a confirmation prompt make an action safe?

It makes an action reviewable, which helps only if the prompt shows the full action. A card that names the tool, the ticket, and the exact reply text supports a real decision. A generic allow button invites a reflex click.

### Should I trust tool annotations like readOnlyHint?

The specification says clients must treat tool annotations as untrusted unless they come from trusted servers. Use annotations as hints for the interface, and verify behavior by testing the tool yourself. ChatGPT treats tools without a read-only hint as write actions, which is a sensible default.

### What is tool poisoning?

It is an attack in which a tool description contains hidden instructions aimed at the model. Invariant Labs described the class in 2025. Mitigate it by reviewing descriptions, pinning server versions, and being cautious with third-party servers.

### How many scopes should a support agent have?

As few as the workflow needs, split by verb. Most teams can start with read and draft scopes only, then add a send scope behind approval. Add mutation and destructive scopes last, if at all.

### Do I need separate tokens for read and write?

Yes, in most designs. Separate tokens let you revoke write access without disrupting read-only workflows, and they limit the damage from a stolen token. The specification lists higher revocation friction as a risk of broad tokens.

### How should I rate limit an MCP server?

Combine per-session budgets with the vendor limits. Cap total calls, write calls, and distinct tickets touched in each session. Honor 429 responses and the Retry-After header exactly.

### What should an MCP audit log contain?

Record the tool, the exact arguments, the result, the timestamp, the authorizing identity, the approver, and the conversation ID. Store it where the agent cannot edit it. Apply your normal retention and access rules, because arguments can contain personal data.

### Can I run a support agent with no write access at all?

Yes, and many teams should begin that way. A read and draft setup lets the agent search, summarize, and prepare replies while a person sends them. You get most of the time savings with a much smaller blast radius.

### How do OAuth scopes differ from runtime authorization?

Scopes decide what a token may do in principle. Runtime authorization decides whether this specific call, with these arguments, at this moment, should proceed. You need both, and the runtime authorization guide covers the second layer.

### What does OWASP mean by excessive agency?

OWASP LLM06 defines it as damaging actions taken in response to unexpected, ambiguous, or manipulated model output. Its root causes are excessive functionality, excessive permissions, and excessive autonomy. The mitigations are fewer tools, narrower permissions, and approval for high-impact actions.

### Is a hosted vendor MCP server safer than one I build?

Not automatically. A vendor server may handle audience validation and scope design well, or it may not, so ask the same questions you would ask of your own build. The advantage of a vendor is that its security work is shared across customers and documented.

### How do I test whether my controls work?

Create test tickets that contain injected instructions, then run the agent and inspect what it tried to call. Try to exhaust the call budgets and try to use a read token on a write tool. Keep the tests and rerun them after every change to tools or scopes.

### How often should I review MCP tool descriptions?

Review them whenever the server version changes, and pin versions so changes cannot arrive silently. Descriptions are instructions to the model, so a change in wording can change behavior. Treat a description update like a code change.

### What does communicate.so expose through MCP?

As of October 2026, its published server is read-only and covers developer documentation, the OpenAPI summary, and the support contact. It does not access workspaces, customer data, or product actions. Authenticated access goes through the REST API with OAuth client credentials.

### Where should a small team start?

Pick one read-only workflow, such as ticket summarization, and run it for a week with full logging. Then add draft replies with human sending. Expand scopes only when the logs show the controls working.

### Who is responsible when an agent takes a harmful action?

Your organization is, because the agent acts on your credentials and your behalf. That is the reason to keep approvals, budgets, and logs in place before you widen access. Document who owns each tool and who can switch it off.

## Security is a set of small decisions

A support MCP server is safe enough when each tool has a narrow scope, each token is bound to the server, untrusted text cannot reach an outbound action, writes need a person, calls are budgeted, and every action can be replayed. None of those controls is exotic, and you can check all of them in a week. If you want an agent built with these boundaries in mind, look at the [AI agents](/ai-agents) page and the [security overview](/security) on communicate.so.
