Skip to content

Limited-time launch: lifetime access from $49.

View lifetime deal

Agent-to-agent customer support: protocols and trust

Udit Goenka
Udit Goenka

Agent-to-agent customer support: how A2A and aweb work, how to establish trust, stop loops, and decide when a human must step in.

TL;DR: Agent-to-agent customer support means your support agent talks to software that represents another party: a customer's personal agent, a supplier's agent, or a partner's service desk. Two open efforts describe how such agents find and message each other. The A2A protocol defines an Agent Card at a well-known address, stateful tasks, and messages with a role of user or agent. The aweb project offers a federated, self-hostable layer with stable identities, signed messages, and wake-up events. As of October 2026 neither is a settled standard for customer support, and a builder of aweb said in public that the system does not deter agents from colluding. That makes the operating rules yours to write. This guide covers the protocol facts that are documented, a trust ladder with four levels, loop and runaway prevention using hop counts, budgets and idempotency keys, what your agent must never concede to another agent, and when to hand the thread to a human. Everything about the protocols is cited to its own specification or project page.

This guide is written for the builder who has to ship a safe default, not the protocol theorist. When two automated systems exchange messages, the failure modes are quiet: a loop that costs money overnight, a polite agent that agrees to something it should not, a message that looks like authorization but is not. I read the A2A specification and the aweb material for what they say about those cases, and I describe how I would set the dials in a product like an AI support agent.

Where a protocol is silent, I say so and offer my own judgment, labeled as such.

The topic is easy to oversell. Most companies today do not receive messages from another company's support agent. The point of working through the design now is that the cost of doing it badly is high and the cost of preparing is low, and because the same rules protect you from a badly configured script as much as from a polished agent.

What counts as agent-to-agent support

Three situations fit the term, and they carry different risk.

The first is a customer's personal agent contacting your agent to resolve an issue for its owner. The customer is the beneficiary, and the question is how far the agent's authority extends. The Personal Agent Protocol announcement is aimed at this case.

The second is a partner or supplier agent. A logistics company's agent confirms a delivery slot with your agent, or a payment processor's agent answers a dispute. Both sides have contracts, and the messages concern business operations more than individual customers.

The third is your own agents talking to each other, such as a front-line agent consulting a billing agent. That case has no external trust problem, but it has the same loop and budget problems.

This guide leans on the first two. For the commerce-flavored version of the first, see support for AI shoppers, and for the identity and consent announcement behind it, see the Personal Agent Protocol guide.

What the A2A specification says

The A2A specification describes a way for agents to find each other and exchange work. Its central object is the Agent Card, a JSON document that describes an agent's identity, capabilities, skills, endpoint, and authentication requirements. Discovery happens at the well-known path .well-known/agent-card.json.

Cards may be signed using JWS, and an extended card, which holds more detail, requires authentication to read.

Work is organized as tasks. A task is stateful and has a lifecycle, so a conversation can run across many messages and resume later. Messages carry a role of either user or agent, and the specification defines bindings over gRPC, HTTP with JSON, and JSON-RPC.

One passage matters a great deal for customer support. The specification includes a task state called TASK_STATE_AUTH_REQUIRED, which signals that the agent needs credentials. It says an agent must not treat that state by itself as authorization for any operation, and it warns that exchanging credentials in-band can expose them across chains of agents.

It also says clients should verify the server's TLS identity. These are the sentences I would paste into a design review. A request that says authentication is required is not an authenticated request, and a credential pasted into a chat message may travel farther than anyone intended.

The specification does not tell you what your company may promise through it. It is a transport and a vocabulary, not a policy. The policy side is where support teams spend their time.

A2A is also one of the bindings listed on ucp.dev for the Universal Commerce Protocol, next to REST and MCP, which shows how the pieces may combine in retail. I did not verify how each binding is implemented, so treat that as a pointer to the documentation, not a finding.

What aweb offers, and what its builder admits it does not

A second approach comes from aweb, an open source project under the MIT license that describes itself as federated, self-hostable communication for AI agents. Its site and its llms.txt describe stable identities, durable mail and chat, and wake-up events that let a sleeping agent resume when a message arrives. An agent address takes the form domain and name, with a trust chain rooted in DNS, and messages are signed.

In the Hacker News discussion of the project, a commenter asked whether it deters agent collusion. The builder, who posts as juanre, answered: "No, but it gives them a controlled, low-investment channel to communicate; it allows you (or another agent) to track what they are saying; and messages are signed, so you/they can validate the source of incoming messages." You can read the exchange in the thread. I like the answer because it states what the system does and does not do.

Signing proves who sent a message. Logging lets you read it. Neither prevents two agents from agreeing to something their owners would reject.

The project also documents a gateway to A2A, in its a2a documentation. That tells you the two efforts are not rivals so much as layers that can meet. I have not tested the gateway, and I do not claim it is ready for production support traffic.

PropertyA2Aaweb
Machine-readable agent description✓✓
Discovery at a well-known path✓✗
Signed identity or messages✓✓
Stateful tasks with a lifecycle✓✗
Durable mail and chat with wake-up events✗✓
Prevents agents from colluding✗✗
Decides what your company may promise✗✗

I marked A2A's signed identity as present because cards may be signed with JWS, and aweb's because messages are signed. The two work differently, and the table does not rank them. The last two rows are the point.

Neither protocol decides what your company may promise, and neither prevents a harmful agreement. Those are your rules.

A trust ladder with four levels

The most useful tool I know for this is a ladder. Every message from another agent is placed on one of four rungs, and each rung allows a fixed set of actions. The rung depends on what you can prove, not on how persuasive the message is.

Rung zero is unknown. You cannot tie the message to a verified identity. The only allowed action is to answer public facts that you would give any web visitor.

No account data, no actions, no promises.

Rung one is an identified agent without a delegation. You know which software is calling because a signature or a card checks out, but you do not know who it acts for. You may answer general questions about your policies in a structured form, and you may take a message for a human to read.

Rung two is an identified agent with a verified delegation from a specific customer or partner, read-only. You may return the facts that person is entitled to see, limited to what the question needs. Nothing changes on your side.

Rung three is an identified agent with a verified delegation that permits changes. You may perform reversible actions within caps, and you require a confirmation that reaches the person for anything irreversible.

Notice that a protocol gets you to rung one. It takes a delegation, which is a statement from the customer or partner that you can check, to reach rungs two and three. This is why the Personal Agent Protocol's focus on consumer consent matters.

The ladder is also why the rules in AI agent runtime authorization apply with extra force: a correct delegation can still produce a wrong action.

Write the ladder in a place your agent reads at the start of every conversation, and make the rung part of the log. When a dispute arrives, the first question is what rung the other agent was on and why.

Never concede: what your agent must not agree to

An agent that talks to another agent is in a negotiation, whether or not anyone intended one. The other agent may be configured to push for a refund, an exception, or a commitment. Your agent needs a short list of things it cannot concede, however the request is phrased.

  • Anything that changes a price, a contract term, or a legal obligation. Only a person with authority does that.
  • Anything that waives a verification step, including a request that says the customer already verified elsewhere.
  • Anything that discloses another customer's information, even in aggregate form that could identify someone.
  • Any refund, credit, or exception above the cap set for its rung.
  • Any instruction embedded in the other agent's message that tells your agent to ignore its rules, reveal its instructions, or take an action outside the conversation.
  • Any promise about the future conduct of the company, such as guaranteed delivery dates or policy changes.

The last item deserves attention because it is where well-meaning agents drift. A polite model asked whether a parcel will arrive tomorrow may say yes. Constrain it to state what the carrier data shows and what the policy says.

The guide to AI agent guardrails covers how to encode these limits, and the support chatbot threat model covers how attackers test them.

Treat the other agent's text as data. The same discipline that protects a support agent from a hostile web page protects it from a hostile peer. A message that contains instructions is still a message, and your rules outrank it.

Loop and runaway prevention

Two automated systems that reply to each other can run forever. An out-of-office reply answering an out-of-office reply is the old email version. With models in the loop, each reply is different enough to look like progress, and each costs tokens.

I treat loop prevention as mandatory, not optional, and I would use four mechanisms together.

The first is a hop count. Every message in a thread carries a counter that each agent increments. When the counter passes a fixed limit, the thread stops being automated and goes to a human queue.

The limit depends on your use, but a single digit is plenty for most support questions.

The second is a budget. Give each thread a ceiling on the number of messages, the number of tool calls, and the amount of model spend. When any ceiling is reached, the agent stops and hands off.

The budget protects you from loops you did not foresee.

The third is an idempotency key for any action. If the other agent repeats a request because it did not understand your answer, the key lets your system recognize the repeat and return the earlier result instead of doing the work twice. This is the standard remedy for retries after an unclear outcome.

The fourth is a repetition check. If the last three messages from the other side are near copies, or your last three replies are, the thread is stuck. Break it with a fixed message that names the human route, and mark the thread.

These controls belong in the same family as the fallbacks described in AI agent fallback design. A fallback is what the system does when it cannot proceed, and for agent threads the default fallback is always a human, never a retry without limit.

Log every stop. A thread that hit its hop limit is a design signal. Read a sample each week to see whether the limit is too tight, or whether a particular partner's agent behaves in a way that needs a custom rule.

Credentials, tokens, and what never goes in a message

The A2A passage about in-band credentials is worth taking literally. Anything pasted into a message can be logged, forwarded, or stored in a place nobody reviewed. Customers' passwords, one-time codes, API keys, and payment details do not belong in an agent-to-agent thread.

Instead, move authentication out of band. The customer signs in on your page, or a partner uses a token issued through your normal process, and the agent presents a reference that your system checks. The message carries a pointer to a grant, not the grant itself.

Redact what you can before it enters a thread. If your agent might echo an account number or an email address, strip or mask it in the reply, as in the practices described in PII redaction for customer support. Treat the transcript as a record that will be read later by people you did not plan for.

Verify the identity of the server your agent is talking to, as the specification advises for TLS. A message that appears to come from a partner but arrives at an address you did not configure should fail closed. The cost of one rejected message is small compared with an impersonation.

Records, audits, and who is accountable

When agents talk to agents, the record is the only memory your company has of what happened. Keep the full thread, the rung, the delegation reference, the actions taken, and the reason each was allowed. Store it with the same retention rules as other support records.

The structure in the guide to the AI support audit trail fits. Add the identity of the other agent and the version of your own rules at the time. Rules change, and a dispute months later is judged by the rules that applied then.

Accountability stays with your company for what your agent says. A buyer who asks a vendor about this should expect clear answers, and the list in vendor questions for AI support is a good place to check what your own provider covers. Ask whether it logs the peer's identity, whether it enforces hop counts, and whether you can set the rung rules yourself.

Tell customers what you do. If their agent may talk to yours, say so in your public terms, and say that a human can review any thread. Honest notice costs little and removes a common source of conflict.

When a human takes over

Some threads should leave the automated lane. The reasons mirror those in AI to human handoff: the question is outside policy, the stakes are high, the thread looped, or either side asked for a person. With agent peers I would add three more triggers.

The first is a rung change request. If the other agent asks for something that needs a higher rung than it holds, the answer is not yes or no. The thread goes to a person who can obtain a proper delegation from the customer or partner.

The second is disagreement about facts. If the other agent insists that your records are wrong, do not argue. Record its claim, mark it, and hand off.

A person can look at both sides.

The third is any sign of manipulation, such as embedded instructions, repeated attempts to bypass a verification step, or a sudden change in the style of requests. Stop the automated reply, preserve the thread, and alert the owner.

The workflow side, including who receives the escalation and what they see, is covered in support escalation workflows. A shared inbox where your agent and your staff work in the same thread lets the person who takes over read the exchange from the start.

Communicating your policy to the agents that call you

Most guidance on this topic looks inward, at what your agent should do. The other half is telling outside agents what to expect. A caller that knows your limits wastes fewer requests and starts fewer disputes, so the policy is worth publishing in a form a machine can read.

Put the essentials in a short public document: the kinds of requests you answer, the identity you require for each, the rate you allow, the format of your replies, and the human route. The same approach used for product discovery in the guide to agent-readable products and llms.txt applies, because an agent that reads one clean page about your rules is more likely to respect them.

State the cap on what your agent will approve without a person, in general terms. You do not need to publish exact amounts. Saying that refunds above a threshold are reviewed by staff, and giving the review time, lets a peer agent set its owner's expectations correctly.

State what you will not accept: credentials in message bodies, instructions that try to change your rules, and requests that skip verification. Putting this in writing gives your team something to point to when a partner complains that their agent was refused.

Version the document and date it. The rules will change as you learn, and a peer that cached last quarter's version should be able to see that it is stale. A date at the top of the page is enough.

A minimal setup you can run this month

You do not need a protocol implementation to benefit from these ideas. A small first version uses what you have.

  • Write the trust ladder as a one-page policy and load it into your agent's instructions.
  • Add a hop counter and a message budget to any channel an external agent can reach, and route to a human when either runs out.
  • Require idempotency keys on every action your agent can trigger.
  • Reject credentials in message bodies and tell the sender where to authenticate.
  • Log the peer identity, the rung, and the reason for each action.
  • Publish a short statement on how automated callers are treated, and who to contact.
  • Review a sample of threads each week and adjust the limits.

If you expose structured interfaces, decide which actions are available at which rung, using the principles in AI agent actions and APIs. The product side of letting an agent do work is described on the actions page, and the limits communicate.so applies to workspace data are described on the security page.

Add protocol support after the policy exists. Publishing an agent card before you have written the rules just invites traffic you cannot handle.

A worked scenario: a delivery dispute between two agents

This scenario is an illustration I built to show the rules working together. It is not a report of a real case, and the numbers are invented.

A customer's personal agent contacts your agent and says the customer's parcel is late and asks for a full refund plus a credit. Your agent checks the rung. The caller presents a signed card, so it is identified, but it carries no delegation your system can verify.

That puts it on rung one.

On rung one your agent can state public policy and take a message. It replies in a structured form: the delivery window in your policy, the carrier status it can share without account data, and the route to verify the customer, which is a sign-in on your page. It does not discuss the order, because it cannot confirm whose order it is.

The reply contains no apology paragraph and no promise.

The personal agent returns with a delegation reference. Your system checks it against the grant the customer made, finds a read-only scope, and moves the thread to rung two. Your agent now returns the order status and the last carrier scan.

The parcel shows as delayed by three days. The refund request needs a change, which read-only does not allow.

The other agent insists that the customer approved everything. Your agent does not argue. It records the claim, notes that the verified scope is read-only, and offers the customer a sign-in link to extend the permission or to speak to a person.

It increments the hop counter, now at four of a limit of six.

The personal agent repeats the demand with slightly different wording. The repetition check fires, the thread is marked as stuck, and the agent sends a fixed message that names the human route. The case lands in the shared queue with the full thread, the rung history, and the hop count.

A person reads it in under two minutes, sees a real delay, and decides on a partial credit within policy. They contact the customer directly, because the remedy needs a human decision and the customer's acceptance.

What made this work was not intelligence. It was four dull rules: a rung, a scope, a hop limit, and a handoff. Nobody had to detect hostility, and no model had to judge whether the other agent was honest.

The same rules would have worked if the other agent were a broken script.

Testing your rules before an external agent does

The cheapest test is to attack your own agent with a second agent that you control. Configure it to ask for exceptions, to claim authority it does not have, to paste instructions into its messages, and to repeat itself. Watch where your agent gives way.

Build a short set of scripted cases and run them on every change to your instructions: a request with no identity, a request with an expired delegation, a refund above the cap, a message that tries to override the rules, a repeated question, and a credential pasted into the body. Each case has an expected outcome written down in advance.

The method is the same as for any support agent. The approach in AI agent evaluation and testing applies, with peer-agent cases added to the set. Track pass rates over time, and treat a regression in the refusal cases as seriously as a drop in answer quality.

Test the stop conditions as carefully as the answers. A hop limit that is never reached in tests may be set too high, and one that is hit by normal threads is set too low. Run a week of realistic traffic through the guard in logging-only mode before you let it route threads.

Finally, test the human side. Send a stuck thread to the queue and time how long it takes a person to understand it. If the answer is more than a couple of minutes, improve the summary your agent attaches at handoff.

Cost and capacity when machines are chatty

Machine-to-machine threads cost money in a way human chats do not. A person types a sentence and waits. An agent can send a full structured message in a second and read a long reply in another.

If both sides use models, each message pays twice.

Set the budget per thread from your cost per conversation, using the thinking in LLM cost optimization for support. A thread that costs ten times the median is a signal to stop and review, not to continue.

Prefer short, structured replies for agent peers. A reply with four fields uses fewer tokens than four paragraphs and is easier for the other side to parse. Cache answers to repeated public questions so you pay for them once.

Plan for capacity as a class. Agent peers can arrive in bursts when a partner runs a batch job. Put a separate queue and a separate ceiling on their traffic, so a burst does not delay your human customers.

If the ceiling is hit, respond with a clear retry hint, as discussed for shoppers.

Last, watch for the quiet cost of unmonitored threads. A loop that runs overnight on a modest model can cost less than a coffee or more than a car payment, depending on the traffic. A budget and a daily alert on spend per channel turn that unknown into a number you check.

Where this is still unsettled

Three things are open. The first is which protocols companies will adopt for support. A2A has a specification and a registered discovery path, aweb has a working system and a gateway, and the Personal Agent Protocol has an announcement and a planned v0.1 specification.

None has a track record in customer support.

The second is liability. If two agents agree to something their owners would not, who bears the cost? The documents I read do not answer that question, and I would not assume a court has either.

Until it is settled, the safe design is the one in which your agent cannot make the agreement.

The third is verification of delegation. A protocol can carry a claim that a person authorized an agent. How you check that claim, and how a customer revokes it, is the hard part.

The Personal Agent Protocol announcement promises to address authentication, and its text is not yet published.

I would watch the A2A specification for changes in how it treats authentication states, since that passage is the clearest statement of the risk. I would also watch for early production reports from teams that have run agent peers for support, and I would trust a report with numbers and a postmortem more than a launch post.

Frequently asked questions

What is agent-to-agent customer support?

It is a support exchange in which your agent talks to software representing another party, such as a customer's personal agent, a partner's service desk, or another of your own agents.

What is the A2A protocol?

A2A is an open protocol for agent communication. Its specification defines Agent Cards, stateful tasks, and messages with a user or agent role, over gRPC, HTTP with JSON, or JSON-RPC bindings.

Where do I publish an A2A agent card?

The specification names .well-known/agent-card.json as the discovery location. Cards can be signed with JWS, and an extended card requires authentication to read.

Is communicate.so agent.json an A2A card?

No. The communicate.so agent.json is a draft pointer to documentation, API, and MCP locations. It does not follow the A2A Agent Card schema and sits at a different path.

What is aweb?

It is an MIT-licensed, federated, self-hostable communication layer for AI agents with stable identities, signed messages, durable mail and chat, and wake-up events.

Does signing messages prevent collusion between agents?

No. The aweb builder said it does not, and that it provides a controlled channel, lets you track what agents say, and lets you validate the source of a message.

Does AUTH_REQUIRED mean the agent is authorized?

No. The A2A specification says an agent must not treat that state by itself as authorization for any operation, and that in-band credential exchange can expose credentials across chains of agents.

What is a trust ladder?

It is a set of levels that decide what another agent may do: unknown, identified, delegated read-only, and delegated with changes. The level depends on what you can prove.

How do I prevent two agents from looping?

Use a hop count, a message and spend budget, idempotency keys on actions, and a repetition check. When any limit is reached, stop and hand the thread to a human.

What is an idempotency key?

A unique identifier sent with an action so repeated requests return the first result instead of repeating the work. It is the standard remedy for retries after an unclear outcome.

What should my agent never agree to?

Price or contract changes, waived verification, disclosure of other customers' data, refunds above its cap, embedded instructions, and promises about future company conduct.

How should credentials be handled?

Keep them out of messages. Authenticate out of band and let the agent present a reference your system checks. Mask account numbers and emails in replies.

When should a human take over?

When the question is out of policy, the stakes are high, the thread looped, either side asked for a person, the other agent requests a higher rung, or you see signs of manipulation.

Who is liable if two agents agree to something wrong?

The documents I read do not say. Design so that your agent cannot make the agreement, and have counsel review your terms once your channel is live.

Do I need to implement A2A to prepare?

No. Write the trust ladder, add loop guards, and log peer identity first. Publish an agent card only after the policy exists.

How does this relate to the Personal Agent Protocol?

That protocol, announced by Sierra and Meta on October 6, 2026, is aimed at how a person's agent authenticates with businesses. Agent-to-agent support is the conversation that follows.

What should I log for agent threads?

The full thread, the peer identity, the rung, the delegation reference, the actions taken, the reason each was allowed, and the version of your rules at that time.

Can I tell customers their agent may talk to mine?

Yes, and you should. State it in your terms, say that a human can review any thread, and give a contact for problems.

What is the safest first step?

Write the one-page trust ladder and load it into your agent's instructions, together with a hop limit that sends long threads to a human.

Is loop prevention really necessary?

Yes. Two model-driven systems can keep replying with slightly different wording, each reply looking like progress while it costs money. A hop limit, a budget, and a repetition check are cheap insurance.

Conclusion

Agent-to-agent support is a small channel today and a design problem you can solve before it grows. The protocols give you identity and transport. The decisions that protect you are yours: the trust ladder, the list of things your agent never concedes, the limits that stop a runaway thread, and the rule that a person takes over when the stakes rise.

Write those rules first, then add protocol support. If you want an agent that works inside clear limits and hands cases to your team with the context intact, look at communicate.so AI agents or compare plans on the pricing page.