When not to use AI customer support: six situations for a human
Communicate.so
Six situations where AI customer support is the wrong tool and a human is required, argued with the mechanism, not just the rule.
TL;DR: AI support has a real, useful range, and it is smaller than most vendor pitches suggest. This guide argues six specific situations where the correct tool is a human, not an AI agent: active harm or crisis situations, active fraud or account compromise, decisions with irreversible financial consequences past a threshold, legal or regulatory disputes, a customer who has already asked for a human and been refused, and any case where the AI has failed twice on the same issue. Each case gets the mechanism, not just the rule, because the reason matters more than the list. An AI agent grounded in real content is a strong tool for the repetitive majority of support volume, and building it into a company also means being honest about where it should not be making the call.
Most writing about AI customer support argues for adoption. This one argues the opposite case on purpose: when not to use ai customer service is a real question with a real answer, and a vendor that will not name its own limits is not being honest with you. Six situations follow where a human is the correct default, not a fallback.
This is written from inside a company that builds AI support tooling, which is exactly why it needs to be specific rather than defensive. An agent that resolves the repetitive majority of questions, the kind covered in customer support automation, earns its place by being honest about the remainder. What follows is that remainder, argued properly rather than hedged.
Each of the six sections below follows the same structure on purpose: what the situation looks like, why an AI agent is the wrong tool for it specifically, and what a working alternative looks like in practice. The goal is not a list of warnings to post on a wall. It is a working definition of scope that a support team can actually build into how their agent is configured, reviewed with the same seriousness a company would apply to any other operational risk.
Active harm or crisis situations
A customer expressing self-harm risk, describing an emergency, or reporting an active safety threat needs a human immediately, with no automated triage step in between. This is not a support ticket in the normal sense. It is a moment where the response has to carry judgment, tone, and possibly a decision to involve emergency services, none of which a support agent should be making alone.
The mechanism here is that AI support agents are built to answer questions from documented content, and a crisis is not a documented-content problem. There is no knowledge base article that tells an agent the right thing to say to someone in distress, and an agent that tries anyway is pattern-matching its way through a situation where the cost of a wrong pattern is severe.
The correct design is a hard trigger, not a soft one. Certain language should route directly to a human without the agent attempting a first response at all, the same immediate-escalation logic covered in support escalation workflow. This is the one category on this list where speed to a human matters more than getting a smooth handoff message right.
Support teams that build this trigger tend to underestimate how narrow the detection window needs to be. A crisis rarely announces itself with an obvious keyword; it shows up in a customer's tone shift mid-conversation, a sudden change in what they are asking for, or language that a generic support script was never trained to recognize. The safest posture treats a wide range of ambiguous signals as worth a human look rather than waiting for certainty, because the cost of one missed case outweighs the cost of many unnecessary handoffs.
Active fraud or account compromise
A customer reporting that their account has been compromised, that a payment method was used without authorization, or that they suspect fraud needs a human who can verify identity and take real security action. An AI agent can gather initial information, but the actual decision to lock an account, reverse a charge, or investigate a breach carries authority a support agent should not hold unsupervised.
The risk is specific: a compromised account is, by definition, a situation where you cannot fully trust the identity of the person you are talking to. An automated agent following a documented password-reset flow can become the exact tool an attacker uses to take over an account, if it applies its normal helpful pattern without the extra verification a human would instinctively add.
The mechanism is that fraud response requires judgment calls that fall outside any script: does this identity claim look right, does this pattern match known attack behavior, does this need to be escalated to a security team rather than resolved in the chat. An agent scoped correctly, per AI agent guardrails, should treat any fraud or compromise signal as an immediate handoff trigger rather than a question to answer.
The documented chatbot incidents covered in AI chatbot failures share a common thread with fraud risk: each one involved an agent acting confidently inside a situation it was not actually equipped to judge, whether that was a jailbroken support bot or a device policy nobody had written down. Fraud is the same failure mode with higher stakes, because the agent is not just wrong about a policy, it is potentially the mechanism an attacker uses to complete the takeover. That is reason enough to treat any compromise signal as a hard stop rather than a prompt to answer more carefully.
Irreversible financial decisions past a threshold
Communicate.soA refund, contract cancellation, or credit decision above a meaningful dollar threshold should get a human check before it executes, even if the AI agent can technically process it. The Air Canada bereavement fare case is the clearest illustration of this: a chatbot stated a financial promise, the customer relied on it, and the company was legally bound to honor what its own tool said.
The mechanism is asymmetry. A correct automated refund saves a few minutes of human time. A wrong one, especially one the company then has to honor because a customer reasonably relied on it, costs real money and sets a precedent other customers can point to, the exact dynamic covered in the Air Canada chatbot ruling.
Below a threshold you are comfortable eating as a cost of doing business, full automation is reasonable. Above it, a human should confirm before the promise becomes binding.
This is not an argument against AI touching money at all. It is an argument for a dollar-value gate: let the agent handle small, low-stakes refunds end to end, and route anything above a set threshold, or anything involving a policy exception, to a person who can weigh the specific case rather than apply a general pattern.
Setting that threshold is a business decision, not a technical one, and it deserves the same rigor a company would apply to any spending authority. A reasonable starting point is looking at the last twelve months of refund requests, sorting by dollar amount, and finding the point past which a wrong automated decision would sting rather than round to a rounding error. That number becomes the gate, and it should move only when the team has evidence the agent is reliable at the amounts already inside it, not on a hunch that things have been going fine.
Legal or regulatory disputes
A customer threatening legal action, referencing a regulatory complaint, or asking a question that touches on compliance obligations needs a human, and usually one connected to legal or compliance, not a frontline agent of either kind.
The mechanism is that a chatbot's statement is a statement your company can be held to, established plainly by the Air Canada ruling's reasoning that a company is responsible for information from a chatbot the same way it is responsible for a static page. In an active dispute, every sentence the agent generates is a potential exhibit. That is not a place for an agent optimizing for a fast, helpful-sounding answer, per the ruling explainer.
The practical trigger is any mention of legal counsel, a regulator, a formal complaint process, or specific legal terms like breach of contract. None of these need to be perfectly detected by keyword matching to be useful; a rough trigger that occasionally over-escalates to a human is a far smaller cost than an AI agent negotiating a legal dispute on the company's behalf.
There is a record-keeping angle worth naming separately. Once a dispute is active, the full conversation history, including what the AI agent said before a human took over, becomes part of the record a company may need to produce. That is a strong argument for a support platform where AI and human replies live in one continuous thread rather than a system where the automated portion disappears into a separate log a legal team has to go hunting for later.
A customer who has already asked for a human
If a customer has explicitly asked to speak with a person and been told no or redirected back to the bot, continuing with AI is the wrong call regardless of whether the AI could technically resolve the issue. This is less about the AI's capability and more about respecting an explicit preference. Twig's research on AI support complaints lists no escalation path as one of the most common frustrations customers report, and refusing a direct request for a human is the sharpest version of that failure.
The mechanism is trust, not competence. A customer who asks for a human and gets deflected back to a bot learns that the request itself does not work, which damages trust in the support channel independent of whether any individual answer was correct. That damage compounds every time it happens, and it costs far more in reputation than the marginal automation it saves.
The fix is a visible, working escalation path, not a hidden one buried behind three more automated prompts. A clean AI to human handoff that a customer can trigger with a plain request, and that actually connects them to a person promptly, is the baseline, not a nice-to-have.
Worth tracking as its own number: how often customers ask for a human and how often that request is honored on the first ask. A team that only measures overall deflection rate can miss a rising refusal pattern entirely, because a conversation where the bot deflects an explicit human request but eventually resolves the issue still counts as a win in the aggregate metric, even though it is exactly the experience this section argues against.
Repeated failure on the same issue
If an AI agent has already failed to resolve a specific issue twice in the same conversation or across follow-up contacts, the third attempt should be a human, not another automated pass. Two failed attempts is a signal that the issue falls outside what the agent is grounded to handle, whether that is a gap in the knowledge base, an edge case the documentation does not cover, or a genuinely unusual situation.
The mechanism is diminishing returns combined with rising frustration. A third automated attempt on an issue the agent has already failed twice is unlikely to succeed by the same mechanism that failed before, and each additional failed attempt makes the eventual human handoff start from a more frustrated customer than the last one would have.
A simple failure counter tied to automatic escalation solves this cleanly: after two unresolved attempts on the same underlying issue, route to a human by default rather than waiting for the customer to ask. This connects directly to the confidence and escalation logic in reducing AI hallucinations in support, where the agent's own uncertainty, not just explicit customer requests, should trigger the handoff.
A repeated failure is also useful data outside the single conversation it happened in. If the same issue keeps triggering a second and third attempt across many customers, that is a signal the underlying content the agent is grounded on has a real gap, not just a signal that one conversation needs a human. The handoff solves the immediate case; the pattern across handoffs is what tells a team what to fix at the source so the same failure stops recurring.
The cost of getting the line wrong in either direction
Draw the line too conservatively and you lose most of the value AI support offers; draw it too loosely and you inherit the specific risks documented in real incidents. Both mistakes are common, and they come from different instincts. A team that has read every AI horror story tends to over-restrict, routing far more to humans than the risk actually justifies, which quietly defeats the point of automating the repetitive majority in the first place.
A team optimizing purely for deflection rate makes the opposite mistake, treating every automated resolution as a win regardless of what was being resolved. That is how a company ends up with an agent confidently stating a bereavement fare policy, a device-limit rule, or a downgrade term that was never actually documented anywhere, the pattern behind every incident covered in AI chatbot failures.
The corrective in both directions is the same: draw the line from the actual cost of being wrong, not from a general feeling about AI risk or a general enthusiasm for automation. A wrong answer to "what are your support hours" costs almost nothing. A wrong answer that promises a refund, alleges fraud, or dismisses a safety concern costs a great deal, and the six categories in this guide are exactly the situations where that cost is highest.
AI agent guardrails covers how to encode that distinction into what the agent is actually allowed to do.
Where the line actually sits
| Situation | AI agent handles it | Human required |
|---|---|---|
| Documented FAQ, order status, account lookup | ✓ | ✗ |
| Small refund under a set dollar threshold | ✓ | ✗ |
| Self-harm, emergency, or safety threat language | ✗ | ✓ |
| Suspected fraud or account compromise | ✗ | ✓ |
| Large refund, contract change, or policy exception | ✗ | ✓ |
| Legal threat or regulatory complaint | ✗ | ✓ |
| Explicit request for a human | ✗ | ✓ |
| Same issue unresolved after two attempts | ✗ | ✓ |
The pattern across all six situations is the same: they are the cases where a wrong or clumsy answer carries a cost that a fast, cheap resolution cannot offset. Everything left of that line is where AI genuinely helps, resolving the repetitive volume that used to consume a support team's entire day.
Building an AI agent honestly means drawing this line before launch, not discovering it after an incident. Communicate's agent is scoped to answer from data you connect and hands off to a human on the same shared inbox when a conversation crosses into any of the six categories above, because the goal is covering the repetitive majority well, not pretending the remainder does not exist.
Frequently asked questions
Does this mean AI customer support is not worth using?
No. The argument is about scope, not value. An AI agent grounded in real content resolves the repetitive majority of support volume effectively, the case made in customer support automation, and the six situations here are deliberately the minority of cases where a human is the better tool, not evidence against automation generally.
What counts as a crisis or safety situation for escalation purposes?
Language indicating self-harm risk, an active emergency, or a threat to someone's safety should trigger an immediate human handoff without the agent attempting a normal first response. The specific keyword and phrase triggers a team builds should be reviewed with input from people experienced in crisis response, not treated as a standard support configuration decision.
It is worth building this list with more entries than feels necessary at first, since the cost of an unneeded human handoff is small compared to the cost of a missed one, and reviewing the list periodically against real conversation transcripts catches phrasing the original list did not anticipate.
Why can't an AI agent just verify identity better to handle fraud cases?
Identity verification helps, but the underlying issue is judgment under uncertainty: deciding whether a specific pattern looks like a known attack, weighing context a script cannot fully capture, and having the authority to take real security action. Better verification reduces risk without eliminating the need for a human decision-maker, which is why fraud stays on the human side of the line described in AI agent guardrails.
What dollar threshold should trigger human review for refunds?
There is no universal number, because it depends on your margins, your typical order size, and how much risk you are comfortable automating. The useful exercise is picking a number you would be comfortable eating as a cost of full automation, then routing anything above it to a human, rather than leaving the threshold undefined and discovering it after a costly mistake.
A business with a five dollar average order size and one with a five hundred dollar average order size should not use the same number, and a threshold copied from a blog post without adjusting for that difference is worse than no threshold at all.
How does the Air Canada ruling relate to refund automation?
The ruling established that a company is bound by what its chatbot promises, the same way it would be bound by a written policy. That is the direct reason large or exceptional refund and credit decisions belong with a human: an AI agent making that promise unsupervised creates a binding commitment nobody reviewed. Full detail is in the Air Canada chatbot ruling explainer.
Should every legal mention trigger an automatic human handoff?
A rough trigger that occasionally over-escalates on a legal-sounding phrase is a far smaller problem than an AI agent continuing to engage in an active legal or regulatory dispute. It is reasonable to accept some false positives here in exchange for near-zero tolerance for false negatives, since the downside of missing a real one is much larger than the downside of a human reviewing an ordinary complaint.
What should happen when a customer explicitly asks for a human?
The request should be honored, with a clear and prompt handoff to a person, not redirected back to the bot or buried behind more automated prompts. Refusing an explicit request for a human is one of the sharpest ways to damage trust in a support channel, and it shows up directly in complaint research like Twig's list of common AI support frustrations.
How many failed attempts should trigger automatic escalation?
Two unresolved attempts on the same underlying issue is a reasonable default: it gives the agent one retry in case the first miss was a fluke, but stops short of letting a customer cycle through repeated failures. This ties directly into the confidence and escalation approach covered in reducing AI hallucinations in support.
Can an AI agent detect a crisis situation reliably?
Keyword and pattern-based detection can catch a meaningful share of clear cases, but it will miss some and occasionally over-trigger on others. The right posture is to accept a wide net that sometimes escalates unnecessarily rather than a narrow one that risks missing a real crisis, because the cost of those two error types is not remotely symmetric.
No detection system replaces a trained human making the actual judgment call once the conversation reaches them. The agent's job is narrow: recognize enough of a signal to get a person into the conversation quickly, not to make any determination about what is actually happening.
What is the difference between escalation and full human handoff?
Escalation can mean flagging a conversation for a human to review after the fact, while a full handoff means a person takes over the live conversation immediately. The six situations in this guide generally call for a full, immediate handoff, not a delayed review, because the cost of an AI agent continuing to engage in the meantime is the whole problem.
Should a company disclose that certain issues always route to a human?
Being upfront that some categories, fraud, safety, legal disputes, always reach a person builds trust rather than undermining confidence in the AI agent for everything else. It pairs naturally with general AI disclosure obligations discussed in the context of the EU AI Act, where transparency about what the customer is interacting with is the core duty.
Does a small business need all six escalation categories from day one?
Most of them, yes, because the risk they address does not scale down with company size. A small business with one customer suffering real financial or safety harm faces the same underlying problem a large company does, just at a smaller volume. The categories worth deferring are the ones tied to complex internal processes, like nuanced legal triage, that a small team may handle more informally at first.
What changes with size is not whether the category matters but how formal the process behind it needs to be. A two-person support team can handle fraud escalation with a shared inbox and a clear rule; a two hundred person team needs a documented workflow and a dedicated queue for the same underlying category.
How do I decide what counts as a fraud signal for my business?
Look at the specific fraud patterns your business has actually seen, not a generic list, since fraud signals vary a great deal by industry. Common starting points include a customer disputing a charge they say they did not authorize, a request to change account details paired with unusual urgency, or login activity the customer says was not theirs, each of which should trigger a human review rather than an automated resolution.
A support team's own historical tickets are the best source for this list, better than any generic industry template. Pulling the last year of chargebacks and account disputes and reading what triggered each one usually surfaces two or three patterns specific to that business that a generic list would have missed entirely.
Is it safe to let an AI agent handle account password resets?
Standard password resets through a verified channel, like a link to an email or phone number already on file, are generally safe to automate because the verification step does the real work. The risk rises sharply when a customer claims they no longer have access to that verified channel, which is exactly the moment a human should take over, consistent with the fraud and compromise reasoning in AI agent guardrails.
What happens if an AI agent is used in one of these six situations anyway?
The risk ranges from a frustrated customer to a real financial, legal, or safety consequence, depending on the category. The Air Canada case shows the financial and legal end of that range concretely: a company held liable for a promise its chatbot made without the review a human would have applied to the same commitment.
The reputational cost tends to outlast the immediate incident. A customer who experiences an AI agent mishandling a fraud report or dismissing a safety concern is unlikely to describe the failure narrowly; they describe it as the company failing them, which is a harder impression to repair than a single wrong answer would suggest.
How does Communicate handle these escalation cases?
Communicate's AI agent is scoped to answer from the data you connect and is designed to hand off to a human on the same shared inbox surface, with full conversation history intact, when a conversation falls outside that scope. The specific triggers, dollar thresholds, fraud signals, explicit human requests, are configured per business rather than applied as one fixed rule, because the right line depends on the specific risks a given company carries.
Does deflection rate matter if a team follows these six exceptions?
It still matters as a measure of how much repetitive volume the agent is absorbing, but it stops being the only number worth watching. A high deflection rate achieved by letting the agent answer questions it should have escalated is the opposite of the goal here; the useful version of the metric is deflection on the categories this guide leaves to AI, measured separately from the six it does not, a distinction covered in support ticket deflection rate.
How do these six situations interact with first response time targets?
A fast automated response to a situation that actually needed a human is not a good outcome, even though it looks good on a first response time report. The honest fix is measuring first response time separately for AI-handled and human-escalated conversations, so a team is not rewarded for speed on cases that needed judgment instead.
Can these six categories change over time as a company grows?
Yes, and they usually should. A company that starts with a low, conservative refund threshold because it cannot yet absorb the cost of an automated mistake can raise that threshold once it has data showing the agent handles smaller refunds reliably. The categories themselves, crisis, fraud, large financial decisions, legal disputes, explicit human requests, repeated failure, stay constant; only the specific thresholds inside them should move with evidence.
Where should I read next on building the human side of this well?
See AI to human handoff for how to make the handoff itself smooth rather than a cold restart, and support escalation workflow for how to structure the triggers and routing logic behind these six categories.
Between the two, the handoff piece matters more than it looks on paper. A well-designed escalation trigger that dumps a customer into a slow, context-free queue undoes most of the trust it was meant to protect, so the routing logic and the receiving experience need to be built and reviewed together, not treated as separate projects.