AI in Customer Support: Measure Resolution, Not Just Deflection
A deflected contact is a closed conversation. A resolved issue is a problem that stayed solved.

Research updated Oct 3, 2026
Key topics
A deflected contact is a closed conversation. A resolved issue is a problem that stayed solved.
The dashboard says the bot handled it. The customer reopened the same ticket three days later, or filed a chargeback, or posted the transcript. Those events land in different systems than the one reporting the win, which is why the win looks clean. Deflection counts exits. It cannot count outcomes, because it never looks past the end of the conversation.
That gap is the whole problem with how most teams measure AI in customer support workflows. The metric that gets reported is the one that is easiest to compute, not the one that tells you whether the customer's problem went away.
Deflection counts exits, not outcomes

Deflection is a routing metric. It answers one question: did this conversation end without a human? That is genuinely useful for capacity planning. If your goal is to keep agents out of password resets, deflection tells you whether the routing worked.
It stops being useful the moment the underlying issue has a downstream state change. A refund, a replacement, an account modification, a dispute — each of these can fail independently of the conversation that promised it. The chat ends. The state change does not happen. Deflection records a success anyway.
Grant the narrow case: for high-volume, low-consequence intents where the customer wanted information rather than intervention, deflection is close to a real outcome. "What are your support hours" has no downstream state. The answer either arrived or it didn't.
The boundary is the state change. If satisfying the customer's goal requires something in the world to be different afterward, then the conversation ending is not evidence that it happened.
So replace the unit. An issue is resolved when the customer's goal is met, the state change that satisfies it is confirmed, and the outcome holds without reopening inside a defined window. Three parts, all observable. If you cannot observe the downstream state, you cannot claim resolution — and the honest move is to say so rather than quietly substituting deflection and calling it success.
Sort issue types before you sort vendors
Most teams evaluate models before they evaluate their own ticket taxonomy. That order is backwards. The properties that determine whether AI can handle an intent belong to the intent, not to the model.
Classify along three axes:
- Consequence if wrong. What breaks when the system gets it wrong?
- Reversibility. Can the action be undone cheaply?
- Verification cost. How expensive is it to confirm the outcome was correct?
Low-consequence, reversible, cheap-to-verify intents are the natural first tier: status lookups, policy questions, address changes, password resets. High-consequence or irreversible intents — refunds, account termination, disputes, identity verification, anything that moves money — need deterministic checks and explicit human review gates. The model can decide what to say. It should not be the thing that decides whether funds move.
Ambiguity deserves its own category. Multi-request messages, unclear intent, and emotionally loaded contacts are where autonomous handling degrades fastest. A single targeted clarifying question usually beats a confident answer built on a guess.
There is a research signal here, and it is worth stating with its boundaries intact. A 2025 study of customer-service agents modeled the interaction as a partially observable decision process and compared agent designs on a banking-support benchmark. Its reported results show that task success varies substantially with how the agent is configured — including how much reasoning effort is applied and how knowledge is retrieved — and that the gap between an idealized "gold" tool sequence and retrieved-knowledge performance is large. Read that narrowly: on that benchmark, under those task conditions, retrieval and procedure design moved outcomes more than the choice of base model. It is not a general ranking, and it is not proof that any of this is production-ready. A benchmark result is a result under agreed test conditions. Your ticket queue is not those conditions.
The decision rule that survives contact: if you cannot write the verification step for an intent, it is not ready for autonomous handling.
Define resolution in a way you can actually measure
"Resolution" is a slogan until you write it down. A workable definition has three parts: the customer's stated goal, the state change that satisfies it, and a hold window during which the issue must not reopen.
Then compute it honestly:
Resolution rate = resolved issues ÷ issues that entered the workflow.
Count issues, not conversations. One customer sending three messages about one broken order is one issue. Inflating the denominator with message threads is how resolution rates get to numbers nobody believes.
Pair resolution with a quality signal, because a fast wrong answer can look resolved. Policy adherence, factual correctness against your knowledge base, and customer-reported satisfaction are separate axes. A system can score well on one and fail another.
Measure cost per resolved issue, not cost per token or cost per contact. A cheaper model that needs retries, extra tool calls, or human correction can cost more per solved problem than a stronger model that finishes on the first attempt. Token counts are diagnostics. They do not tell you whether the customer's problem was solved.
Track escalation as a first-class outcome rather than a failure. Escalation rate, escalation reason, and post-escalation resolution tell you where the boundary of autonomous handling actually sits — which is the thing you are trying to find.
And watch the gaming risk. Any metric tied to a bonus will be optimized. If a system or an agent can close a ticket without satisfying the goal, your resolution rate becomes fiction with a decimal point.
Design escalation as a product surface, not a fallback
Escalation quality determines whether AI support preserves customer outcomes or quietly damages them. It needs three things: a trigger, a handoff contract, and a recovery path. Miss one and escalation becomes a restart.
Triggers should be explicit and testable: confidence thresholds, intent classes on the high-consequence list, repeated failed attempts, detected frustration, and any direct request for a human.
The handoff contract is the part teams skip. The human agent needs the customer's goal, what was already attempted, the tool calls made, and the current state of the case. Without it, the customer repeats themselves, the AI's work is wasted, and the interaction costs more than if a human had answered first.
Preserve deterministic authorization and refund checks outside the model. Instrument the escalation path itself — time to human, context completeness, whether the human had to redo work — because that is where the hidden cost lives. A bounded number of clarification attempts prevents loops; an unbounded one turns a confused customer into a stuck one.
Quality checks that catch the quiet failures
Sample and review contacts that were resolved without a human. The deflected-and-forgotten case is exactly the one nobody audits, because the dashboard already called it a win.
Build a small labeled evaluation set from real conversations, with human annotators marking the correct action for each. That is the ground truth your offline metrics get measured against. Twenty reviewed conversations teach more than a dashboard of aggregate rates.
Separate offline evaluation from live behavior. A model that scores well on a static set can still fail on retrieval, tool ordering, or missing context in production. Track a short list of failure modes explicitly: wrong policy cited, action taken without required verification, hallucinated account state, premature closure, clarification loops.
Then feed findings back into knowledge representation. Support knowledge written for human agents — dense tables, nested conditions, rich text — is often hard for a model to use correctly. Research on LLM-friendly knowledge formats suggests that reformatting policy and workflow documents is frequently higher-leverage than swapping models. That matches what I have seen building retrieval systems: the model is rarely the bottleneck. The corpus is.
What to instrument in the first ninety days
Start narrow. One intent tier, one resolution definition, one hold window. Narrow scope is what makes the measurement trustworthy, and it is the only way to know whether a result means anything.
Instrument four numbers from day one: resolution rate, escalation rate with reasons, cost per resolved issue, and reopen rate within the hold window.
Write the escalation contract before you write the prompts. The handoff is the part that determines whether customers stay.
Build the labeled evaluation set early, even if it is small. The skills worth developing on your team are intent taxonomy design, evaluation set construction, retrieval and knowledge-base restructuring, and cost-per-outcome modeling. Those are the reusable assets. The model choice is the replaceable part.
One watchpoint worth naming precisely. In November 2025, Microsoft announced general availability of an autonomous case-resolution capability inside its Dynamics 365 customer service product, describing an agent that reads an inbound case, checks warranty status through a custom agent, and sends a replacement confirmation without human involvement — escalating when a message is ambiguous or unanswered. That is one vendor's shipped capability, not a market-wide arrival. My read is that as features like this move from pilot to general availability, the pressure to report deflection-style numbers will grow, because those numbers are flattering and easy to produce. Treat that as an interpretation, not an established trend. The teams that hold a written resolution definition will be the ones who can tell whether the deployment actually worked.
If you cannot state the customer's goal, the state change that satisfies it, and the window in which it must hold, you are not measuring resolution. You are measuring silence. Pick one intent tier. Write the definition and the contract. Instrument cost per resolved issue before you scale anything.
References
Want practical AI trend signal in one place?
Use the AI Trend Brief Starter Pack to turn fast-moving AI news into a clearer builder-focused reading path.
AITrendFast Monthly — September 2026
A focused September 2026 AITrendFast briefing covering open-weight adaptation, multimodal generation, agent interoperability, permissions, memory, and ecosystem security.
- 86-page Illustrated PDF edition
- 6 curated reports
- Enhanced PDF edition with bundle-only briefing guidance
- Offline-friendly format for focused review
- Source report links for future online updates
Coming soon


