Measure true resolution, not just deflection, for an AI support agent with an escalation policy
A small open-source evaluation shows a naive support agent deflecting every ticket while resolving about half, and a policy-gated agent that deflects fewer but resolves what it handles.
Evidence: The source shows its work: steps, screenshots, code or data you can inspect.
The business problem
Support dashboards often report deflection as success, hiding tickets that left the queue without actually being solved.
What was tried
A LangGraph pipeline classifies intent, retrieves knowledge-base context, decides an action, checks escalation triggers (low confidence, an explicit request for a person, negative sentiment, regulated topics) and either resolves or escalates with a handoff summary. A rule-based judge compares what the agent did with ground truth, and the gap between deflection rate and true resolution rate is the headline metric. Backends are simulated with local files and a mock model runs by default, with an optional Claude key.
What was reported (mixed)
On its own simulated data the naive agent showed 100% deflection but 53% true resolution (a 47-point gap), while the policy-gated agent deflected 53% and resolved 100% of those, a gap of zero.
Limitations
The results come from a simulated environment with a mock model, and the ticket count, test composition and variance are not stated. It handles only single-message tickets, uses keyword search, has an uncalibrated optional language-model judge, and does not report latency or cost. It is a one-commit portfolio project under the MIT licence, and external statistics in the README were not verified.
What you need
Python and the listed packages. An Anthropic API key is optional, and tests run without one.
Sources
- GitHub (aniruddhabasu1985-tech/trueresolve-AI) ↗ Code repository, publication date unknown
Source published: unknown. Last reviewed here: October 11, 2026. Spot a mistake? Tell us.
Related workflows
Roll out OpenClaw for inbox triage, meeting notes and CRM updates in phases with approvals and isolation
An infrastructure vendor's guide sets out business uses for OpenClaw and a phased rollout: pilot a low-risk workflow, add guardrails and approvals, then scale once value and safety are shown.
Draft help-centre articles automatically from resolved tickets and keep a human in charge of publishing
A merged helpdesk change that, when a ticket is resolved, asks an AI whether other customers are likely to ask the same thing and writes a draft article for managers to review if none exists.
Draft help-centre-cited support replies and propose refunds that a person must approve
An open-source support workbench classifies tickets into three tiers, drafts replies that cite help-centre articles, reads billing data in read-only test mode, and only proposes refunds that need a named approver.
Deflect support tickets only when a calibrated confidence score is high, and queue the middle band for human review
An open-source ticket deflection engine escalates risky tickets directly, answers procedural ones by rule, and uses retrieval over documentation with three confidence bands for auto-reply, human review and escalation.