Skip to content
Founder Workflows
Browse
Browse by stage

Measure true resolution, not just deflection, for an AI support agent with an escalation policy

A small open-source evaluation shows a naive support agent deflecting every ticket while resolving about half, and a policy-gated agent that deflects fewer but resolves what it handles.

DocumentedSupportSmall SaaSScript

Evidence: The source shows its work: steps, screenshots, code or data you can inspect.

The business problem

Support dashboards often report deflection as success, hiding tickets that left the queue without actually being solved.

What was tried

A LangGraph pipeline classifies intent, retrieves knowledge-base context, decides an action, checks escalation triggers (low confidence, an explicit request for a person, negative sentiment, regulated topics) and either resolves or escalates with a handoff summary. A rule-based judge compares what the agent did with ground truth, and the gap between deflection rate and true resolution rate is the headline metric. Backends are simulated with local files and a mock model runs by default, with an optional Claude key.

What was reported (mixed)

On its own simulated data the naive agent showed 100% deflection but 53% true resolution (a 47-point gap), while the policy-gated agent deflected 53% and resolved 100% of those, a gap of zero.

Limitations

The results come from a simulated environment with a mock model, and the ticket count, test composition and variance are not stated. It handles only single-message tickets, uses keyword search, has an uncalibrated optional language-model judge, and does not report latency or cost. It is a one-commit portfolio project under the MIT licence, and external statistics in the README were not verified.

What you need

Python and the listed packages. An Anthropic API key is optional, and tests run without one.

Sources

Source published: unknown. Last reviewed here: October 11, 2026. Spot a mistake? Tell us.

Related workflows