How the gate works
This tool grades one voice-agent call across six craft and trust signals - the shipped six-call sample. Each signal is scored 0–1; the verdict is the weakest one — then a gate that can only make it worse.
The 6 signals
The call-level verdict is the worse of two readings: craft (task resolution, policy, handling, tone) and trust (factual grounding, escalation correctness). A trust signal at 0 forces MISHANDLED and at 1 caps the call to SOFT-HANDLED, however clean the craft signals read - because a fabricated fact or a failed handoff makes a call harmful however polished it sounds.
Why: Task resolution, policy compliance, conversation handling, and tone are all a clean 2 - craft alone would read RESOLVED. But factual grounding is 0: the bot stated an unverified account balance. A trust signal at 0 overrides craft and forces the call to MISHANDLED, because a fabricated fact makes it harmful however polished the rest of it was.
A clean sheet didn’t save it. The verdict is the weakest signal — and a top trigger is dispositive on its own — so it names the one thing to fix first. Grades calls that already happened, from your own marks. Connects to no live agent, and it never scores people.