Voice-Agent Conversation QA & Escalation Auditor
Your dashboard says the call was resolved — was it? Grade a batch of voice-agent call transcripts and catch the calls a human would call mishandled: the hallucinated balance, the angry caller who never got a handoff, the polite call that ended with the goal unmet. Mark six 0/1/2 signals per call — a craft MIN over task resolution, policy, handling, and tone sets the base, then a graduated trust gate over factual grounding and escalation correctness worsens it (a trust signal at 1 caps to SOFT-HANDLED, at 0 forces MISHANDLED). Per call RESOLVED / SOFT-HANDLED / MISHANDLED; the batch rolls up CLEAN RUN / SPOT-CHECK / MISHANDLING FOUND with a mishandle rate and the one call to review first. The post-launch QA layer beside the Voice Agent Deployment Kit (build) and the Go-Live Readiness Gate (pre-launch). Runnable Python engine + workbook + four Claude Skills + 729-combo verifier + 2 playbooks. Deterministic, offline; grades the conversation, never people.






