How the gate works
This tool grades a live support bot across six fitness signals before you trust it in production. Each signal is scored 0–1; the verdict is the weakest one — then a gate that can only make it worse.
The 6 signals
A weakest-signal read of 1 would normally land at RESTRICT SCOPE. But if grounded answering AND verified resolution are both only partial at the same time, the gate forces PULL BACK regardless - because a bot that's improvising some answers while also miscounting merely-contained chats as resolved is unsafe in a way the weakest-signal read alone doesn't capture.
Why: The weakest-signal read alone would land at RESTRICT SCOPE - mean 83/100, with only grounding and verified resolution sitting at partial. But grounding and verified resolution are both partial at once - the bot improvises some answers off training data AND miscounts some merely-contained chats as resolved - so the vanity-deflection gate forces PULL BACK instead.
A clean sheet didn’t save it. The verdict is the weakest signal — and a top trigger is dispositive on its own — so it names the one thing to fix first. Your own marks, scored offline. No log access, and it grades the bot, not people.