How the math works
Every step in the chain looks fine on its own. Multiply them and the true end-to-end number is the one nobody watches.
The inputs (one refund-processing agent)
Eight steps that are each 85% reliable run clean end-to-end just 27% of the time. The kit computes true end-to-end reliability for every agent and flags the long chains before they quietly fail.
Expected end-to-end = per-step reliability raised to the number of steps. A monitoring aid, not a guarantee of agent behavior.