The Compound Failure Problem
Why "almost reliable" isn't reliable at all
0.9520 = 0.358 — long autonomy is a coin flip dressed as a plan.
35.8%
chance of full success after 20 steps at 95.0% per-step accuracy
Expected failures per 100 runs: 64
At 95% accuracy, a 20-step task succeeds only 35.8% of the time. One in three. Would you board that flight?
At 99% accuracy, a 100-step task succeeds only 36.6% of the time. Even "near-perfect" fails at scale.
At 99.9%, you need fewer than 70 steps to drop below 93%. No pipeline is immune.
At 90% accuracy, 10 steps gives you 34.9%. Most real-world agent workflows exceed 10 steps.
This is why multi-step AI verification catches errors humans miss.
See how →