89% of organizations have agent observability, while 52.4% run offline evaluations on a test set (LangChain, 2026).
Degrading one layer of an agent moved the end-to-end pass rate by only 1.7 to 5.9 points, while the matching test slice fell by 25 to 91 points.
A parameter-level error reached a wrong final answer with a human-calibrated probability of about 0.62.