In a production AI system, most failures are never judged at all. Evaluating an agent’s work costs money — a per-trace inference call to a large language model, or a human’s time, both too expensive to run on every trace. Production systems fall back to sampling a fraction and hoping the fraction is representative.
A study of 4,671 traces spanning five agent domains argues there is a cheaper witness already in the room. Failing agents, the researchers found, leave a detectable behavioral signature in standard observability telemetry: disproportionate effort relative to outcome. The agents that go wrong do not usually do less. They do more, to less effect, and the system logs record it.
The researchers formalized four signatures. The dominant one, present in all five domains tested, they call Compensatory Effort: failing runs expend 17 to 159 percent more effort on each domain’s primary signal (statistically significant, p<0.006, everywhere tested). Then there is Output Inflation — 38 percent larger patches on the SWE-rebench software-engineering benchmark, though the direction and reliability of the signal vary by domain.
Unfocused Orchestration shows up as token entropy — a measure of how scattered the agent’s output tokens are — running 5.98 to 12.3 percent higher in four of the five domains. And Correlation Decoupling means the effort-signal correlations simply shift structure between pass and fail traces (p<0.001, three domains), as if the failing agent’s work stops bearing on its result.
What that buys, in practice, is a simple threshold check over the signals: no judge, no label, no model call. Applied to the telemetry alone, it catches 52.9 to 61.5 percent of failures. That is not everything, but it is not sampled hope, either.
The judges themselves, it turns out, leave things on the floor. Across three human-labeled benchmarks, judge models missed 6.5 to 18.1 percent of confirmed failures on the two directly comparable ones. A trained telemetry classifier, with an AUROC of 0.821, recovered 61.5 to 75.4 percent of each judge’s misses. Put the two together and the miss rate for the weakest of the eight judges tested — GPT-4o-mini — falls from 18.1 percent to 4.4 percent, a 4.1-fold reduction. Stronger judges start with fewer misses, the researchers note, so the gain is correspondingly smaller.
None of this requires new instrumentation. The signatures live in the exhaust the systems already produce: the effort, the entropy, the inflated output of an agent working harder and getting less done. In 4,671 traces, the failures had announced themselves. The telemetry was just waiting for someone able to read it.

