LLM Agent Reliability Follows Geometric Decay, Collapsing Within 16 Steps on Production Tasks
Empirical evidence shows LLM agent performance degrades geometrically with horizon length due to per-step error rates that do not reach 1.0 at any current scale. Context limits worsen rather than improve outcomes. Horizon-aware evaluation and per-step reliability targets are required for production viability.
These results close the gap between short-horizon academic benchmarks and production workloads. Projected reliability falls from 0.42 at GAIA-scale horizons to 0.24 at 100-step enterprise flows, implying that aggregate pass-rate metrics systematically overstate deployability. The released trajectories enable direct replication and reliability budgeting.
Mittal et al.: By Q3 2027 at least two frontier labs will publish per-step reliability scores above 0.95 on 50-step agentic benchmarks or publicly retract claims of production-ready agents.
Sources (3)
- [1]Primary Source(https://arxiv.org/abs/2609.01660)
- [2]Supporting Source(https://arxiv.org/abs/2305.11730)
- [3]Supporting Source(https://arxiv.org/abs/2402.19450)