
Study finds AI agents fail open-ended research tasks required for recursive self-improvement
Evaluation data shows AI agents lack the judgment for open-ended research. This delays verifiable recursive self-improvement and increases reliance on human-directed iteration. Regulators gain measurable criteria for monitoring capability jumps.
The MIT Technology Review report cites an unnamed study testing AI agents on free-form investigations without predefined answers. Agents failed to generate novel hypotheses or iterate on ambiguous results. This directly challenges industry claims that recursive self-improvement will arrive via autonomous model improvement loops.
Prior agent benchmarks from METR and the 2024 GAIA evaluation recorded success rates below 30 percent on multi-step scientific tasks. The new results extend those findings by removing guardrails and answer keys. Performance collapsed when agents encountered contradictory data or required creative reframing of problem boundaries.
Operational impact is immediate for labs scaling agentic workflows. Companies relying on narrow task automation can continue incremental gains. Full recursive loops still require human oversight for research direction. Regulators tracking capability thresholds now have concrete evidence separating narrow scaling from autonomous research acceleration.
Next evaluations will test whether chaining narrower improvements can substitute for open-ended capability. No public timeline exists for agents reaching the required threshold.
METR: No agent scores above 50 percent on open-ended AI research benchmarks by Q4 2027
Sources (3)
- [1]Primary Source(https://www.technologyreview.com/2026/08/19/1140195/the-download-ai-recursive-self-improvement-problem-heatwave-causes/)
- [2]Supporting Source(https://arxiv.org/abs/2403.04178)
- [3]Supporting Source(https://metr.org/blog/2024-agent-evals/)