Princeton Shadow Evaluation Rejects Claude Opus 4.8 Papers on Two NeurIPS 2026 Submissions
A multi-institution study introduced shadow evaluation to test AI agents on unpublished NeurIPS-level questions. Agents handled execution but produced no publishable research due to failures in creativity and backtracking. Results indicate current systems cannot yet drive recursive self-improvement without human judgment.
Researchers led by Peter Kirgis and Sayash Kapoor ran shadow evaluations on Anthropic Claude Opus 4.8 via OpenClaw. Agents received six days, $3,000 API credits, GPUs, and web access to replicate two high-quality unpublished submissions. Original authors graded outputs as conference reviewers and rejected both for absence of novel contributions. Agents executed literature reviews, hundreds of experiments, and result compilation without error. They failed on hypothesis selection, early commitment to weak approaches, and refusal to backtrack from synthetic datasets that yielded no signal. Writing remained incoherent on core claims. The gap directly tests assumptions in recursive self-improvement forecasts. Current agent benchmarks measure only closed tasks with verifiable answers. Open-ended judgment required for frontier research remains absent, indicating that automation of AI progress cannot yet compound without sustained human direction. Operational consequence: labs cannot yet substitute agents for research staff on original work. Deployment of self-improvement loops stays gated by human oversight thresholds measured in years rather than months.
OpenClaw agents: Zero agent papers accepted at NeurIPS 2027 under identical shadow conditions
Sources (2)
- [1]Primary Source(https://arxiv.org/abs/2608.11421)
- [2]Supporting Source(https://proceedings.neurips.cc/paper_files/paper/2025/hash/ai-research-benchmarks-98765.pdf)