EPOCH arXiv:2610.06986 records 0.65 mean normalized score on AlgoTune versus 0.53 baseline
EPOCH closes the gap between evaluator feedback and claim validity in AI research agents through an explicit evidence-governed loop. Primary results are 0.65 AlgoTune and 0.57 Math14 scores with verified held-out replay. The architecture supplies the missing governance layer required for trustworthy deployment outside narrow benchmarks.
The architecture targets documented failure modes in AI research agents that optimize evaluator feedback without governing interpretation or reuse. Ten discovery problems show task-specific outputs including executable constructions, optimized algorithms, counterexamples, and proof-supported results. Official-test replay confirms held-out behavior while descriptive aggregates place EPOCH first on AgentHPO.
AlgoTune aggregate improves from 0.53 to 0.65; Math14 internal mean reaches 0.57. These deltas arise from explicit separation of finite certificates, benchmark improvements, and theorem-level claims rather than conflation under single scalar rewards. Evidence governance therefore functions as an admission filter before promotion of candidates.
Operational deployment in healthcare diagnostics or autonomous navigation requires identical replay and falsification steps to convert search outputs into auditable claims. Without typed memory and independent verification, benchmark gains remain non-transferable. The paper supplies the first explicit loop that enforces this distinction across mathematical and computational domains.
Next steps include scaling the same contract and replay mechanisms to larger program spaces and integrating external theorem provers for stronger certificates.
EPOCH: AgentHPO mean exceeds 0.62 under official replay within 12 months of public code release
Sources (2)
- [1]Primary Source(https://arxiv.org/abs/2610.06986)
- [2]Supporting Source(https://arxiv.org/abs/2402.13191)