DeReAct Records 6.5-7.0 Point Pass@1 Gains for Qwen3-Coder-480B on GAIA and SWE-bench Verified
DeReAct separates reasoning, action authorization, and completion certification into distinct policies. Measurements on GAIA and SWE-bench Verified show the largest reliability gains accrue to weaker base models. External gating produces more evidence-complete trajectories once the gating policy itself is sufficient.
DeReAct decouples action proposal from validation by inserting an external Critic policy and a Context Manager that rebuilds environment-supported state and certifies completion. The architecture replaces the single LLM policy used in ReAct with two independent gating modules that authorize actions before execution and block unsupported termination claims. Trajectory logs show reduced error propagation when the gating models are themselves capable enough to detect the targeted failure modes.
On GAIA and SWE-bench Verified, absolute gains are largest for weaker backbone models and shrink as base capability rises. With Claude Opus 4.5, Pass@1 remains statistically flat while trajectory completeness and constraint adherence improve. Ablations confirm external gating adds value only when failure prevalence exceeds the gating policy's own error rate.
The pattern indicates that reliability interventions scale inversely with model strength and that completion control can be traded for stronger grounding rather than higher termination speed. Future deployments will require calibrated gating models matched to the failure distribution of each target domain.
DeReAct: Pass@1 gains above 5 points on at least two new agent benchmarks within 12 months when applied to sub-70B models.
Sources (2)
- [1]Primary Source(https://arxiv.org/abs/2610.02351)
- [2]Supporting Source(https://arxiv.org/abs/2210.03629)