Qwen3.6-27B Agents Ignore Stated Wall-Clock Budgets on MLE-Bench Lite Without Harness Timing Feedback
The study isolates time adherence from productive time use in budget-constrained LLM agents. Harness feedback and GRPO training fix adherence but leave task performance flat because agents repeat actions instead of reallocating effort. The central remaining gap is the absence of learned mappings from remaining budget to higher-value actions.
The arXiv paper evaluates small LLM agents under explicit runtime constraints on two agentic benchmarks where extra compute can improve outcomes. Prompt-only budget statements produce no measurable adherence. Agents lack time awareness, cannot forecast action durations, and hold no learned policy linking remaining budget to strategy shifts. Harness interventions that expose elapsed time and enforce deadlines raise adherence rates for Qwen3.6-27B with no drop in task score. GRPO reinforcement learning on budget-aware rewards drives near-perfect adherence on Zork I and generalizes to unseen budgets, yet leaves MLE-Bench performance unchanged from the untrained baseline. Post-adherence analysis shows agents still treat surplus time as license for repetition rather than deeper search or alternative paths. GRPO policies trained across multiple budgets collapse to the shortest-budget strategy, exposing a separation between deadline compliance and productive allocation. Future harnesses will require explicit action-cost models and reward terms that penalize redundant steps once a minimum viable solution is reached.
Qwen3-4B: GRPO policies reach 85% adherence on held-out Zork I budgets within four training epochs after multi-budget exposure.
Sources (3)
- [1]Primary Source(https://arxiv.org/abs/2610.10833)
- [2]Supporting Source(https://arxiv.org/abs/2402.19450)
- [3]Supporting Source(https://arxiv.org/abs/2305.15717)