TEAM-Design Allocates Replays Across 6 Clinical Settings Where No Human-AI Workflow Beats Both Baselines
TEAM-Design supplies the first decision-targeted sampling rule for human-agent evaluation. It reallocates limited replay resources to the harder counterfactual and maintains statistical guarantees. Reanalyses show most published clinical teams do not justify deployment.
The paper introduces TEAM-Design, a sampling rule that computes two replay probabilities per task. Probability rises when the missing baseline outcome is difficult to predict from observed data and when the comparison is closer to failing the superiority threshold. It falls when replay cost is high. The rule is proven to solve the budgeted design problem exactly and to maintain error-rate control even under random draw from the computed probabilities.
Reanalysis of six chest X-ray and radiology studies showed every human-AI workflow failed to beat both human-alone and model-alone performance. A coding benchmark yielded one workflow that succeeded. Synthetic and semi-synthetic experiments confirmed TEAM-Design outperforms variance-based allocation when one comparison is markedly harder than the other, yet underperforms when difficulties are balanced.
Existing agent benchmarks and Bayesian methods optimize parameter estimation rather than the binary keep-or-replace decision. By forcing explicit measurement of both counterfactuals, TEAM-Design surfaces cases where collaboration increases error rates or masks accountability, directly addressing ethical concerns over diffused responsibility in clinical and engineering teams.
Next deployments will require pre-specified replay budgets and documented probability vectors before any human-AI workflow receives production approval. Regulators can treat the resulting superiority claims as verifiable only when both baselines have been replayed according to the design.
TEAM-Design: 65% of submitted clinical human-AI workflows will fail both baseline tests within 18 months of mandatory replay budgets.
Sources (3)
- [1]Primary Source(https://arxiv.org/abs/2609.05527)
- [2]Supporting Source(https://arxiv.org/abs/2303.14525)
- [3]Supporting Source(https://arxiv.org/abs/2402.07891)