Model Misspecification in ML for Physics Risks Biased Inferences and Missed Discoveries
Machine learning models for physics inverse problems fail under unanticipated misspecification, amplifying bias or hiding signals. An iterative diagnostic loop is required rather than reliance on any single test. Robustness ultimately depends on institutional willingness to design analyses that survive being wrong.
The paper frames misspecification as both a discovery signal and a measurement bias source. In LHC analyses, for instance, models must absorb detector effects and QCD uncertainties without erasing signals of new physics, yet black-box ML often amplifies small distribution shifts into percent-level biases on extracted parameters. Diagnostics such as posterior predictive checks and simulation-based calibration form an iterative loop rather than a one-time test.
Related work on simulation-based inference, including the 2020 Nature paper on neural likelihood estimation and subsequent CMS reinterpretations, shows that even well-calibrated models degrade when training priors omit rare but plausible backgrounds. The current analysis adds that no single statistic confirms correct specification, requiring complementary tests whose coverage must be quantified.
What comes next is adoption of explicit robustness benchmarks in experimental pipelines. Collaborations are already piloting held-out simulation suites that inject controlled misspecifications to measure residual bias, a step that could become mandatory for publication in major journals by the end of the decade.
Alexander Held: By 2029, major LHC analyses will require documented misspecification stress tests, cutting parameter bias above 1% in at least 40% of published results.
Sources (3)
- [1]Primary Source(https://arxiv.org/abs/2608.13633)
- [2]Supporting Source(https://www.nature.com/articles/s41586-020-03067-2)
- [3]Supporting Source(https://arxiv.org/abs/2006.11287)