RL Agent Using GP Surrogates Records 2x Improvement Over Manual Tuning on APOLLO Polarized Target Data
Reinforcement learning with Gaussian-process uncertainty estimates yields a 2x gain over expert manual tuning of microwave frequency in dynamically polarized nuclear targets. The approach transfers across target samples via a standardized simulator but lacks published live-beam or long-term drift results.
The framework trains multilayer perceptrons and Gaussian-process regressors on microwave frequency, beam current, and accumulated dose to predict polarization. Gaussian-process models supply calibrated uncertainty that flags out-of-distribution regions; MLPs do not. The GP approximation is placed inside a Gym-style simulator so a single agent can be trained across successive target samples without retraining the surrogate from scratch.
Operational logs show the trained agent selects frequency steps that raise polarization roughly twice as far as the recorded manual adjustments under identical beam and dose conditions. The lower-confidence-bound reward explicitly trades off expected polarization against predictive variance, preventing the policy from exploiting high-uncertainty regions that would be unsafe in a real experiment.
Prior control literature on accelerator RF tuning (e.g., arXiv:2006.12687) and nuclear magnetic resonance feedback (Phys. Rev. Accel. Beams 24, 2021) used model-predictive or PID loops without distributional-shift detection. The present work adds explicit uncertainty gating and cross-sample transfer, yet reports no live-beam validation or failure-mode analysis under sudden dose spikes.
Next steps require closed-loop deployment at an operating facility with independent safety interlocks and a minimum six-month dataset to test whether the 2x margin holds when material properties drift beyond the original training envelope.
Kasparian et al.: Six-month closed-loop deployment at APOLLO will sustain at least 1.8x median improvement versus recorded operator actions with zero safety-interlock trips.
Sources (3)
- [1]Primary Source(https://arxiv.org/abs/2610.02452)
- [2]Supporting Source(https://arxiv.org/abs/2006.12687)
- [3]Supporting Source(https://journals.aps.org/prab/abstract/10.1103/PhysRevAccelBeams.24.104601)