Reinforcement Learning Agent Forms Dark Holes in Simulated Coronagraph Testbed
Preprint demonstrates RL-based post-coronagraphic wavefront control that creates dark holes in simulation. Evidence is limited to synthetic data; real-hardware confirmation is absent. Relevance lies in potential 12-month path to improved contrast on space observatories.
The arXiv preprint from Laura Manuela Castaneda Medina describes a fully learned agent that actuates a deformable mirror from raw focal-plane images augmented with physics-informed phase estimates. Training occurred entirely in simulation of a post-coronagraphic testbed; the policy learned to minimize speckle intensity in target regions without an explicit wavefront reconstruction step. Sample size consisted of repeated episodes on the same simulated optical train; no real hardware data were used.
Performance approached classical focal-plane wavefront control in final contrast but required no hand-crafted sensing matrix. The approach directly targets the sub-nanometer stability needed for 10^-10 contrast exoplanet imaging on future space telescopes. Because the agent interacts continuously with the system, it can in principle adapt to slowly drifting aberrations that defeat static controllers.
The study remains limited to simplified simulations that omit realistic detector noise, non-common-path errors, and thermal drifts present on flight hardware. Validation on a physical testbed with closed-loop deformable-mirror control is the required next measurement before any claim of operational readiness.
Within twelve months the method could be tested on existing ground-based extreme adaptive-optics benches such as SPHERE or GPI-2, providing the first hardware benchmark against model-based controllers already flying.
Castaneda Medina: First closed-loop dark-hole contrast below 10^-7 on a physical testbed will be reported within 18 months.
Sources (2)
- [1]Primary Source(https://arxiv.org/abs/2609.20880)
- [2]Supporting Source(https://arxiv.org/abs/2309.12345)