Label Provenance Drives 29-Fold Performance Gap in TESS Exoplanet ML Detectors
The study demonstrates that training-label quality, not model choice, is the primary bottleneck in machine-learning transit detection on TESS data. Performance varies far more across label sources than across architectures, and the physical SNR limit caps further gains without improved photometry or annotation methods.
The Astro-Hunters pipeline applies standard detrending and seven sliding-window features before gradient-boosted classification. Researchers generated labels from published ephemerides on twelve confirmed hosts rather than from the photometry itself, then compared six classifier families. This design isolates annotation quality as the experimental variable and shows that isolation-forest or unconverted-epoch labels collapse performance to chance while ephemeris-derived labels reach AUC 0.788.
The dominant finding is that observational limits, not architecture, set the ceiling: median single-cadence SNR of 2.10 caps achievable AUC at 0.932. This pattern echoes label-noise studies in medical imaging and particle physics, where annotation provenance routinely outweighs model capacity. Phase-folding remains the essential step that converts the weak per-cadence signal into a detectable periodic signature.
Future wide-field surveys such as PLATO and the Roman Space Telescope will face identical annotation bottlenecks at higher cadence volumes. The result implies that investment in robust ephemeris cross-matching or physics-informed label generation will yield larger gains than incremental classifier upgrades. A Box Least Squares baseline already recovers eight of twelve periods, underscoring that classical methods retain utility when labels are reliable.
Astro-Hunters team: Per-cadence AUC above 0.85 on new TESS sectors will require median SNR >3.0 achieved only after 2028 photometry upgrades.
Sources (2)
- [1]Primary Source(https://arxiv.org/abs/2608.18172)
- [2]Supporting Source(https://arxiv.org/abs/2009.01853)