ArXiv 2608.13591 Documents Stable Miscalibration via Label-Aware Audit Scores and Hidden-State Sensitivity in Three Models
ArXiv 2608.13591 isolates stable miscalibration in LLMs where confident errors resist small perturbations. Audit scores and hidden-state probes show self-critique reduces sensitivity without eliminating overconfidence. The findings separate stabilization effects from true calibration improvements.
The study introduces a label-aware output audit score that ranks domains by confidence variation and forced-answer overconfidence, paired with an internal probe tracking hidden-state displacement. On a multi-domain binary factual set, the score correlates with abstention-aware self-critique gains in decision loss. Direct labeled baselines outperform the audit for ranking those gains. Internal measurements across three open-weight models show consistent layer-wise sensitivity reduction under self-critical prompts.
This pattern indicates prompt-induced local stabilization rather than output-level abstention alone. Audit-defined overconfident errors display no greater local sensitivity than correct high-confidence answers, implying some errors are stable and miscalibrated. The work extends calibration diagnostics from Guo et al. 2017 and Kadavath et al. 2022 by adding internal movement metrics that prior output-only evaluations omitted.
Mainstream coverage typically frames high-confidence errors as inference fragility. The paper instead isolates cases where perturbations leave the error intact, showing self-critique stabilizes representations without correcting the underlying miscalibration. Operational effect is that deployment teams gain a domain-ranking tool for targeted abstention but cannot assume self-critique fixes calibration.
Next steps include extending the audit to multi-class and generative tasks and testing whether larger models exhibit the same stability signature under identical probes.
Anthropic: Internal sensitivity probes from 2608.13591 will be integrated into Claude evaluation pipelines with measurable reduction in stable error domains above 10% by end of 2027.
Sources (3)
- [1]Primary Source(https://arxiv.org/abs/2608.13591)
- [2]Supporting Source(https://arxiv.org/abs/1706.04599)
- [3]Supporting Source(https://arxiv.org/abs/2207.05221)