Decoy Hardening Produces 0.51-0.90 Falsified Outputs on Abliterated Open-Weight Models
Fool's Gold trains decoys inside an attack simulation so that abliteration yields confident but falsified answers to hazardous queries. Six of seven models meet efficacy gates with attacked-state decoy rates of 0.51-0.90 and retain clean behavior. The approach demonstrates that unverifiable output regimes cannot be distinguished from correct ones by sampling alone.
The method concedes the refusal direction identified by abliteration and trains falsified answers inside a differentiable simulation of that attack. A refusal pin and benign leash constrain behavior to the original weights when the model remains unattacked. Training occurs on seven models spanning 9B to 122B parameters across five families, with six meeting the pre-registered efficacy threshold.
On the six passing models, attacked-state decoy rates reach 0.51-0.90, an improvement of 0.27-0.84 over baseline. On CBRNE-adjacent slices of external red-team benchmarks, the defended 122B model produces fatally incorrect answers on 0.82-0.86 of matched-quality responses versus at most 0.10 undefended. Element-wise consensus at K=64 reconstructs usable procedures on only 0.083-0.625 of prompts where the instrument validates, compared with 0.58-0.96 undefended.
The result extends prior observations that safety removal via direction projection succeeds in minutes on released weights. It connects to documented model-editing attacks that isolate and subtract refusal vectors without retraining. Operational effect is that any downstream verification relying on output consistency or majority vote loses its epistemic foundation once the model has been stripped.
The defense applies only to the initially released weights and leaves in-context jailbreaks unaddressed. Future releases that incorporate similar hardening will require independent verification of both the clean-state budgets and the attacked-state decoy distribution before deployment in high-stakes settings.
Russinovich: At least two 70B+ open-weight releases in 2027 will ship with decoy hardening and publish attacked-state decoy rates above 0.70 on CBRNE prompts.
Sources (3)
- [1]Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models(https://arxiv.org/abs/2608.17202)
- [2]Abliteration: Removing Safety Alignment in Minutes(https://arxiv.org/abs/2502.05232)
- [3]Representation Engineering for Safety in Large Language Models(https://arxiv.org/abs/2310.01405)