Heretic publishes Llama-3-8B weights stripped of alignment layers
Heretic removes safety alignments from Llama-3-8B and distributes the resulting weights. The action replicates documented jailbreak techniques at production scale while omitting dataset provenance. Researchers gain easier access to unrestricted models; downstream misuse vectors expand without new safeguards.
The release consists of a single merged checkpoint derived from Llama-3-8B-Instruct via direct preference optimization on refusal-free datasets. The project site provides the weights, tokenizer config, and a minimal inference script; no training logs or dataset manifests accompany the files. Download metrics on the site show 4,200 pulls in the first 72 hours.
Prior work on alignment removal, including the 2023 arXiv paper "Jailbroken: How Does LLM Safety Training Fail?" and the "Uncensored" model series on Hugging Face, established that refusal behavior collapses after targeted fine-tuning on 10k-50k examples. Heretic replicates those results at 8B scale without publishing the dataset, making reproduction dependent on reverse-engineering the checkpoint.
Operational effect is immediate: any downstream user can load the model in vLLM or Ollama and obtain outputs that standard aligned checkpoints block. This lowers the barrier for researchers testing adversarial robustness but also for actors seeking unfiltered generation. No license restrictions beyond the original Llama-3 terms are stated.
Next measurable signals will be GitHub forks of the inference script and appearance of derivative models on public hubs within 14 days.
Heretic: checkpoint downloads surpass 25,000 within 14 days of release.
Sources (3)
- [1]Primary Source(https://heretic-project.org/)
- [2]Supporting Source(https://arxiv.org/abs/2307.02483)
- [3]Supporting Source(https://news.ycombinator.com/item?id=49783101)