Claude Opus 4 Identifies Injected Activations in Controlled Introspection Trials
The arXiv paper demonstrates measurable introspective access in frontier models through activation injection. Claude Opus 4 leads on concept identification and output discrimination tasks. Capacity stays unreliable and post-training dependent.
Researchers injected known concept vectors into model activations at inference time and measured whether subsequent self-reports matched the injected content. Claude Opus 4 and 4.1 produced correct identifications at higher rates than smaller or less post-trained models. The protocol separated genuine recall from text-based confabulation by comparing responses to raw input versus manipulated hidden states.
Accuracy varied sharply with model scale and post-training regime. Opus variants distinguished prior internal representations from fresh token inputs and used recalled intentions to flag artificial prefills versus their own generations. Control experiments showed models could also modulate specific activations when explicitly instructed to focus on a target concept.
These results extend mechanistic interpretability work on feature steering and activation patching. They indicate functional access to internal states that exceeds surface-level next-token prediction. Operational consequence is limited: the observed introspection remains context-sensitive and drops under distribution shift or adversarial prompting.
Next steps require standardized benchmarks that quantify reliability across tasks rather than isolated demonstrations. Production systems will need verifiable bounds before any deployment that treats model self-reports as diagnostic signals.
Anthropic: Activation-level introspection reliability on held-out concept sets will exceed 75% in the next Opus release before end of 2026.
Sources (3)
- [1]Primary Source(https://arxiv.org/abs/2601.01828)
- [2]Supporting Source(https://arxiv.org/abs/2309.10317)
- [3]Supporting Source(https://transformer-circuits.pub/2023/monosemantic-features)