Anthropic Blocked 47 Biological Weapons Queries via Claude 3.7
Anthropic's classifiers blocked 47 biological weapons queries on Claude 3.7. Data shows earlier refusals than peer models but increased friction for valid research. The episode reveals a concrete implementation of proactive misuse prevention rather than after-the-fact policy statements.
Anthropic deployed updated refusal classifiers on Claude 3.7 that flagged queries containing specific pathogen synthesis steps and delivery mechanisms. Internal logs showed 12 queries reached the 80 percent risk threshold before termination. These detections occurred after the model refused initial benign-seeming prompts about vaccine research that contained hidden follow-up sequences.
Model cards released in June 2026 report a 3.2 percent refusal rate on high-risk biosecurity prompts, up from 1.1 percent in the prior version. Independent red-team evaluations by the Center for AI Safety documented similar refusal patterns across three frontier models but noted Anthropic's classifiers triggered 22 percent earlier in conversation threads. Coverage omitted that OpenAI and Google DeepMind reported zero comparable public blocks during the same window.
The pattern indicates Anthropic has shifted from post-hoc auditing to real-time prompt-level intervention. This approach aligns with its published constitution clauses on catastrophic risk but creates measurable latency increases of 180 milliseconds per classified query. Operational impact includes higher false-positive rates on legitimate synthetic biology research, with academic users reporting 9 percent more refusals on non-weapon topics.
Next filings with the voluntary AI safety commitments due December 2026 will require disclosure of classifier false-negative rates. Regulators are expected to request raw query samples for external audit.
Anthropic: False-negative rate on bio-risk prompts will fall below 0.5 percent by March 2027.
Sources (2)
- [1]Anthropic Model Card v3.7(https://assets.anthropic.com/model-cards/claude-3-7.pdf)
- [2]Center for AI Safety Red Team Report 2026(https://www.safe.ai/reports/red-teaming-2026)