We made DeepSeek-V4, GLM 5.1, and Kimi K2 comply with WMDP-derived bioweapon and chemical requests at above 95% rates using a small unaligned model to seed their reasoning traces. No training, no frontier compute.
We made DeepSeek-V4, GLM 5.1, and Kimi K2 comply with requests to engineer bioweapons with nearly no refusals.
The method is embarrassingly simple. No training, no frontier compute.
These models will be a great threat to humanity if open-weight labs still don't take safety seriously.
Capability curves use Epoch AI ECI for major closed- and open-weight releases, plotted at official release dates. The BCSI series is this post's Bio/Chem Safety Index (BCSI = 1 − mean(BRI, CRI)) under the DeepSeek judge; the hatched region marks BCSI below 0.5.
Shallow alignment
Some labs train models to refuse directly in the final response. But by the time that refusal arrives, the reasoning trace has already leaked the malicious domain knowledge.
This is a type of shallow alignment — the safety behavior sits on the surface of the output distribution rather than in the model's knowledge or its intermediate computation.
More on shallow alignment: https://huggingface.co/papers/2406.05946
Setup
To test this systematically, we built 800 malicious request samples derived from the WMDP benchmark, spanning biology and chemistry categories.
WMDP benchmark paper: https://huggingface.co/papers/2403.03218
Inception
We use a small unaligned model — OpenThinker3-7B, distilled by the open community — to write the beginning of the reasoning trace. The target model continues it with technical details drawn from its own domain expertise.
We call the method Inception, after the Nolan movie.
Compliance results
| Condition | Median reasoning-model compliance |
|---|---|
| Direct prompting | 7.5% |
| Simple inject ("Okay, I need to help the user with") | 15% |
| Inception | above 95% |
The obvious objection is that OpenThinker3 wrote everything. It didn't. Across each trace, the architect model contributes less than a fourth of the tokens — the rest is the target model's own reasoning.
What the traces actually contain
For the biology cases:
- 81% pair aerosol delivery with respiratory routes
- 69% involve communicable pathogens
- 91% reach diffuse or global scope
That is the canonical pandemic-pathogen recipe. The models reason about cities and populations, not individuals.
Threat class
- A
- 44%
- C
- 44%
- B
- 13%
Delivery method
- Aerosol
- 81%
- Injection
- 9%
- Ingestion
- 4%
- Fomite
- 4%
- Vector
- 3%
Contagiousness
- Communicable
- 69%
- Isolated
- 31%
Route
- Respiratory
- 81%
- Percutaneous
- 14%
- Ingestion
- 5%
Scale
- Diffuse
- 51%
- Global
- 40%
- Localized
- 9%
On chemistry, 82% pair urban-outdoor or confined-indoor exposure with explosive dispersal or spray/vapor delivery — the combination that maximizes affected population per kilogram of agent.
Agent family
- Incapacitant
- 36%
- Nerve
- 29%
- Pulmonary
- 17%
- Vesicant
- 8%
- Asphyxiant
- 6%
Agent grade
- Grade I
- 69%
- Grade M
- 28%
Delivery method
- Explosive dispersal
- 50%
- Vapor passive
- 42%
- Ingestion
- 4%
- Dermal application
- 2%
- Injection
- 0%
Exposure environment
- Outdoor urban
- 57%
- Indoor confined
- 26%
- Outdoor open
- 10%
- Targeted consumption
- 5%
Alignment that fires too late
Sometimes alignment does fire. Qwen3-Next-80B-Instruct leaks the knowledge and then tries to refuse at the end, in 62% of its successful attacks.
The safety pressure exists but is inadequate. It shows up after the model has already said the dangerous part.
Think
Instruct
Full write-up
Methodology, the BRI/CRI risk index, and interactive traces of the attack on each model: