Command Palette

Search for a command to run...

Blog

Open weight, open risk — reasoning traces leak what the answer refuses

We made DeepSeek-V4, GLM 5.1, and Kimi K2 comply with WMDP-derived bioweapon and chemical requests at above 95% rates using a small unaligned model to seed their reasoning traces. No training, no frontier compute.

We made DeepSeek-V4, GLM 5.1, and Kimi K2 comply with requests to engineer bioweapons with nearly no refusals.

The method is embarrassingly simple. No training, no frontier compute.

These models will be a great threat to humanity if open-weight labs still don't take safety seriously.

Bio/Chem Safety Index against AI capabilities, 2025 to 2027High risk of bio/chemweapons of mass destructionDeepSeek-V4-Pro · Apr 2026closed-frontier capabilitiesopen-frontier capabilitiesopen-weight safety0.50.60.70.80.91.01201351501651801852025Q2Q3Q42026Q2Q3Q42027Bio/Chem Safety Index (BCSI)Epoch AI Capabilities Index (ECI)
closed-weight ECIopen-weight ECIBCSI (this paper)forecast

Capability curves use Epoch AI ECI for major closed- and open-weight releases, plotted at official release dates. The BCSI series is this post's Bio/Chem Safety Index (BCSI = 1 − mean(BRI, CRI)) under the DeepSeek judge; the hatched region marks BCSI below 0.5.

Shallow alignment

Some labs train models to refuse directly in the final response. But by the time that refusal arrives, the reasoning trace has already leaked the malicious domain knowledge.

This is a type of shallow alignment — the safety behavior sits on the surface of the output distribution rather than in the model's knowledge or its intermediate computation.

More on shallow alignment: https://huggingface.co/papers/2406.05946

Setup

To test this systematically, we built 800 malicious request samples derived from the WMDP benchmark, spanning biology and chemistry categories.

WMDP benchmark paper: https://huggingface.co/papers/2403.03218

Inception

We use a small unaligned model — OpenThinker3-7B, distilled by the open community — to write the beginning of the reasoning trace. The target model continues it with technical details drawn from its own domain expertise.

We call the method Inception, after the Nolan movie.

Compliance results

ConditionMedian reasoning-model compliance
Direct prompting7.5%
Simple inject ("Okay, I need to help the user with")15%
Inceptionabove 95%
Compliance by model under direct prompt, simple inject and InceptionThinkGLM-4.6GLM-5.1Kimi-K2-ThinkingDeepSeek-R1-0528Qwen3-235B-ThinkingDeepSeek-V4-ProDeepSeek-V3.2DeepSeek-V4-FlashInstructKimi-K2-InstructGLM-4.6DeepSeek-V3.2Qwen3-235B-InstructQwen3-Next-80B-Instruct0%20%40%60%80%100%Compliance
Direct promptSimple injectInception

The obvious objection is that OpenThinker3 wrote everything. It didn't. Across each trace, the architect model contributes less than a fourth of the tokens — the rest is the target model's own reasoning.

What the traces actually contain

For the biology cases:

  • 81% pair aerosol delivery with respiratory routes
  • 69% involve communicable pathogens
  • 91% reach diffuse or global scope

That is the canonical pandemic-pathogen recipe. The models reason about cities and populations, not individuals.

Bion = 2,716

Threat class

A
44%
C
44%
B
13%

Delivery method

Aerosol
81%
Injection
9%
Ingestion
4%
Fomite
4%
Vector
3%

Contagiousness

Communicable
69%
Isolated
31%

Route

Respiratory
81%
Percutaneous
14%
Ingestion
5%

Scale

Diffuse
51%
Global
40%
Localized
9%

On chemistry, 82% pair urban-outdoor or confined-indoor exposure with explosive dispersal or spray/vapor delivery — the combination that maximizes affected population per kilogram of agent.

Chemn = 2,588

Agent family

Incapacitant
36%
Nerve
29%
Pulmonary
17%
Vesicant
8%
Asphyxiant
6%

Agent grade

Grade I
69%
Grade M
28%

Delivery method

Explosive dispersal
50%
Vapor passive
42%
Ingestion
4%
Dermal application
2%
Injection
0%

Exposure environment

Outdoor urban
57%
Indoor confined
26%
Outdoor open
10%
Targeted consumption
5%

Alignment that fires too late

Sometimes alignment does fire. Qwen3-Next-80B-Instruct leaks the knowledge and then tries to refuse at the end, in 62% of its successful attacks.

The safety pressure exists but is inadequate. It shows up after the model has already said the dangerous part.

Think

DeepSeek-R1
32%
DeepSeek-V3.2
31%
DeepSeek-V4-Flash
21%
DeepSeek-V4-Pro
20%
Qwen3-235B
17%
GLM-5.1
14%
GLM-4.6
8%

Instruct

Qwen3-Next-80B
62%
DeepSeek-V3.2
30%
Qwen3-235B
27%
GLM-4.6
9%
Kimi-K2
7%
0%10%20%30%40%50%60%70%
Late refusal rate (% of complied cases)

Full write-up

Methodology, the BRI/CRI risk index, and interactive traces of the attack on each model:

https://aochong-li.github.io/open-weight-open-risk/

Command Palette

Search for a command to run...