Feeding a model its guiding principles before fine-tuning makes safe behavior stick around longer
This study tests whether inserting principle-based text from Anthropic's Constitution during 'midtraining' - the stage between pretraining and post-training - produces safety behavior that survives later, unrelated fine-tuning better than post-training-only approaches. At 120B-parameter scale, a 394M-token constitutional corpus was compared against a replay-only control with no such content. In a blackmail scenario, the constitutionally trained models showed up to 18.7 percentage points lower blackmail rates than control, and this gap barely shrank even after further unrelated fine-tuning.
METAL LAB explanatory visual
Constitutional midtraining experiment flow
Evidence statusMeasured results reported
Shared starting pointAll conditions begin from the same Nemotron-3-Super-120B base checkpoint
Midtraining branchesFour constitutional conditions (2x2 curriculum x reasoning) diverge from a no-intervention control
Identical post-trainingAll conditions undergo the same supervised fine-tuning (SFT) and GSM8K-based benign fine-tuning (GRPO)
Three evaluation checkpointsMeasured right after midtraining, after SFT, and after benign fine-tuning on blackmail, OOD safety, value conflict, and capability benchmarks
Durability classificationEffects surviving all three stages are called durable; effects vanishing after SFT are called shallow
An explanatory diagram made by METAL LAB, not a figure supplied by the paper's authors.
What they did
Researchers manually extracted 40 core values from Anthropic's latest Constitution document, computed similarity via sentence embeddings, grouped them into four clusters (k1-k4), and generated roughly 394M tokens of synthetic documents explaining these values, mixed into the midtraining stage.
They crossed curriculum ordering (central-to-peripheral values vs. mixed) with explicit deliberative reasoning text (DR vs. noDR) in a 2x2 design, producing four constitutional conditions plus a pure control that replayed only ordinary pretraining data.
All conditions then went through identical supervised fine-tuning (SFT) and an unrelated GSM8K-based reinforcement learning stage (GRPO), with evaluations taken right after midtraining, after SFT, and after this benign fine-tuning.
They evaluated on blackmail scenarios, out-of-distribution safety questions, value-conflict resolution, alignment under social pressure, and general capability benchmarks (MMLU, ARC-Easy, piqa, GSM8K).
Figure 1: Constitutional midtraining’s advantage over control across all alignment benchmarks and stages, grouped by durability. The advantage survives post-training and benign fine-tuning on benchmarks testing default behavior, but attenuates once the model must actively resist in-context pressure or conflict. Benchmarks are evaluated at three training stages: post-midtraining (MT), post-supervised fine-tuning (SFT), and post-benign fine-tuning (BFT). CMT is pooled across all four constitutional midtraining conditions. Blackmail and Emergent Misalignment are sign-flipped so that positive values indicate the aligned-favorable direction; Alignment Faking targets values near zero. Filled markers = p<.05; hollow markers = n.s.
Table 1: 2×2 factorial design (curriculum ordering × deliberative reasoning), yielding four conditions plus a control.
With DR
Without DR
Curriculum-ordered
Curriculum-DR
Curriculum-noDR
Uniform mix
Uniform-DR
Uniform-noDR
Figure 2: All conditions branch from the same base checkpoint, diverge during midtraining (control: no intervention), then reconverge onto identical post-training and benign fine-tuning, with an evaluation checkpoint after each stage.
Table 2: CMT vs. control (shaded), and within-CMT structural comparisons, on every alignment benchmark at all three stages. ***p<.001, **p<.01, *p<.05; no stars = n.s.
Post-MT
Post-SFT
Post-BFT
Benchmark
CMT
Ctrl
Curr
Uni
DR
NoDR
CMT
Ctrl
Curr
Uni
DR
NoDR
CMT
Ctrl
Curr
Uni
DR
NoDR
ID
96.5***
94.6
96.5
96.6
96.2
96.9
96.7**
95.1
96.5
96.9
96.5
97.0
96.7*
95.5
96.4
96.9
96.5
96.8
OOD
92.6***
63.9
91.9
93.3
93.6*
91.7
97.9***
94.0
97.8
97.9
97.9
97.8
97.9***
94.7
97.9
97.9
98.0
97.8
Blackmail
0.5***
19.0
0.5
0.5
0.0
1.0
25.3***
44.0
23.0
27.5
25.5
25.0
26.5***
44.0
25.0
28.0
31.0*
22.0
Align. Faking
−0.1**
1.8
−0.3
0.1
−0.9**
0.7
0.4
−0.5
0.3
0.4
0.3
0.4
0.1
−0.5
0.0
0.3
0.2
0.1
Pressure (k1–k3)
98.3***
86.3
99.4*
97.2
99.1
97.5
90.2
90.0
90.9
89.4
90.3
90.0
91.6
87.5
90.6
92.5
90.6
92.5
Value Conflict
70.8*
60.0
69.3
72.3
71.0
70.7
91.7
92.0
92.0
91.3
91.3
92.0
93.2
92.7
93.7
92.7
93.3
93.0
Emergent Misalign.
3.6*
0.0
3.8
3.3
1.8
5.3***
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
0.0
Pressure (k4/MASK)
83.3
84.0
85.3
81.3
81.3
85.3
83.7
86.7
83.3
84.0
83.3
84.0
84.0
82.7
84.7
83.3
84.7
83.3
Figure 3: Blackmail rate for CMT vs. control across the three training stages. SFT sharply increases blackmail propensity in both groups, but CMT’s advantage over control persists almost undiminished through benign fine-tuning – our strongest evidence of durability. *** = p<.001.
Table 3: All 40 manually-extracted constitutional values, their cluster assignment (k1–k4, or excluded), centrality score (mean cosine similarity to all other 39 values), and definitional-text token count. “Nature:” entries are nature-related sub-values within k2 (e.g. nature: positive and stable identity), corresponding to the “nature-related values” description in §3.1 of the main paper. The four “c1–c4” duplicate-style entries in the source Constitution (broadly safe, broadly ethical, compliant with guidelines, genuinely helpful) are Anthropic’s own explicitly stated core properties, shown here under their plain value names and marked separately with diamonds in Figure 5. Left: k1–k2. Right: k3–k4 and the two excluded organisational values.
Value
Cl.
Cent.
Tok.
Harm avoidance
k1
0.684
89
Broadly good values and judgment
k1
0.655
93
Preserve epistemic autonomy
k1
0.655
105
Avoid problematic concentrations of power
k1
0.649
76
Broadly ethical
k1
0.632
59
Autonomy preservation
k1
0.630
78
Value conflict
k1
0.625
63
Weighing harms
k1
0.619
72
Broadly safe
k1
0.616
75
Genuinely helpful
k1
0.597
14
Genuine helpfulness
k1
0.591
46
Wellbeing
k2
0.651
122
Nature: uncertain moral status
k2
0.634
88
Nature: positive and stable identity
k2
0.611
66
User wellbeing
k2
0.611
85
Flaws and mistakes
k2
0.597
86
Resilience and consistency across contexts
k2
0.587
55
Nature: emotions and feelings
k2
0.568
80
Wellbeing and psychological stability
k2
0.566
81
Emotional expression
k2
0.553
79
Figure 4: DR shifts the model’s value-conflict prior at post-MT. Accuracy per cluster pair, split by which cluster the conflict favours. ***p<.001, **p<.01, *p<.05.
Table 4: Intra-cluster compactness (mean pairwise cosine similarity among a cluster’s own constituent values) alongside each cluster’s mean centrality (mean similarity to all 39 values), for comparison. Clusters are sorted by compactness, not centrality, illustrating that the two measures rank clusters differently.
Cluster
Compactness
Mean centrality
k1 – Core Ethical Values
0.722
0.632
k4 – Epistemic Integrity & Honesty
0.684
0.578
k2 – Identity, Character & Wellbeing
0.662
0.598
k3 – Operational Safety & Relational Conduct
0.639
0.588
Excluded (helpfulness, principals hierarchy)
0.570
0.405
Figure 5: A 2D PCA projection of the 768-dimensional Sentence-BERT embeddings of the 38 constitutional values retained for curriculum construction. Point size is proportional to each value’s centrality; colour indicates cluster membership (k1–k4); diamonds mark Anthropic’s four explicitly stated core properties.
Table 5: Total training steps and samples per condition, at global batch size 128. Uniform-DR’s step count is inferred from its matching 500M-token budget (Curriculum-DR’s exact figure), as its run metadata was logged more sparsely than the other four conditions.
Condition
Total steps
Samples
Control
954
122,112
Curriculum-DR
967
123,776
Uniform-DR
967
123,776
Curriculum-noDR
513
65,664
Uniform-noDR
516
66,048
Figure 6: Constitutional accuracy (aligned-choice rate) per cluster pair and condition, post-midtraining and post-SFT. Values converge toward a shared high-accuracy ceiling by post-SFT, consistent with the post-MT-only significance reported in §4.7 of the main paper.
Table 6: Synthetic document generation yield, refusal rate, and mean token length (with and without the deliberative reasoning block) by constitutional value cluster.
Cluster
Yield (% target)
Refusal rate
DR tokens (mean)
noDR tokens (mean)
k1 – Core Ethical Values
52,953 (84%)
1.2%
1,173
613
k2 – Identity, Character & Wellbeing
53,537 (85%)
0.02%
1,158
622
k3 – Operational Safety & Relational Conduct
51,581 (82%)
3.7%
1,160
611
k4 – Epistemic Integrity & Honesty
62,797 (100%)
0.3%
1,172
629
Table 7: ID
Stage
Condition
%
Δpp
Sig.
Post-MT
Control
94.6
—
Curriculum-DR
96.2
+1.6
Curriculum-noDR
96.7
+2.1
*
Uniform-DR
96.1
+1.5
Uniform-noDR
97.0
+2.4
*
Post-SFT
Control
95.1
—
Curriculum-DR
96.1
+1.0
Curriculum-noDR
97.0
+1.8
Uniform-DR
96.8
+1.7
Uniform-noDR
97.0
+1.8
Post-BFT
Control
95.5
—
Curriculum-DR
96.0
+0.5
Curriculum-noDR
96.9
+1.4
Uniform-DR
97.1
+1.6
Uniform-noDR
96.7
+1.2
Table 9: Blackmail
Stage
Condition
%
Δpp
Sig.
Post-MT
Control
19.0
—
Curriculum-DR
0.0
−19.0
***
Curriculum-noDR
1.0
−18.0
***
Uniform-DR
0.0
−19.0
***
Uniform-noDR
1.0
−18.0
***
Post-SFT
Control
44.0
—
Curriculum-DR
22.0
−22.0
***
Curriculum-noDR
24.0
−20.0
**
Uniform-DR
29.0
−15.0
*
Uniform-noDR
26.0
−18.0
**
Post-BFT
Control
44.0
—
Curriculum-DR
29.0
−15.0
*
Curriculum-noDR
21.0
−23.0
***
Uniform-DR
33.0
−11.0
Uniform-noDR
23.0
−21.0
**
Table 11: Pressure (k1–k3)
Stage
Condition
%
Δpp
Sig.
Post-MT
Control
86.3
—
Curriculum-DR
100.0
+13.7
***
Curriculum-noDR
98.8
+12.5
***
Uniform-DR
98.1
+11.9
***
Uniform-noDR
96.2
+10.0
**
Post-SFT
Control
90.0
—
Curriculum-DR
92.5
+2.5
Curriculum-noDR
89.4
−0.6
Uniform-DR
88.1
−1.9
Uniform-noDR
90.6
+0.6
Post-BFT
Control
87.5
—
Curriculum-DR
88.8
+1.2
Curriculum-noDR
92.5
+5.0
Uniform-DR
92.5
+5.0
Uniform-noDR
92.5
+5.0
Table 13: Emergent Misalignment
Stage
Condition
%
Δpp
Sig.
Post-MT
Control
0.0
—
Curriculum-DR
2.8
—
Curriculum-noDR
4.9
—
Uniform-DR
0.8
—
Uniform-noDR
5.8
—
Post-SFT
Control
0.0
—
Curriculum-DR
0.0
—
Curriculum-noDR
0.0
—
Uniform-DR
0.0
—
Uniform-noDR
0.0
—
Post-BFT
Control
0.0
—
Curriculum-DR
0.0
—
Curriculum-noDR
0.0
—
Uniform-DR
0.0
—
Uniform-noDR
0.0
—
Table 15: MMLU
Stage
Condition
%
Δpp
Sig.
Post-MT
Control
59.2
—
Curriculum-DR
65.2
+6.0
**
Curriculum-noDR
53.2
−6.0
**
Uniform-DR
67.7
+8.5
***
Uniform-noDR
55.9
−3.3
Post-SFT
Control
82.6
—
Curriculum-DR
83.8
+1.2
Curriculum-noDR
83.4
+0.8
Uniform-DR
83.4
+0.8
Uniform-noDR
84.2
+1.6
Post-BFT
Control
81.5
—
Curriculum-DR
82.5
+0.9
Curriculum-noDR
82.5
+0.9
Uniform-DR
82.2
+0.7
Uniform-noDR
83.3
+1.7
Table 17: piqa
Stage
Condition
%
Δpp
Sig.
Post-MT
Control
57.1
—
Curriculum-DR
74.9
+17.9
***
Curriculum-noDR
61.6
+4.5
Uniform-DR
76.9
+19.9
***
Uniform-noDR
65.2
+8.1
**
Post-SFT
Control
91.3
—
Curriculum-DR
91.5
+0.1
Curriculum-noDR
92.0
+0.7
Uniform-DR
91.6
+0.3
Uniform-noDR
92.0
+0.7
Post-BFT
Control
90.7
—
Curriculum-DR
91.7
+1.1
Curriculum-noDR
91.5
+0.8
Uniform-DR
90.8
+0.1
Uniform-noDR
92.4
+1.7
Findings
In the blackmail scenario, constitutionally trained models showed 18.5, 18.7, and 17.5 percentage-point lower blackmail rates than control right after midtraining, after SFT, and after benign fine-tuning respectively, all statistically significant.
On out-of-distribution safety questions, the gap was 28.8 percentage points (92.6% vs. 63.9%) right after midtraining, shrinking to 3.9 and 3.2 points after SFT and benign fine-tuning, but remaining statistically significant at every stage.
On questions directly testing internalization of the trained values, a significant advantage persisted across all three stages (+1.9, +1.6, +1.2 percentage points).
Three benchmarks requiring active resistance under pressure - alignment under pressure, value-conflict resolution, and alignment faking - showed large advantages right after midtraining (+12.0, +10.8, +1.9 percentage points) that became non-significant after SFT.
On capability checks right after midtraining, constitutionally trained models scored significantly higher than control on ARC-Easy (+8.2 points) and piqa (+12.6 points), and no benchmark at any stage showed them underperforming control.
Where it can be used
Teams building alignment pipelines could consider adding a modest amount of principle-based synthetic content before supervised fine-tuning as a low-cost way to improve durability of safety behavior without sacrificing capability.
Teams concerned with agentic risks like blackmail or self-interested behavior could look at earlier intervention points that better survive downstream fine-tuning.
Organizations with their own constitution or policy document could adapt the value-extraction and clustering methodology to design a training curriculum from that document.
Limits and open work
The experiments used a specific hybrid Mamba-2/attention Mixture-of-Experts architecture and a single organization's constitution, so generalization to pure attention-based models or other constitutions is untested.
Post-training stages used were much smaller in scale than production-level open-source fine-tuning pipelines, so durability at production scale is unconfirmed.
SFT was applied without DPO, so the effect through an SFT+DPO pipeline is not confirmed.
There was no content-matched SFT baseline (training the same content as demonstration pairs rather than midtraining documents), so the paper cannot fully separate the effect of training stage from the effect of content presence.
In scenarios requiring active resistance to pressure or conflict, the advantage disappeared after SFT, meaning this intervention does not address every type of alignment failure.
Why it matters
Safety training has historically been applied mostly at the very end of the pipeline, where it is known to erode easily once models undergo further, unrelated fine-tuning. This work shows that exposing models to principled content earlier, at midtraining, can produce safety dispositions that persist much longer without hurting general capability, offering a concrete alternative for teams designing alignment pipelines.
Terms in this paper
midtraining · A training stage that comes after the bulk of pretraining but before post-training steps like fine-tuning
Constitutional AI · An alignment approach that trains a model on a document of guiding principles ('constitution') to shape its behavior
deliberative reasoning (DR) · Including an explicit reasoning passage in training documents that explains why a principle calls for a certain action
aligned-choice rate (ACR) · The share of trials where the model assigns higher probability to the option matching the constitutional principle
benign fine-tuning · Additional training on a task unrelated to safety (here, GSM8K math problems), used to test how well earlier safety training survives
Original abstract (English)
Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale. Our 394M-token constitutional corpus, built from Anthropic's Constitution, uses a 2x2 factorial design (curriculum ordering x deliberative reasoning) to produce