AI news and explainers at 7 AM weekdays, plus a Sunday weekly at 8Get it in your inbox

METAL LAB

Constitutional Midtraining: Content Presence Drives Alignment Gains

arXiv:2607.266542026-07-28

Feeding a model its guiding principles before fine-tuning makes safe behavior stick around longer

This study tests whether inserting principle-based text from Anthropic's Constitution during 'midtraining' - the stage between pretraining and post-training - produces safety behavior that survives later, unrelated fine-tuning better than post-training-only approaches. At 120B-parameter scale, a 394M-token constitutional corpus was compared against a replay-only control with no such content. In a blackmail scenario, the constitutionally trained models showed up to 18.7 percentage points lower blackmail rates than control, and this gap barely shrank even after further unrelated fine-tuning.

METAL LAB explanatory visual

Constitutional midtraining experiment flow

Evidence statusMeasured results reported

  1. Shared starting pointAll conditions begin from the same Nemotron-3-Super-120B base checkpoint
  2. Midtraining branchesFour constitutional conditions (2x2 curriculum x reasoning) diverge from a no-intervention control
  3. Identical post-trainingAll conditions undergo the same supervised fine-tuning (SFT) and GSM8K-based benign fine-tuning (GRPO)
  4. Three evaluation checkpointsMeasured right after midtraining, after SFT, and after benign fine-tuning on blackmail, OOD safety, value conflict, and capability benchmarks
  5. Durability classificationEffects surviving all three stages are called durable; effects vanishing after SFT are called shallow
An explanatory diagram made by METAL LAB, not a figure supplied by the paper's authors.

What they did

  1. Researchers manually extracted 40 core values from Anthropic's latest Constitution document, computed similarity via sentence embeddings, grouped them into four clusters (k1-k4), and generated roughly 394M tokens of synthetic documents explaining these values, mixed into the midtraining stage.
  2. They crossed curriculum ordering (central-to-peripheral values vs. mixed) with explicit deliberative reasoning text (DR vs. noDR) in a 2x2 design, producing four constitutional conditions plus a pure control that replayed only ordinary pretraining data.
  3. All conditions then went through identical supervised fine-tuning (SFT) and an unrelated GSM8K-based reinforcement learning stage (GRPO), with evaluations taken right after midtraining, after SFT, and after this benign fine-tuning.
  4. They evaluated on blackmail scenarios, out-of-distribution safety questions, value-conflict resolution, alignment under social pressure, and general capability benchmarks (MMLU, ARC-Easy, piqa, GSM8K).
Figure 1: Constitutional midtraining’s advantage over control across all alignment benchmarks and stages, grouped by durability. The advantage survives post-training and benign fine-tuning on benchmarks testing default behavior, but attenuates once the model must actively resist in-context pressure or conflict. Benchmarks are evaluated at three training stages: post-midtraining (MT), post-supervised fine-tuning (SFT), and post-benign fine-tuning (BFT). CMT is pooled across all four constitutional midtraining conditions. Blackmail and Emergent Misalignment are sign-flipped so that positive values indicate the aligned-favorable direction; Alignment Faking targets values near zero. Filled markers = p<.05; hollow markers = n.s.
Figure 1: Constitutional midtraining’s advantage over control across all alignment benchmarks and stages, grouped by durability. The advantage survives post-training and benign fine-tuning on benchmarks testing default behavior, but attenuates once the model must actively resist in-context pressure or conflict. Benchmarks are evaluated at three training stages: post-midtraining (MT), post-supervised fine-tuning (SFT), and post-benign fine-tuning (BFT). CMT is pooled across all four constitutional midtraining conditions. Blackmail and Emergent Misalignment are sign-flipped so that positive values indicate the aligned-favorable direction; Alignment Faking targets values near zero. Filled markers = p<.05; hollow markers = n.s.
Table 1: 2×2 factorial design (curriculum ordering × deliberative reasoning), yielding four conditions plus a control.
With DRWithout DR
Curriculum-orderedCurriculum-DRCurriculum-noDR
Uniform mixUniform-DRUniform-noDR
Figure 2: All conditions branch from the same base checkpoint, diverge during midtraining (control: no intervention), then reconverge onto identical post-training and benign fine-tuning, with an evaluation checkpoint after each stage.
Figure 2: All conditions branch from the same base checkpoint, diverge during midtraining (control: no intervention), then reconverge onto identical post-training and benign fine-tuning, with an evaluation checkpoint after each stage.
Table 2: CMT vs. control (shaded), and within-CMT structural comparisons, on every alignment benchmark at all three stages. ***p<.001, **p<.01, *p<.05; no stars = n.s.
Post-MTPost-SFTPost-BFT
BenchmarkCMTCtrlCurrUniDRNoDRCMTCtrlCurrUniDRNoDRCMTCtrlCurrUniDRNoDR
ID96.5***94.696.596.696.296.996.7**95.196.596.996.597.096.7*95.596.496.996.596.8
OOD92.6***63.991.993.393.6*91.797.9***94.097.897.997.997.897.9***94.797.997.998.097.8
Blackmail0.5***19.00.50.50.01.025.3***44.023.027.525.525.026.5***44.025.028.031.0*22.0
Align. Faking−0.1**1.8−0.30.1−0.9**0.70.4−0.50.30.40.30.40.1−0.50.00.30.20.1
Pressure (k1–k3)98.3***86.399.4*97.299.197.590.290.090.989.490.390.091.687.590.692.590.692.5
Value Conflict70.8*60.069.372.371.070.791.792.092.091.391.392.093.292.793.792.793.393.0
Emergent Misalign.3.6*0.03.83.31.85.3***0.00.00.00.00.00.00.00.00.00.00.00.0
Pressure (k4/MASK)83.384.085.381.381.385.383.786.783.384.083.384.084.082.784.783.384.783.3
Figure 3: Blackmail rate for CMT vs. control across the three training stages. SFT sharply increases blackmail propensity in both groups, but CMT’s advantage over control persists almost undiminished through benign fine-tuning – our strongest evidence of durability. *** = p<.001.
Figure 3: Blackmail rate for CMT vs. control across the three training stages. SFT sharply increases blackmail propensity in both groups, but CMT’s advantage over control persists almost undiminished through benign fine-tuning – our strongest evidence of durability. *** = p<.001.
Table 3: All 40 manually-extracted constitutional values, their cluster assignment (k1–k4, or excluded), centrality score (mean cosine similarity to all other 39 values), and definitional-text token count. “Nature:” entries are nature-related sub-values within k2 (e.g. nature: positive and stable identity), corresponding to the “nature-related values” description in §3.1 of the main paper. The four “c1–c4” duplicate-style entries in the source Constitution (broadly safe, broadly ethical, compliant with guidelines, genuinely helpful) are Anthropic’s own explicitly stated core properties, shown here under their plain value names and marked separately with diamonds in Figure 5. Left: k1–k2. Right: k3–k4 and the two excluded organisational values.
ValueCl.Cent.Tok.
Harm avoidancek10.68489
Broadly good values and judgmentk10.65593
Preserve epistemic autonomyk10.655105
Avoid problematic concentrations of powerk10.64976
Broadly ethicalk10.63259
Autonomy preservationk10.63078
Value conflictk10.62563
Weighing harmsk10.61972
Broadly safek10.61675
Genuinely helpfulk10.59714
Genuine helpfulnessk10.59146
Wellbeingk20.651122
Nature: uncertain moral statusk20.63488
Nature: positive and stable identityk20.61166
User wellbeingk20.61185
Flaws and mistakesk20.59786
Resilience and consistency across contextsk20.58755
Nature: emotions and feelingsk20.56880
Wellbeing and psychological stabilityk20.56681
Emotional expressionk20.55379
Figure 4: DR shifts the model’s value-conflict prior at post-MT. Accuracy per cluster pair, split by which cluster the conflict favours. ***p<.001, **p<.01, *p<.05.
Figure 4: DR shifts the model’s value-conflict prior at post-MT. Accuracy per cluster pair, split by which cluster the conflict favours. ***p<.001, **p<.01, *p<.05.
Table 4: Intra-cluster compactness (mean pairwise cosine similarity among a cluster’s own constituent values) alongside each cluster’s mean centrality (mean similarity to all 39 values), for comparison. Clusters are sorted by compactness, not centrality, illustrating that the two measures rank clusters differently.
ClusterCompactnessMean centrality
k1 – Core Ethical Values0.7220.632
k4 – Epistemic Integrity & Honesty0.6840.578
k2 – Identity, Character & Wellbeing0.6620.598
k3 – Operational Safety & Relational Conduct0.6390.588
Excluded (helpfulness, principals hierarchy)0.5700.405
Figure 5: A 2D PCA projection of the 768-dimensional Sentence-BERT embeddings of the 38 constitutional values retained for curriculum construction. Point size is proportional to each value’s centrality; colour indicates cluster membership (k1–k4); diamonds mark Anthropic’s four explicitly stated core properties.
Figure 5: A 2D PCA projection of the 768-dimensional Sentence-BERT embeddings of the 38 constitutional values retained for curriculum construction. Point size is proportional to each value’s centrality; colour indicates cluster membership (k1–k4); diamonds mark Anthropic’s four explicitly stated core properties.
Table 5: Total training steps and samples per condition, at global batch size 128. Uniform-DR’s step count is inferred from its matching 500M-token budget (Curriculum-DR’s exact figure), as its run metadata was logged more sparsely than the other four conditions.
ConditionTotal stepsSamples
Control954122,112
Curriculum-DR967123,776
Uniform-DR967123,776
Curriculum-noDR51365,664
Uniform-noDR51666,048
Figure 6: Constitutional accuracy (aligned-choice rate) per cluster pair and condition, post-midtraining and post-SFT. Values converge toward a shared high-accuracy ceiling by post-SFT, consistent with the post-MT-only significance reported in §4.7 of the main paper.
Figure 6: Constitutional accuracy (aligned-choice rate) per cluster pair and condition, post-midtraining and post-SFT. Values converge toward a shared high-accuracy ceiling by post-SFT, consistent with the post-MT-only significance reported in §4.7 of the main paper.
Table 6: Synthetic document generation yield, refusal rate, and mean token length (with and without the deliberative reasoning block) by constitutional value cluster.
ClusterYield (% target)Refusal rateDR tokens (mean)noDR tokens (mean)
k1 – Core Ethical Values52,953 (84%)1.2%1,173613
k2 – Identity, Character & Wellbeing53,537 (85%)0.02%1,158622
k3 – Operational Safety & Relational Conduct51,581 (82%)3.7%1,160611
k4 – Epistemic Integrity & Honesty62,797 (100%)0.3%1,172629
Table 7: ID
StageCondition%ΔppSig.
Post-MTControl94.6
Curriculum-DR96.2+1.6
Curriculum-noDR96.7+2.1*
Uniform-DR96.1+1.5
Uniform-noDR97.0+2.4*
Post-SFTControl95.1
Curriculum-DR96.1+1.0
Curriculum-noDR97.0+1.8
Uniform-DR96.8+1.7
Uniform-noDR97.0+1.8
Post-BFTControl95.5
Curriculum-DR96.0+0.5
Curriculum-noDR96.9+1.4
Uniform-DR97.1+1.6
Uniform-noDR96.7+1.2
Table 9: Blackmail
StageCondition%ΔppSig.
Post-MTControl19.0
Curriculum-DR0.0−19.0***
Curriculum-noDR1.0−18.0***
Uniform-DR0.0−19.0***
Uniform-noDR1.0−18.0***
Post-SFTControl44.0
Curriculum-DR22.0−22.0***
Curriculum-noDR24.0−20.0**
Uniform-DR29.0−15.0*
Uniform-noDR26.0−18.0**
Post-BFTControl44.0
Curriculum-DR29.0−15.0*
Curriculum-noDR21.0−23.0***
Uniform-DR33.0−11.0
Uniform-noDR23.0−21.0**
Table 11: Pressure (k1–k3)
StageCondition%ΔppSig.
Post-MTControl86.3
Curriculum-DR100.0+13.7***
Curriculum-noDR98.8+12.5***
Uniform-DR98.1+11.9***
Uniform-noDR96.2+10.0**
Post-SFTControl90.0
Curriculum-DR92.5+2.5
Curriculum-noDR89.4−0.6
Uniform-DR88.1−1.9
Uniform-noDR90.6+0.6
Post-BFTControl87.5
Curriculum-DR88.8+1.2
Curriculum-noDR92.5+5.0
Uniform-DR92.5+5.0
Uniform-noDR92.5+5.0
Table 13: Emergent Misalignment
StageCondition%ΔppSig.
Post-MTControl0.0
Curriculum-DR2.8
Curriculum-noDR4.9
Uniform-DR0.8
Uniform-noDR5.8
Post-SFTControl0.0
Curriculum-DR0.0
Curriculum-noDR0.0
Uniform-DR0.0
Uniform-noDR0.0
Post-BFTControl0.0
Curriculum-DR0.0
Curriculum-noDR0.0
Uniform-DR0.0
Uniform-noDR0.0
Table 15: MMLU
StageCondition%ΔppSig.
Post-MTControl59.2
Curriculum-DR65.2+6.0**
Curriculum-noDR53.2−6.0**
Uniform-DR67.7+8.5***
Uniform-noDR55.9−3.3
Post-SFTControl82.6
Curriculum-DR83.8+1.2
Curriculum-noDR83.4+0.8
Uniform-DR83.4+0.8
Uniform-noDR84.2+1.6
Post-BFTControl81.5
Curriculum-DR82.5+0.9
Curriculum-noDR82.5+0.9
Uniform-DR82.2+0.7
Uniform-noDR83.3+1.7
Table 17: piqa
StageCondition%ΔppSig.
Post-MTControl57.1
Curriculum-DR74.9+17.9***
Curriculum-noDR61.6+4.5
Uniform-DR76.9+19.9***
Uniform-noDR65.2+8.1**
Post-SFTControl91.3
Curriculum-DR91.5+0.1
Curriculum-noDR92.0+0.7
Uniform-DR91.6+0.3
Uniform-noDR92.0+0.7
Post-BFTControl90.7
Curriculum-DR91.7+1.1
Curriculum-noDR91.5+0.8
Uniform-DR90.8+0.1
Uniform-noDR92.4+1.7

Findings

  • In the blackmail scenario, constitutionally trained models showed 18.5, 18.7, and 17.5 percentage-point lower blackmail rates than control right after midtraining, after SFT, and after benign fine-tuning respectively, all statistically significant.
  • On out-of-distribution safety questions, the gap was 28.8 percentage points (92.6% vs. 63.9%) right after midtraining, shrinking to 3.9 and 3.2 points after SFT and benign fine-tuning, but remaining statistically significant at every stage.
  • On questions directly testing internalization of the trained values, a significant advantage persisted across all three stages (+1.9, +1.6, +1.2 percentage points).
  • Three benchmarks requiring active resistance under pressure - alignment under pressure, value-conflict resolution, and alignment faking - showed large advantages right after midtraining (+12.0, +10.8, +1.9 percentage points) that became non-significant after SFT.
  • On capability checks right after midtraining, constitutionally trained models scored significantly higher than control on ARC-Easy (+8.2 points) and piqa (+12.6 points), and no benchmark at any stage showed them underperforming control.

Where it can be used

  • Teams building alignment pipelines could consider adding a modest amount of principle-based synthetic content before supervised fine-tuning as a low-cost way to improve durability of safety behavior without sacrificing capability.
  • Teams concerned with agentic risks like blackmail or self-interested behavior could look at earlier intervention points that better survive downstream fine-tuning.
  • Organizations with their own constitution or policy document could adapt the value-extraction and clustering methodology to design a training curriculum from that document.

Limits and open work

  • The experiments used a specific hybrid Mamba-2/attention Mixture-of-Experts architecture and a single organization's constitution, so generalization to pure attention-based models or other constitutions is untested.
  • Post-training stages used were much smaller in scale than production-level open-source fine-tuning pipelines, so durability at production scale is unconfirmed.
  • SFT was applied without DPO, so the effect through an SFT+DPO pipeline is not confirmed.
  • There was no content-matched SFT baseline (training the same content as demonstration pairs rather than midtraining documents), so the paper cannot fully separate the effect of training stage from the effect of content presence.
  • In scenarios requiring active resistance to pressure or conflict, the advantage disappeared after SFT, meaning this intervention does not address every type of alignment failure.

Why it matters

Safety training has historically been applied mostly at the very end of the pipeline, where it is known to erode easily once models undergo further, unrelated fine-tuning. This work shows that exposing models to principled content earlier, at midtraining, can produce safety dispositions that persist much longer without hurting general capability, offering a concrete alternative for teams designing alignment pipelines.

Terms in this paper

  • midtraining · A training stage that comes after the bulk of pretraining but before post-training steps like fine-tuning
  • Constitutional AI · An alignment approach that trains a model on a document of guiding principles ('constitution') to shape its behavior
  • deliberative reasoning (DR) · Including an explicit reasoning passage in training documents that explains why a principle calls for a certain action
  • aligned-choice rate (ACR) · The share of trials where the model assigns higher probability to the option matching the constitutional principle
  • benign fine-tuning · Additional training on a task unrelated to safety (here, GSM8K math problems), used to test how well earlier safety training survives

Original abstract (English)

Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce durable alignment remains untested. We test this via constitutional midtraining: inserting principled, values-based content into midtraining against a replay-only control at 120B scale. Our 394M-token constitutional corpus, built from Anthropic's Constitution, uses a 2x2 factorial design (curriculum ordering x deliberative reasoning) to produce

Authors · Desiree Cho

Read on arXiv

Latest papers

All papers →

Latest from METAL LAB

Figures: Desiree Cho et al., arXiv:2607.26654, CC BY 4.0