On-Policy Self-Distillation without Any Supervision
정답 없이도 AI가 스스로 채점하고 스스로 가르쳐 수학 실력을 올린다
언어모델을 후련(post-training)시킬 때 보통 정답이나 더 큰 모델의 도움이 필요했는데, 이 연구는 모델이 같은 문제를 여러 번 풀어본 뒤 다수결로 나온 답을 임시 정답처럼 쓰는 u-OPSD 방법을 제안한다. 다수결과 다른 답을 낸 풀이만 골라 다수결에 맞춘 스스로의 예측 분포를 가르치는 방식으로, 정답 없이도 수학 벤치마크에서 정답 기반 기존 방법과 비슷하거나 더 좋은 성적을 냈다. 코드는 공개되어 있다.
METAL LAB 해설 도표
u-OPSD: 정답 없이 스스로 교사가 되는 구조
증거 상태측정 결과가 보고됨
1. 다중 롤아웃 생성같은 문제에 대해 모델이 G=8번 독립적으로 답을 생성한다
2. 다수결 투표생성된 답들 중 가장 많이 나온 답을 임시 정답으로 삼고, 이 답과 일치/불일치하는 롤아웃을 나눈다(임계값 τ=0.5 미달 시 학습에서 제외)
3. 교사 조건화다수결과 일치하는 풀이 중 가장 긴 것을 정답 풀이 대신 사용해 교사 분포를 만든다
4. 불일치 롤아웃에 증류다수결과 다른 답을 낸 학생 풀이의 각 토큰 위치에서 교사의 다음 토큰 분포를 학생에게 가르쳐(순방향 KL) 스스로 틀린 지점을 고치게 한다
METAL LAB이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.
무엇을 했나
기존 온폴리시 자기증류(OPSD) 방법은 모델이 스스로 만든 답을 학습에 쓰지만, 여전히 정답 풀이나 외부 피드백, 더 큰 모델의 도움 같은 외부 정보가 필요했다.
u-OPSD는 한 문제에 대해 모델이 여러 번(G=8) 답을 생성한 뒤, 답들이 일정 비율(임계값 τ=0.5) 이상 일치하면 다수결로 나온 답을 정답 대신 쓴다.
다수결과 일치하는 풀이 중 가장 긴 것을 교사 역할로 삼아, 다수결과 다른 답을 낸 학생 풀이에 그 교사의 다음 토큰 확률 분포를 가르쳐(distill) 학생이 스스로 틀린 지점을 고치게 한다.
수학 경시 벤치마크 5종(AIME24, AIME25, HMMT25, MATH500, AMC23)에서 Qwen3 4B/8B 비추론(non-thinking) 모드 기준 기본 모델보다 8.5~10.7점 올랐고, 정답을 쓰는 기존 OPSD보다도 2.3~3.2점 더 높았다.
추론(thinking) 모드에서는 정답 기반 OPSD와 비슷한 수준(4B에서 0.9점 우위, 8B에서 동률)이었고, 정답 기반 GRPO보다는 0.7~1.1점 높았다.
Figure 1: Comparison between OPSD / SDFT with ground-truth solution or ICLs (left), SDPO with rich feedback from the environment (middle) and our u-OPSD without any supervision.
Table 1: Performance comparison on math reasoning benchmarks for Qwen3 models with non-thinking mode.
Method
AIME24
AIME25
HMMT25
MATH500
AMC23
Avg.
Qwen3-4B
Base
25.83
17.78
10.83
84.10
66.25
40.96
w/ GT.
+ SFT
26.67
19.72
13.06
84.85
69.38
42.73
+ GRPO
25.00
22.50
15.00
86.20
80.62
45.86
+ OPSD
32.22
20.83
16.39
85.75
76.25
46.29
[1.2pt/1.4pt] w/o GT.
+ TTRL
25.00
20.83
11.94
83.60
67.50
41.77
+ RENT
22.22
20.28
11.39
84.00
74.38
42.45
+ Intuitor
23.89
20.00
11.67
83.70
70.62
41.98
+ u-OPSD
37.50
27.78
14.44
86.50
81.25
49.49
Qwen3-8B
Base
27.50
23.33
13.61
84.05
69.38
43.57
w/ GT.
+ SFT
26.94
21.67
11.94
84.10
72.50
43.43
+ GRPO
30.56
21.94
13.06
87.85
73.75
45.43
+ OPSD
41.67
28.06
18.33
87.15
85.00
52.04
[1.2pt/1.4pt] w/o GT.
+ TTRL
27.22
21.11
13.06
84.45
71.88
43.54
+ RENT
28.33
21.67
10.83
84.00
70.00
42.97
+ Intuitor
26.11
22.50
11.67
84.20
73.12
43.52
+ u-OPSD
45.56
34.72
18.61
89.55
83.12
54.31
Figure 2: Overview of Unsupervised On-policy Self-Distillation (u-OPSD), which replaces ground-truth supervision in On-Policy Self-distillation with pseudo-labels generated from the model’s own majority-vote consensus.
Table 2: Thinking mode, per benchmark at step 150, under the protocol of Table 1. Shading and bold as in that table.
Method
AIME24
AIME25
HMMT25
MATH500
AMC23
Avg.
Qwen3-4B
Base
74.17
64.72
45.56
94.80
95.00
74.85
w/ GT.
+ SFT
73.06
69.17
41.67
95.40
94.38
74.73
+ GRPO
73.89
69.44
43.61
95.45
99.38
76.35
+ OPSD
75.28
68.06
43.06
95.20
99.38
76.20
[1.2pt/1.4pt] w/o GT.
+ TTRL
72.78
68.33
45.56
95.90
96.25
75.76
+ RENT
74.72
65.83
43.06
95.25
99.38
75.65
+ Intuitor
76.39
68.33
42.78
95.35
97.50
76.07
+ u-OPSD
76.39
68.06
46.94
95.75
98.12
77.05
Qwen3-8B
Base
75.56
66.67
45.00
96.35
96.88
76.09
w/ GT.
+ SFT
76.39
69.72
43.89
95.80
95.00
76.16
+ GRPO
76.94
69.17
47.78
95.70
95.00
76.92
+ OPSD
80.83
69.72
46.67
95.75
96.88
77.97
[1.2pt/1.4pt] w/o GT.
+ TTRL
77.22
68.61
46.94
95.75
96.25
76.95
+ RENT
77.50
70.28
45.83
95.95
96.25
77.16
+ Intuitor
76.94
70.28
44.17
96.20
95.62
76.64
+ u-OPSD
76.94
71.39
47.50
96.00
98.12
77.99
Figure 3: Training curves of the Qwen3-4B runs quoted in Tables 1 and 2, on AIME24, AIME25 and MATH500 over the first four checkpoints, in thinking (left three columns) and non-thinking mode (right three); the dashed line is the base model. Top: u-OPSD against the supervised arms. Bottom: u-OPSD against the label-free ones. The axis counts checkpoints: u-OPSD saves every 25 steps and GRPO every 50, so the GRPO points span its steps 50–200. The two modes separate at a glance—in non-thinking mode u-OPSD breaks away from every other arm, while in thinking mode all methods stay within about two points of base.
Table 3: Qwen3-30B-A3B-Instruct-2507, non-thinking, scored as pass@1 rather than average@n: generation i of every problem forms one single-sample run, and i=1,2,3 give three estimates, reported as mean ± population standard deviation. Each arm is shown at the checkpoint with the best five-benchmark mean under this metric. Because the metric differs from Tables 1 and 2, the two are not comparable and the numbers are kept apart. Bold marks the best value in each column.
Model
Method
AIME24
AIME25
HMMT25
MATH500
AMC23
Avg.
Qwen3-30B-A3B -Instruct-2507
Base
80.00±0.00
63.33±2.72
43.33±2.72
96.33±0.25
95.83±1.18
75.77±0.85
OPSD
78.89±3.14
61.11±3.14
47.78±1.57
97.20±0.43
96.67±2.36
76.33±1.58
u-OPSD
75.56±5.67
65.56±4.16
50.00±0.00
97.00±0.43
99.17±1.18
77.46±2.09
Qwen3-4B -Instruct-2507
Base
66.67±7.20
53.33±5.44
27.78±4.16
93.87±0.34
93.33±1.18
67.00±2.25
OPSD
62.22±5.09
52.22±5.09
31.11±5.09
94.93±0.64
95.00±0.00
67.10±1.13
u-OPSD
68.89±5.67
57.78±1.57
28.89±3.14
94.20±0.86
94.17±1.18
68.78±1.70
Table 4: Performance under combinations of the teacher-reference selection (rows) and the distillation-target selection (columns), on Qwen3-8B non-thinking, G=8, τ=0.5, k=1. Each cell is the five-benchmark average at the best checkpoint in 25–150, with the step given in parentheses. “label-only” strips the teacher’s reference down to the boxed pseudo-label.
Teacher ref. \ distill
longest
random
shortest
label-only
43.40 (125)
43.00 (100)
41.55 (25)
shortest (default)
57.10 (125)
54.37 (75)
55.23 (150)
random
57.96 (75)
56.93 (75)
55.77 (75)
longest
59.00 (75)
57.90 (100)
55.01 (125)
Table 5: Ablation of the disagreeing-rollout selection policy (matched decay schedule). Each cell shows step 150 / best checkpoint in 25–150. All variants use G=8, τ=0.5.
Variant
AIME24
AIME25
HMMT25
MATH500
AMC23
OPSD (supervised)
27.50 / 27.50
20.00 / 23.61
12.50 / 13.33
83.80 / 84.65
71.88 / 72.50
disagree-1 (default)
33.89 / 35.28
27.78 / 27.78
14.72 / 16.11
87.60 / 87.60
79.38 / 79.38
disagree-2
29.72 / 31.11
23.33 / 27.50
15.28 / 16.67
86.20 / 86.30
75.62 / 80.00
disagree-3
27.78 / 34.72
25.83 / 31.11
13.33 / 18.33
85.45 / 87.10
72.50 / 76.25
disagree-all (no cap)
29.44 / 30.83
22.50 / 25.83
10.00 / 13.33
85.25 / 86.00
73.75 / 74.38
longest-1
33.06 / 34.72
23.33 / 26.67
14.72 / 18.33
85.50 / 86.80
73.75 / 78.12
Table 6: Comparison of divergence computation strategy: Full vocabulary is logit distillation over every token (2); sampled token evaluates the two policies only at the token the student drew (26); top-k rows truncate the teacher to its k largest entries. We report on Qwen3-8B non-thinking at the best checkpoint.
Variant
AIME24
AIME25
HMMT25
MATH500
AMC23
Avg.
Student token
29.44
19.72
12.78
84.70
70.62
43.45
top-k
20
50.00
37.78
19.44
90.20
88.12
56.94
50
48.06
37.50
24.17
89.45
85.62
55.85
100
53.89
43.33
20.56
91.65
85.62
59.01
200
51.67
40.00
22.50
90.15
88.75
58.11
Full-vocabulary
53.89
37.50
20.28
89.90
85.62
57.10
Table 7: Divergence family Dβ under u-OPSD, on Qwen3-8B non-thinking, longest-1, G=8, τ=0.5. Each arm is shown at its best checkpoint in 25–150. Objective names follow 59, who report the same comparison under gold supervision.
Objective
AIME24
AIME25
HMMT25
MATH500
AMC23
Avg.
Base
27.50
23.33
13.61
84.05
69.38
43.57
Forward KL KL(πT∥πS), β=0 (default)
53.89
37.50
20.00
89.75
84.38
57.10
Reverse KL KL(πS∥πT), β=1
training diverges (not scored)
JSD (β=0.5)
29.44
21.94
12.22
83.70
69.38
43.34
Table 8: Reproduction of published OPSD results with the released code and hyperparameters (avg@12, temperature 1.0). “pub.” denotes the numbers published in the official OPSD repository; “ours” our rerun. Checkpoint columns are steps 50/75/100/150, matching the checkpoints published for this configuration.
Config
Bench
base
checkpoints
Qwen3-4B (non-thinking)
AIME24 pub.
23.1
20.3
27.5
31.1
32.8
AIME24 ours
22.2
23.1
27.2
32.2
31.9
AIME25 pub.
21.4
21.4
20.8
21.1
21.9
AIME25 ours
17.8
20.8
23.1
20.8
21.1
HMMT25 pub.
10.8
11.1
13.1
16.4
14.4
HMMT25 ours
12.2
10.6
12.8
16.4
11.7
Table 9: GRPO under matched and mismatched reasoning modes. Each row is the best of ten checkpoints (steps 50–500) by five-benchmark mean, scored at temperature 1.0 against the base model in the evaluation mode of that row. The two “thinking → non-thinking” rows reuse the checkpoints of the rows above them; only the evaluation prompt differs.
Train
Eval
Step
AIME24
AIME25
HMMT25
MATH500
AMC23
Avg.
Δ base
Qwen3-4B
thinking
thinking
300
73.89
69.44
43.61
95.45
99.38
76.35
+1.79
thinking
non-thinking
500
25.00
21.39
13.89
83.75
76.25
44.06
+2.62
non-thinking
non-thinking
200
25.00
22.50
15.00
86.20
80.62
45.86
+4.43
Qwen3-8B
thinking
thinking
250
76.94
69.17
47.78
95.70
95.00
76.92
+0.66
thinking
non-thinking
100
27.22
22.50
13.33
85.20
67.50
43.15
+1.20
non-thinking
non-thinking
500
30.56
21.94
13.06
87.85
73.75
45.43
+3.48
Table 10: Self-consistency threshold τ on Qwen3-8B non-thinking, longest-1, G=8. Each cell shows step 150 / best checkpoint in 25–150. τ is the fraction of valid rollouts that must agree for a prompt to receive a pseudo-label; τ=0 accepts every prompt.
τ
AIME24
AIME25
HMMT25
MATH500
AMC23
Avg.
0
53.33 / 57.50
41.39 / 46.67
22.50 / 27.22
90.70 / 92.15
88.12 / 89.38
59.21 / 61.83
0.3
50.83 / 53.33
34.44 / 41.67
20.28 / 23.06
89.85 / 90.65
85.00 / 88.12
56.08 / 58.59
0.5
43.06 / 53.89
35.83 / 37.50
20.28 / 20.28
88.00 / 89.90
82.50 / 85.62
53.93 / 57.10
0.7
32.78 / 37.50
23.61 / 26.67
13.06 / 14.44
85.95 / 87.70
74.38 / 76.25
45.96 / 48.18
0.9
32.22 / 32.22
20.83 / 22.50
13.61 / 13.61
84.70 / 84.85
70.62 / 72.50
44.40 / 44.40
Table 11: Number of rollouts per prompt G on Qwen3-8B non-thinking, longest-1, τ=0.5. Each cell shows step 150 / best checkpoint in 25–150.
G
AIME24
AIME25
HMMT25
MATH500
AMC23
Avg.
4
47.22 / 50.28
35.83 / 37.78
18.06 / 21.11
88.15 / 90.20
84.38 / 87.50
54.73 / 56.99
8
43.06 / 53.89
35.83 / 37.50
20.28 / 20.28
88.00 / 89.90
82.50 / 85.62
53.93 / 57.10
12
53.61 / 57.22
40.83 / 43.33
20.56 / 26.39
91.00 / 92.00
86.25 / 90.00
58.45 / 61.79
16
51.67 / 55.28
41.67 / 44.17
22.22 / 22.22
89.90 / 89.90
87.50 / 87.50
58.59 / 59.37
Table 12: How the teacher is updated, on Qwen3-8B non-thinking, longest-1, G=8, τ=0.5. Each cell shows step 150 / best checkpoint in 25–150. “fixed” freezes the teacher at the initial policy (the base model with the adapter disabled), which is OPSD’s own setting and ours everywhere else; the EMA rows let the teacher track the student at the given decay.
Teacher
AIME24
AIME25
HMMT25
MATH500
AMC23
Avg.
fixed
43.06 / 53.89
35.83 / 37.50
20.28 / 20.28
88.00 / 89.90
82.50 / 85.62
53.93 / 57.10
EMA 0.999
43.61 / 52.78
36.67 / 41.39
18.33 / 23.33
88.40 / 90.50
87.50 / 87.50
54.90 / 58.82
EMA 0.99
52.78 / 54.17
39.72 / 43.89
23.33 / 23.33
89.85 / 91.00
83.75 / 86.88
57.89 / 58.88
EMA 0.995
54.17 / 55.00
40.28 / 43.33
21.94 / 22.50
90.10 / 90.75
83.75 / 89.38
58.05 / 59.47
Table 13: Learning-rate schedule ablation: best AIME24 checkpoint (avg@12) per method under an effectively-constant LR (30-epoch horizon, ≈5×10−6 throughout) vs. the matched 150-step linear decay used throughout Tables 1 and 2. The constant-LR column is OPSD’s native configuration, so the released figure is directly comparable there; no released run exists under the decayed schedule.
Method
Constant LR
Linear decay
OPSD (supervised)
32.22
27.50
as released
32.8
–
disagree-1
37.50
35.28
disagree-2
32.78
31.11
disagree-3
33.61
34.72
disagree-all
33.61
30.83
longest-1
31.94
34.72
실제로 확인된 결과
수학 벤치마크 5종 평균에서 Qwen3-4B/8B 비추론 모드 기준 기본 모델보다 각각 8.5점, 10.7점 향상되었고, 정답 기반 OPSD보다 각각 3.2점, 2.3점 더 높았다.
추론 모드에서는 4B에서 OPSD보다 0.9점, 8B에서 동률이었으며, GRPO보다는 각각 0.7점, 1.1점 높았다.
명령어 튜닝 모델 Qwen3-30B-A3B-Instruct-2507과 Qwen3-4B-Instruct-2507에서도 각각 75.77→77.46, 67.00→68.78로 향상되어 OPSD를 1.1점, 1.7점 앞섰다.
훈련 프롬프트 64개를 분석한 결과 롤아웃의 96.3%에서 답을 추출할 수 있었고, 94.0%의 문제가 자기일치 임계값을 넘겼으며, 그 중 86.7%가 실제 정답과 일치했다.
교사가 짧은 답(정답 라벨만)만 보고 학습하면 11.4~15.6점 성능이 떨어졌고, 순방향 KL 대신 대칭 발산(JSD)을 쓰면 13.8점 떨어져 훈련 전 모델 수준에 머물렀다.
어디에 쓸 수 있나
정답 채점 라벨을 구하기 힘든 영역에서, 정답 대신 모델 스스로의 다수결 합의를 임시 정답으로 활용해 후련시키는 파이프라인 설계
경시 수학처럼 정답이 명확히 추출 가능한 도메인에서 정답 없이 모델 성능을 끌어올리는 훈련 루프 구축
기존 OPSD/GRPO 같은 정답 기반 방법과 비교 실험을 할 때 라벨 없는 기준선으로 참고
한계와 남은 검증
실험은 Qwen3 계열(4B, 8B)과 경시 수학 도메인 한 곳에만 한정되어 검증되었고, 답을 명확히 추출·정규화할 수 있는 문제에만 다수결 투표가 적용된다.
비추론 모드에서는 이득이 크지만 추론 모드에서는 이득이 작아, 기본 모델이 이미 강할수록 개선 여지가 줄어드는 경향이 있다.
다수결이 만드는 임시 정답은 기본 모델이 가장 많이 내는 답을 그대로 반영하므로, 기본 모델의 정확도에 성능이 묶여 있고 훈련 중 오답 다수결(13.3% 사례에서 확인)이 섞일 위험이 있다.
체크포인트별 12개 샘플과 5개 벤치마크로 분산을 줄였지만, 시드를 여러 번 반복한 오차범위(seed-replicated error bar)는 아직 보고되지 않아 후속 개정에서 추가될 예정이다.
GRPO를 훈련한 추론 모드 그대로 비추론 평가로 옮기는 실험은 했지만, 반대 방향(비추론 훈련 후 추론 평가)은 아직 수행되지 않았다.
왜 중요한가
이 방법이 통하면, 정답이 없거나 채점하기 어려운 문제(코딩, 서술형 등)에서도 모델이 스스로 훈련 데이터를 만들어 실력을 키울 수 있는 길이 열린다. 정답 라벨링에 드는 비용과 정답 없는 영역에서의 확장성 문제를 동시에 줄일 잠재력이 있다는 점에서 중요하다.
이 논문의 용어
온폴리시 자기증류(OPSD) · 모델이 자기 자신이 만든 답변으로 스스로를 다시 가르치는 학습 방식
다수결 투표(majority vote) · 같은 문제를 여러 번 풀게 한 뒤 가장 많이 나온 답을 정답처럼 취급하는 방법
자기일치 임계값(τ) · 여러 번 푼 답 중 얼마나 많은 비율이 일치해야 그 답을 믿고 학습에 쓸지 정하는 기준값
GRPO · 정답과 비교한 보상값을 이용해 정책을 업데이트하는 강화학습 기법
순방향 KL 발산(forward KL) · 교사 모델과 학생 모델의 예측 분포 차이를 계산해 학생을 교사 쪽으로 맞추는 척도