Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress
AI가 선생님 모델을 따라 배우다가, 정답에 다가가는 '좋은 생각'까지 억누르는 문제를 잡아낸다
학생 모델이 스스로 만든 답변을 선생님 모델과 비교해 학습시키는 On-Policy Distillation(OPD) 방식은, 학생이 선생님과 다른 방식으로 생각해도 무조건 감점을 주는 맹점이 있다. 연구팀은 학생이 실제로 정답에 가까워지고 있는지를 별도로 측정해, 선생님과 다르다는 이유만으로 억눌린 좋은 추론 구간을 찾아내 그 구간의 감점을 걸러내는 R2-OPD를 제안했다. DeepSeek-R1-Distill-Qwen-1.5B 실험에서 기존 OPD 대비 평균 정확도 2.51점, 정답 중 하나라도 맞힐 확률 4.46점을 끌어올렸다.
무엇을 했나
- OPD는 학생이 만든 답변을 선생님 모델의 확률분포와 비교해 토큰 단위로 점수를 매기는데, 이 점수가 실제 추론 진전을 제대로 반영하지 못하는 경우가 있다는 걸 확인했다
- 이를 해결하기 위해 학생이 중간 지점까지 쓴 내용만 가지고 정답을 맞힐 확률을 여러 번 시뮬레이션(Neval=8회)해서 '진짜 진전 점수(process reward)'를 별도로 계산했다
- 같은 방향으로 진전을 보이는 인접 구간들을 하나로 합쳐(sign-consistent merging) 노이즈를 줄인 뒤, 이 진전 점수 순위와 선생님 유사도 점수 순위가 어긋나는 구간을 찾아 해당 구간의 학습 신호를 걸러냈다(masking)
- DeepSeek-R1-Distill-Qwen-1.5B에서 JustRL을 선생님으로 쓴 실험 결과 기존 OPD 대비 평균 정확도(avg@4) 2.51점, 4개 중 하나라도 맞힐 확률(pass@4) 4.46점 향상했고, AIME 대회 수학 문제에서 특히 큰 개선을 보였다
- Qwen3-1.7B와 e3-1.7B 조합으로도 실험해 다른 모델 계열에서도 같은 방법이 통한다는 것을 확인했다

| Method | AIME 24 | AIME 25 | Olympiad | Avg. | ||||
|---|---|---|---|---|---|---|---|---|
| avg@4 | pass@4 | avg@4 | pass@4 | avg@4 | pass@4 | avg@4 | pass@4 | |
| Student | 22.50 | 43.33 | 23.33 | 36.67 | 43.19 | 58.31 | 29.67 | 46.10 |
| Teacher | 41.67 | 56.67 | 30.83 | 43.44 | 53.28 | 68.43 | 41.92 | 56.18 |
| OPD (1) | 28.33 | 50.00 | 22.50 | 30.00 | 46.86 | 62.10 | 32.55 | 47.37 |
| E-OPD (17) | 18.33 | 36.67 | 12.50 | 23.33 | 49.91 | 64.56 | 26.91 | 41.52 |
| TIP-OPD (36) | 17.50 | 33.33 | 11.67 | 20.00 | 48.24 | 62.83 | 25.80 | 38.72 |
| IW-OPD (34) | 20.83 | 36.67 | 17.50 | 26.67 | 43.59 | 59.09 | 27.31 | 40.81 |
| Uni-OPD (13) | 20.00 | 43.33 | 19.17 | 26.67 | 53.16 | 69.97 | 30.78 | 46.66 |
| R2-OPD (Ours) | 32.50 | 56.67 | 25.83 | 36.67 | 46.86 | 62.19 | 35.06 | 51.83 |

| Dataset | Base | OPD | R2-OPD | |||
|---|---|---|---|---|---|---|
| avg@4 | pass@4 | avg@4 | pass@4 | avg@4 | pass@4 | |
| AIME 24 | 24.17 | 30.00 | 22.50 | 36.67 | 25.00 | 40.00 |
| AIME 25 | 18.75 | 20.00 | 27.50 | 33.33 | 25.83 | 36.67 |
| Olympiad | 50.10 | 63.87 | 53.82 | 67.09 | 54.31 | 67.91 |
| Avg. | 31.01 | 37.96 | 34.61 | 45.70 | 35.04 | 48.19 |

| Dataset | No Merge | R2-OPD | Δ |
|---|---|---|---|
| AIME 24 | 17.5 | 32.50 | 15.0 |
| AIME 25 | 11.67 | 25.83 | 14.16 |
| Olympiad | 47.8 | 46.86 | -0.96 |

| Setting | Value |
|---|---|
| Optimization and sequence settings | |
| Training epochs | 1 |
| Optimizer | AdamW |
| Learning rate | 5×10−6 |
| Global batch size | 64 |
| Maximum prompt length | 1,024 tokens |
| Maximum response length | 7,168 tokens |
| KL support size | Student top-16 (H=16) |
| Process-reward estimation | |
| Rollouts per evaluated boundary | Neval=8 |
| Sampling temperature | 0.7 |
| Top-k | 50 |
| Top-p | 1.0 |
| Maximum rollout length | 300 tokens |
| Segment filtering | |
| Minimum sentences between boundaries | Smin=3 |
| Minimum merged segments | nmin=3 |
| Segment masking ratio | q=30% |

| Condition | Action |
|---|---|
| The string-level pre-check does not find gi in yi, or segmentation yields fewer than two segments. | Skip process-reward rollouts and retain all token-level OPD supervision. |
| The response has fewer than max(3,nmin) merged segments. | Do not apply segment masking. |
| No strict adjacent pair is available after process-reward ranking. | Do not apply segment masking. |
| All candidate segments have zero inconsistency score. | Do not apply segment masking. |
| At least one candidate has a positive inconsistency score. | Mask up to the response-specific budget; retain every other token. |
| Category | Matched markers |
|---|---|
| Reconsideration | wait, hold on, let me reconsider, hmm |
| Correction | actually |
| Verification | let me check |
| Alternative reasoning | alternatively |
왜 중요한가
선생님 모델을 그대로 흉내 내게 하는 기존 지식 증류 방식은 학생이 독창적이지만 옳은 풀이 경로를 찾아도 벌점을 줄 수 있다는 구조적 약점을 드러낸 연구다. 추론 능력이 중요한 수학·코딩 AI를 더 저렴하게, 더 잘 훈련시키려는 회사나 연구자에게 실질적인 훈련 전략 개선안을 제시한다.
이 논문의 용어
- On-Policy Distillation(OPD) · 학생 모델이 직접 만든 답변을 선생님 모델에게 채점받아 학습하는 지식 증류 방식
- 역방향 KL 발산(reverse KL divergence) · 학생의 확률분포가 선생님의 확률분포와 얼마나 다른지 측정하는 수치로, 학습 신호로 쓰인다
- 프로세스 리워드(process reward) · 중간 단계까지의 풀이만으로 정답을 맞힐 확률이 앞 단계 대비 얼마나 올랐는지를 나타내는 점수
- sign-consistent merging(부호 일치 병합) · 진전 방향(플러스/마이너스)이 같은 인접 구간들을 하나로 묶어 노이즈를 줄이는 기법
- avg@4 / pass@4 · avg@4는 4번 시도한 정답률의 평균, pass@4는 4번 중 한 번이라도 맞혔는지를 보는 지표
논문 원문 초록 (영문)
On-policy distillation (OPD) has emerged as an effective framework for post-training language models by pairing student-generated trajectories with dense token-level supervision from a teacher. However, OPD implicitly assumes that teacher-derived rewards are an appropriate proxy for reasoning progress, and therefore treats all teacher feedback equally during policy optimization. While in practice, this assumption does not always hold. We observe that teacher-derived rewards often conflict with genuine reasoning progress, as reasoning steps with clear reasoning advancement may still receive lower distillation rewards, simply due to deviation from teacher's outputs. To address this mismatch, we propose Reasoning-Progress-Aware Reward Filtering for On-Policy Distillation (R2-OPD), which constructs two within-trajectory rankings of reasoning spans, one from teacher-derived rewards and the other from independently estimated progress reward. Distillation rewards are selectively suppressed whenever the two rankings disagree, reducing supervision that conflicts with reasoning progress while preserving effective teacher guidance. Our approach shows consistent improvement over standard OPD especially regarding reasoning performances.
arXiv에서 원문 보기최신 논문
- Specification-delta-driven data governance: an empirical study of the {\guillemotleft}spec-delta{\guillemotright} as the unit of change in lakehouse data platforms데이터 플랫폼 변경도 코드처럼 '설계도 조각'을 붙여서 검토하면 어떨까: 실험 설계 논문
- Are LLMs becoming similarly creative? Evidence from three years of models최신 AI 챗봇일수록 서로 비슷한 답을 내놓는다는 3년치 조사 결과
- Auditing Cross-Lingual Fairness in Language Model WatermarkingAI 생성 텍스트를 잡아내는 워터마크 기술이 영어 아닌 언어에서는 훨씬 부실하게 작동하고, 그 격차는 개별 언어가 아니라 언어 계열 단위로 나타난다
- TESTNAV: Pareto-Guided Search for Compositional Robustness TestingAI 모델을 여러 손상이 겹친 입력으로 시험할 때, 굳이 다 테스트하지 않고도 '진짜 위험한 실패'만 골라내는 탐색법
- Optimal Skill Selection for LLM Agents with Provable Bicriteria GuaranteesAI 에이전트에게 어떤 '스킬 문서'를 몇 개나 줘야 잘 작동하는지, 수학적으로 최적해를 보장하며 골라주는 방법
- Reliable Financial Named Entity Recognition under Domain Shift금융 AI가 서류체 문장에서 배운 자신감은 트위터로 가면 거짓말이 된다
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving논문 속 시연이 아니라 실제 서비스에 넣을 수 있는 희소 어텐션 만들기
- Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction텍스트가 빠지거나 망가져도, AI가 그 자리를 대신할 '가짜 텍스트'를 한 번에 만들지 않고 여러 번 고쳐가며 감정을 더 정확히 읽어낸다
METAL LAB 최신 기사
그림 출처: Chen Yang et al., arXiv:2608.19408, arxiv-nonexclusive