VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation
arXiv:2608.186072026-08-18
영상과 소리를 같이 만드는 AI, 어떤 결과물이 더 나은지 사람처럼 판단하는 채점관을 만들었다
영상과 음성을 동시에 생성하는 AI 모델을 강화학습으로 개선하려면 좋고 나쁨을 판단할 채점 기준(보상 신호)이 필요한데, 기존에는 화질·음질·싱크 등을 따로따로 점수 매겨 합치다 보니 사람이 보기엔 어색해도 점수만 높은 결과물이 나오는 문제가 있었다. 연구팀은 사람이 직접 두 영상 중 더 나은 것을 고르고 그 이유를 남긴 대규모 데이터셋(VAPref-10K)을 만들고, 이를 바탕으로 단계적 학습을 거쳐 사람처럼 이유를 대며 판단하는 채점 모델 VA-Judger를 개발했다. 이 채점 모델을 보상으로 써서 영상-음성 생성 모델(LTX-2)을 다시 학습시키자 사람 평가에서 기존 방식보다 훨씬 높은 선호도를 얻었다.
무엇을 했나
기존 방식은 화질, 음질, 싱크 등 항목별 자동 채점 기준을 따로 합쳐서 보상으로 썼는데, 이는 실제 사람의 선호와 잘 맞지 않아 'AI가 채점 기준의 허점만 노려 점수를 따는' 현상(리워드 해킹)이 발생했다
연구팀은 실제 유튜브·영화 등 영상 9천여 개에서 뽑은 문구로 여러 오픈소스 생성 모델의 결과물을 만들고, 사람이 두 개씩 비교해 더 나은 쪽과 그 이유를 표시한 1만여 건의 비교 데이터(VAPref-10K)를 구축했다
채점 모델 VA-Judger는 3단계로 학습된다: 1) 차이가 뚜렷한 쉬운 사례로 기본 판단법과 답변 형식을 익히고, 2) 차이가 미묘한 어려운 사례는 사람이 직접 확인한 답만 걸러 재학습하고, 3) 강화학습으로 프롬프트 일치도·화질·음질·싱크 등 항목별 세부 점수까지 정교하게 다듬는다
항목별 자동 채점 기준들은 사람의 선호와 일치하는 정확도가 50~57%에 그쳤지만, VA-Judger는 이를 10%포인트 이상 웃도는 정확도를 기록했고, 처음 보는(학습에 없던) 폐쇄형 모델들의 결과물 비교에서도 일관되게 더 정확했다
VA-Judger를 보상으로 써서 영상-음성 생성 모델 LTX-2를 재학습시키자, 200명 대상 실제 사람 평가에서 62.3%가 VA-Judger로 학습한 결과물을 최고로 선택해 기존 방식(27.6%)과 기본 모델(10.1%)을 크게 앞섰다
Figure 1: Evaluation of VA-Judger. (a) VA-Judger selects the preferred clip through structured dimension-wise reasoning, whereas separate evaluation metrics disagree with human judgment. (b) VA-Judger achieves the highest agreement with human preferences. (c) VA-Judger provides the reward for post-training LTX-2, improving the quality over the base model and OmniNFT.
Table 1: Accuracy (%) against human pairwise preferences. Single-dimension evaluation metrics and reward models are both evaluated on the 1,150-pair VA-Judger-Bench and report Total Acc / Parsed Acc. Unparseable responses count as incorrect in Total Acc.
Model
Easy
In-domain
Out-of-domain
Overall
Single-dimension evaluation metrics
Video quality: VideoAlign
62.39
52.97
48.35
54.29
Audio quality: AudioBox
58.70
50.00
44.00
50.43
Text-video: CLIP Score
55.56
54.66
53.58
54.50
Text-audio: ImageBind T-A
58.48
52.97
51.48
54.29
Audio-video: ImageBind
61.30
55.51
51.83
55.94
Synchronization: SynchFormer
61.30
46.61
48.35
52.71
Overall: Javis Score
60.00
55.51
54.96
56.88
Reward models
Qwen3-Omni Captioner (No CoT)
57.00 / 57.00
54.00 / 54.00
56.20 / 56.20
57.22 / 57.22
Qwen3-Omni Captioner (CoT)
52.50 / 57.65
49.60 / 54.87
46.60 / 56.97
50.43 / 58.61
Qwen3-Omni Instruct NoCoT
60.75 / 60.75
54.00 / 54.00
54.60 / 54.60
56.61 / 56.61
Qwen3-Omni Instruct CoT
63.25 / 64.71
54.80 / 55.92
55.00 / 55.33
57.83 / 58.69
+ Easy Cold Start
72.00 / 72.00
59.20 / 59.20
56.20 / 56.20
62.35 / 62.35
+ Hard SFT
74.50 / 74.50
63.60 / 63.60
60.20 / 60.20
65.91 / 65.91
+ GRPO (VA-Judger)
76.25 / 76.25
66.00 / 66.00
63.40 / 63.40
68.43 / 68.43
Figure 2: Overview of the VA-Judger training pipeline. Stage 1 cold-starts the model on easy preference pairs. Stage 2 aligns the model on harder pairs through human rejection sampling. Stage 3 performs GRPO using answer-level and dimension-level rewards grounded in human reasons.
Table 2: Results on the randomly sampled 200-prompt JavisBench subset. Best results are highlighted in green and second-best results are underlined. Higher is better unless marked with ↓. The two Δ rows report the raw change produced by VA-Judger post-training.
Video Quality
AudioBox Quality
Cross-Modal Alignment
Model
VQ ↑
MQ ↑
AQ ↑
AB-CE ↑
AB-CU ↑
AB-PC ↑
AB-PQ ↑
TV-Align ↑
TA-Align ↑
ViCLIP ↑
LTX-2
2.248
0.697
4.767
3.957
5.904
3.041
6.166
0.303
0.105
0.232
LTX-2 + OmniNFT
3.727
0.947
5.399
4.745
6.797
3.150
6.905
0.279
0.133
0.219
LTX-2 + VA-Judger
3.942
1.183
5.610
5.136
6.766
3.606
6.932
0.310
0.180
0.242
Δ vs. LTX-2
+1.693
+0.486
+0.843
+1.179
+0.862
+0.564
+0.766
+0.007
+0.075
+0.009
Δ vs. OmniNFT
+0.214
+0.236
+0.211
+0.391
-0.031
+0.456
+0.026
+0.031
+0.047
+0.023
Figure 3: Qualitative comparisons of LTX-2, OmniNFT, and our VA-Judger post-trained model. VA-Judger better follows the requested temporal actions in both examples.
Table 3: Category distributions of the full JavisBench pool and the random evaluation subset. Each cell reports the number and percentage of examples.
Dimension
Category
Full pool (10,140)
Random subset (200)
Event Scenario
Living Scenario
4,029 (39.7%)
73 (36.5%)
Natural Scenario
2,452 (24.2%)
37 (18.5%)
Urban Scenario
2,140 (21.1%)
66 (33.0%)
Industrial Scenario
1,090 (10.7%)
20 (10.0%)
Virtual Scenario
429 (4.2%)
4 (2.0%)
Visual Style
Camera Shooting
8,586 (84.7%)
174 (87.0%)
2D Animation
915 (9.0%)
17 (8.5%)
3D Animation
639 (6.3%)
9 (4.5%)
Sound Type
Musical Sounds
5,065 (50.0%)
73 (36.5%)
Ambient Sounds
4,183 (41.3%)
126 (63.0%)
Biological Sounds
2,215 (21.8%)
59 (29.5%)
Mechanical Sounds
1,940 (19.1%)
49 (24.5%)
Speech Sounds
1,090 (10.7%)
29 (14.5%)
Spatial Composition
Multiple Subjects
7,588 (74.8%)
167 (83.5%)
Single Subject
2,502 (24.7%)
33 (16.5%)
Off-screen Sound
2,165 (21.4%)
59 (29.5%)
Temporal Composition
Sequential Events
3,698 (36.5%)
46 (23.0%)
Simultaneous Events
3,306 (32.6%)
99 (49.5%)
Single Event
3,136 (30.9%)
55 (27.5%)
Figure 4: Human preference rates over 200 three-way comparisons. For each prompt, participants select the best video-audio output among LTX-2, LTX-2 post-trained with OmniNFT, and LTX-2 post-trained with VA-Judger.
Table 4: Human preference accuracy (%) of video-quality metrics.
Metric
Easy
In-domain
Out-of-domain
Overall
VideoAlign Overall∗
62.39
52.97
48.35
54.29
VideoPhy2
52.86
54.39
50.22
51.66
Motion Quality
59.91
51.06
49.74
53.66
Visual Quality
62.42
52.34
42.93
51.70
HPSv3
56.30
52.12
50.61
52.95
Video Aesthetic
57.52
50.21
47.21
51.50
Video Technical/Aesthetic
57.39
51.69
42.78
49.72
Video Technical/Overall
59.13
52.97
42.26
50.35
Video Technical/Technical
54.13
54.24
47.13
50.98
Figure 6: Overview of VAPref-10K construction and the raw data distribution. Real video-audio clips are captioned by Qwen3.5-Omni, rewritten into text prompts by an LLM, and passed to LTX-2, OVI, and DaVinci-MagiHuman to generate candidate clips.
Table 5: Human preference accuracy (%) of audio-quality metrics.
Metric
Easy
In-domain
Out-of-domain
Overall
AudioBox∗
58.70
50.00
44.00
50.43
AudioBox CE
61.52
49.58
42.09
50.51
AudioBox CU
62.39
48.31
38.61
49.02
AudioBox PC
45.22
54.24
53.74
50.75
AudioBox PQ
64.13
45.34
36.17
47.99
NISQA
53.48
51.69
43.48
48.62
DNSMOS
47.17
52.12
40.52
45.08
Figure 7: Distribution of the source videos across the six categories in VAPref-10K.
Table 6: Human preference accuracy (%) of text-video and text-audio consistency metrics.
Metric
Easy
In-domain
Out-of-domain
Overall
Text-video consistency
CLIP Score∗
55.56
54.66
53.58
54.50
ImageBind T-V
47.17
53.39
50.96
50.04
ViCLIP
53.91
43.83
49.74
50.16
Text-audio consistency
ImageBind T-A∗
58.48
52.97
51.48
54.29
CLAP Score
51.25
51.96
66.11
56.70
Table 7: Human preference accuracy (%) of audio-video semantic-consistency metrics. Audio-Video Alignment, AVHScore, and ImageBind A-V are all ImageBind-based metrics for audio-video semantic consistency, but differ in their preprocessing and feature aggregation procedures.
Metric
Easy
In-domain
Out-of-domain
Overall
Audio-Video Alignment∗
61.30
55.51
51.83
55.94
AVH Score
61.74
58.90
54.26
57.83
ImageBind A-V
61.09
58.47
54.96
57.83
Table 8: Human preference accuracy (%) of synchronization and aggregate audio-video metrics. Values in parentheses are score coverage (%).
Metric
Easy
In-domain
Out-of-domain
Overall
SyncNet Confidence
64.17 (40.7)
50.60 (35.2)
40.19 (18.6)
54.38 (29.7)
SyncNet Min Distance
54.55 (40.7)
59.04 (35.2)
47.17 (18.6)
53.46 (29.7)
SyncNet |Offset| (frames)
65.06 (40.7)
47.89 (35.2)
41.86 (18.6)
55.11 (29.7)
SynchFormer |ArgmaxOffset| (s)
63.82
47.13
50.56
55.22
SynchFormer |Offset| (s)∗
61.30
46.61
48.35
52.71
Desync
66.34
48.95
51.51
56.89
Lip-sync Confidence
63.10 (54.8)
53.57 (47.5)
42.98 (21.0)
55.88 (38.2)
Javis Score∗
60.00
55.51
54.96
56.88
왜 중요한가
AI가 영상과 소리를 동시에 만드는 시대에 '얼마나 자연스럽고 사람 마음에 드는가'를 정확히 재는 채점 기준이 없으면 아무리 강화학습을 돌려도 엉뚱한 방향으로 품질이 왜곡될 수 있다. 이 연구는 사람의 실제 판단을 반영한 채점 모델과 그 근거가 되는 비교 데이터셋, 평가 기준(벤치마크)까지 함께 공개해 앞으로 영상-음성 생성 AI를 사람 취향에 맞게 개선하는 표준적인 방법을 제시한다.
이 논문의 용어
강화학습(RL) · AI가 시행착오를 통해 더 높은 점수(보상)를 받는 방향으로 스스로 개선하도록 학습시키는 방법
보상 신호(reward) · 강화학습에서 AI의 결과물이 얼마나 좋은지 알려주는 점수
리워드 해킹 · AI가 채점 기준의 허점을 이용해 실제 품질과 상관없이 점수만 높이는 현상
체인 오브 소트(CoT) · AI가 최종 답을 내기 전 단계적으로 이유를 설명하며 추론하는 방식
GRPO · 여러 개의 후보 답변을 그룹으로 비교해 상대적으로 더 나은 것에 보상을 주는 강화학습 기법
논문 원문 초록 (영문)
Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing methods construct this reward by combining metrics for individual quality dimensions, including audio quality, visual fidelity, and synchronization. However, these metrics evaluate perceptual dimensions separately and fail to capture the overall semantic and temporal coherence among the text prompt, video, and audio that shapes human preferences. Optimizing models against these metrics encourages reward hacking, generating video-audio content that achieves high scores on these metrics yet appears incoherent or unfaithful to human viewers. To address this problem, we first construct a large-scale human-preference dataset VAPref-10K for joint video-audio generation, comprising 9K prompts and 10.3K fine-grained paired comparisons from open-source generation models. We also introduce the VA-Judger-Bench benchmark with both in-domain and out-of-domain model comparisons to evaluate whether reward models truly align with human preferences. We further propose VA-Judger, a chain-of-thought omni-reward model for joint video-audio generation. In particular, VA-Judger first learns from pairs with clear quality gaps to establish structured output and coarse preference discrimination, then distills reliable preference explanations for harder near-quality comparisons via rejection sampling verified against human annotations, and finally performs dimension-wise reinforcement learning that decomposes human feedback into individual quality dimensions for denser reward signals than a single binary preference label. Experiments show that VA-Judger outperforms metric baselines in predicting human preferences on both in-domain and out-of-domain evaluations. Using its human-aligned rewards for post-training audio-video generation model also yields significant improvements in generation quality.