VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation
arXiv:2608.186072026-08-18
A judge model that scores AI-made video-and-sound clips the way humans actually prefer them
AI models that generate video and audio together need a scoring system to improve via reinforcement learning, but past approaches added up separate scores for image quality, sound quality, and sync, which often disagreed with what people actually liked. The authors built a large human-preference dataset (VAPref-10K) where annotators compared pairs of generated clips and explained their choice, then trained a reasoning-based judge model called VA-Judger through three stages. Using VA-Judger as the reward to retrain a video-audio generator (LTX-2) led to outputs that a 20-person human panel preferred far more often than existing methods.
What they did
Prior methods combined separate scores for video quality, audio quality, and synchronization into a single reward, but these scores poorly matched real human judgment, letting models exploit the scoring gaps rather than genuinely improve (reward hacking)
The team built VAPref-10K, a dataset of about 9,000 text prompts drawn from real YouTube, Bilibili, film, and TV clips, paired with 10.3K human comparisons of outputs from open-source video-audio generation models, where annotators picked the better clip and the specific quality reasons behind it
VA-Judger is trained in three stages: first on easy pairs with obvious quality gaps to learn a structured scoring format, then on harder near-tie pairs filtered to match human-verified answers, and finally with reinforcement learning (GRPO) that rewards correct dimension-by-dimension scoring, not just the final pick
Standalone automatic metrics only matched human preference 50-57% of the time, while VA-Judger beat the best of them by over 10 percentage points, staying accurate even on outputs from closed-source models it never saw during training
When VA-Judger's scores were used to retrain the LTX-2 video-audio generator, 20 human participants preferred the VA-Judger-trained outputs 62.30% of the time, versus 27.63% for the prior method (OmniNFT) and just 10.08% for the untrained base model
Figure 1: Evaluation of VA-Judger. (a) VA-Judger selects the preferred clip through structured dimension-wise reasoning, whereas separate evaluation metrics disagree with human judgment. (b) VA-Judger achieves the highest agreement with human preferences. (c) VA-Judger provides the reward for post-training LTX-2, improving the quality over the base model and OmniNFT.
Table 1: Accuracy (%) against human pairwise preferences. Single-dimension evaluation metrics and reward models are both evaluated on the 1,150-pair VA-Judger-Bench and report Total Acc / Parsed Acc. Unparseable responses count as incorrect in Total Acc.
Model
Easy
In-domain
Out-of-domain
Overall
Single-dimension evaluation metrics
Video quality: VideoAlign
62.39
52.97
48.35
54.29
Audio quality: AudioBox
58.70
50.00
44.00
50.43
Text-video: CLIP Score
55.56
54.66
53.58
54.50
Text-audio: ImageBind T-A
58.48
52.97
51.48
54.29
Audio-video: ImageBind
61.30
55.51
51.83
55.94
Synchronization: SynchFormer
61.30
46.61
48.35
52.71
Overall: Javis Score
60.00
55.51
54.96
56.88
Reward models
Qwen3-Omni Captioner (No CoT)
57.00 / 57.00
54.00 / 54.00
56.20 / 56.20
57.22 / 57.22
Qwen3-Omni Captioner (CoT)
52.50 / 57.65
49.60 / 54.87
46.60 / 56.97
50.43 / 58.61
Qwen3-Omni Instruct NoCoT
60.75 / 60.75
54.00 / 54.00
54.60 / 54.60
56.61 / 56.61
Qwen3-Omni Instruct CoT
63.25 / 64.71
54.80 / 55.92
55.00 / 55.33
57.83 / 58.69
+ Easy Cold Start
72.00 / 72.00
59.20 / 59.20
56.20 / 56.20
62.35 / 62.35
+ Hard SFT
74.50 / 74.50
63.60 / 63.60
60.20 / 60.20
65.91 / 65.91
+ GRPO (VA-Judger)
76.25 / 76.25
66.00 / 66.00
63.40 / 63.40
68.43 / 68.43
Figure 2: Overview of the VA-Judger training pipeline. Stage 1 cold-starts the model on easy preference pairs. Stage 2 aligns the model on harder pairs through human rejection sampling. Stage 3 performs GRPO using answer-level and dimension-level rewards grounded in human reasons.
Table 2: Results on the randomly sampled 200-prompt JavisBench subset. Best results are highlighted in green and second-best results are underlined. Higher is better unless marked with ↓. The two Δ rows report the raw change produced by VA-Judger post-training.
Video Quality
AudioBox Quality
Cross-Modal Alignment
Model
VQ ↑
MQ ↑
AQ ↑
AB-CE ↑
AB-CU ↑
AB-PC ↑
AB-PQ ↑
TV-Align ↑
TA-Align ↑
ViCLIP ↑
LTX-2
2.248
0.697
4.767
3.957
5.904
3.041
6.166
0.303
0.105
0.232
LTX-2 + OmniNFT
3.727
0.947
5.399
4.745
6.797
3.150
6.905
0.279
0.133
0.219
LTX-2 + VA-Judger
3.942
1.183
5.610
5.136
6.766
3.606
6.932
0.310
0.180
0.242
Δ vs. LTX-2
+1.693
+0.486
+0.843
+1.179
+0.862
+0.564
+0.766
+0.007
+0.075
+0.009
Δ vs. OmniNFT
+0.214
+0.236
+0.211
+0.391
-0.031
+0.456
+0.026
+0.031
+0.047
+0.023
Figure 3: Qualitative comparisons of LTX-2, OmniNFT, and our VA-Judger post-trained model. VA-Judger better follows the requested temporal actions in both examples.
Table 3: Category distributions of the full JavisBench pool and the random evaluation subset. Each cell reports the number and percentage of examples.
Dimension
Category
Full pool (10,140)
Random subset (200)
Event Scenario
Living Scenario
4,029 (39.7%)
73 (36.5%)
Natural Scenario
2,452 (24.2%)
37 (18.5%)
Urban Scenario
2,140 (21.1%)
66 (33.0%)
Industrial Scenario
1,090 (10.7%)
20 (10.0%)
Virtual Scenario
429 (4.2%)
4 (2.0%)
Visual Style
Camera Shooting
8,586 (84.7%)
174 (87.0%)
2D Animation
915 (9.0%)
17 (8.5%)
3D Animation
639 (6.3%)
9 (4.5%)
Sound Type
Musical Sounds
5,065 (50.0%)
73 (36.5%)
Ambient Sounds
4,183 (41.3%)
126 (63.0%)
Biological Sounds
2,215 (21.8%)
59 (29.5%)
Mechanical Sounds
1,940 (19.1%)
49 (24.5%)
Speech Sounds
1,090 (10.7%)
29 (14.5%)
Spatial Composition
Multiple Subjects
7,588 (74.8%)
167 (83.5%)
Single Subject
2,502 (24.7%)
33 (16.5%)
Off-screen Sound
2,165 (21.4%)
59 (29.5%)
Temporal Composition
Sequential Events
3,698 (36.5%)
46 (23.0%)
Simultaneous Events
3,306 (32.6%)
99 (49.5%)
Single Event
3,136 (30.9%)
55 (27.5%)
Figure 4: Human preference rates over 200 three-way comparisons. For each prompt, participants select the best video-audio output among LTX-2, LTX-2 post-trained with OmniNFT, and LTX-2 post-trained with VA-Judger.
Table 4: Human preference accuracy (%) of video-quality metrics.
Metric
Easy
In-domain
Out-of-domain
Overall
VideoAlign Overall∗
62.39
52.97
48.35
54.29
VideoPhy2
52.86
54.39
50.22
51.66
Motion Quality
59.91
51.06
49.74
53.66
Visual Quality
62.42
52.34
42.93
51.70
HPSv3
56.30
52.12
50.61
52.95
Video Aesthetic
57.52
50.21
47.21
51.50
Video Technical/Aesthetic
57.39
51.69
42.78
49.72
Video Technical/Overall
59.13
52.97
42.26
50.35
Video Technical/Technical
54.13
54.24
47.13
50.98
Figure 6: Overview of VAPref-10K construction and the raw data distribution. Real video-audio clips are captioned by Qwen3.5-Omni, rewritten into text prompts by an LLM, and passed to LTX-2, OVI, and DaVinci-MagiHuman to generate candidate clips.
Table 5: Human preference accuracy (%) of audio-quality metrics.
Metric
Easy
In-domain
Out-of-domain
Overall
AudioBox∗
58.70
50.00
44.00
50.43
AudioBox CE
61.52
49.58
42.09
50.51
AudioBox CU
62.39
48.31
38.61
49.02
AudioBox PC
45.22
54.24
53.74
50.75
AudioBox PQ
64.13
45.34
36.17
47.99
NISQA
53.48
51.69
43.48
48.62
DNSMOS
47.17
52.12
40.52
45.08
Figure 7: Distribution of the source videos across the six categories in VAPref-10K.
Table 6: Human preference accuracy (%) of text-video and text-audio consistency metrics.
Metric
Easy
In-domain
Out-of-domain
Overall
Text-video consistency
CLIP Score∗
55.56
54.66
53.58
54.50
ImageBind T-V
47.17
53.39
50.96
50.04
ViCLIP
53.91
43.83
49.74
50.16
Text-audio consistency
ImageBind T-A∗
58.48
52.97
51.48
54.29
CLAP Score
51.25
51.96
66.11
56.70
Table 7: Human preference accuracy (%) of audio-video semantic-consistency metrics. Audio-Video Alignment, AVHScore, and ImageBind A-V are all ImageBind-based metrics for audio-video semantic consistency, but differ in their preprocessing and feature aggregation procedures.
Metric
Easy
In-domain
Out-of-domain
Overall
Audio-Video Alignment∗
61.30
55.51
51.83
55.94
AVH Score
61.74
58.90
54.26
57.83
ImageBind A-V
61.09
58.47
54.96
57.83
Table 8: Human preference accuracy (%) of synchronization and aggregate audio-video metrics. Values in parentheses are score coverage (%).
Metric
Easy
In-domain
Out-of-domain
Overall
SyncNet Confidence
64.17 (40.7)
50.60 (35.2)
40.19 (18.6)
54.38 (29.7)
SyncNet Min Distance
54.55 (40.7)
59.04 (35.2)
47.17 (18.6)
53.46 (29.7)
SyncNet |Offset| (frames)
65.06 (40.7)
47.89 (35.2)
41.86 (18.6)
55.11 (29.7)
SynchFormer |ArgmaxOffset| (s)
63.82
47.13
50.56
55.22
SynchFormer |Offset| (s)∗
61.30
46.61
48.35
52.71
Desync
66.34
48.95
51.51
56.89
Lip-sync Confidence
63.10 (54.8)
53.57 (47.5)
42.98 (21.0)
55.88 (38.2)
Javis Score∗
60.00
55.51
54.96
56.88
Why it matters
As AI systems increasingly generate video and audio together, having a scoring method that truly reflects human taste is essential, otherwise reinforcement learning can push models toward metrics rather than real quality. This work offers a reusable judge model, dataset, and benchmark that others can use to align future video-audio generators with genuine human preference.
Terms in this paper
reinforcement learning (RL) · a training method where an AI improves by trial and error to earn higher scores, called rewards
reward signal · the score given to an AI's output during reinforcement learning to indicate how good it is
reward hacking · when an AI finds a shortcut to score well on a metric without actually improving real quality
chain-of-thought (CoT) · a method where the AI explains its reasoning step by step before giving a final answer
GRPO · a reinforcement learning technique that compares a group of candidate answers and rewards the relatively better ones
Original abstract (English)
Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing methods construct this reward by combining metrics for individual quality dimensions, including audio quality, visual fidelity, and synchronization. However, these metrics evaluate perceptual dimensions separately and fail to capture the overall semantic and temporal coherence among the text prompt, video, and audio that shapes human preferences. Optimizing models against these metrics encourages reward hacking, generating video-audio content that achieves high scores on these metrics yet appears incoherent or unfaithful to human viewers. To address this problem, we first construct a large-scale human-preference dataset VAPref-10K for joint video-audio generation, comprising 9K prompts and 10.3K fine-grained paired comparisons from open-source generation models. We also introduce the VA-Judger-Bench benchmark with both in-domain and out-of-domain model comparisons to evaluate whether reward models truly align with human preferences. We further propose VA-Judger, a chain-of-thought omni-reward model for joint video-audio generation. In particular, VA-Judger first learns from pairs with clear quality gaps to establish structured output and coarse preference discrimination, then distills reliable preference explanations for harder near-quality comparisons via rejection sampling verified against human annotations, and finally performs dimension-wise reinforcement learning that decomposes human feedback into individual quality dimensions for denser reward signals than a single binary preference label. Experiments show that VA-Judger outperforms metric baselines in predicting human preferences on both in-domain and out-of-domain evaluations. Using its human-aligned rewards for post-training audio-video generation model also yields significant improvements in generation quality.