Figure 1: Evaluation of VA-Judger. (a) VA-Judger selects the preferred clip through structured dimension-wise reasoning, whereas separate evaluation metrics disagree with human judgment. (b) VA-Judger achieves the highest agreement with human preferences. (c) VA-Judger provides the reward for post-training LTX-2, improving the quality over the base model and OmniNFT.
Table 1: Accuracy (%) against human pairwise preferences. Single-dimension evaluation metrics and reward models are both evaluated on the 1,150-pair VA-Judger-Bench and report Total Acc / Parsed Acc. Unparseable responses count as incorrect in Total Acc.
Model
Easy
In-domain
Out-of-domain
Overall
Single-dimension evaluation metrics
Video quality: VideoAlign
62.39
52.97
48.35
54.29
Audio quality: AudioBox
58.70
50.00
44.00
50.43
Text-video: CLIP Score
55.56
54.66
53.58
54.50
Text-audio: ImageBind T-A
58.48
52.97
51.48
54.29
Audio-video: ImageBind
61.30
55.51
51.83
55.94
Synchronization: SynchFormer
61.30
46.61
48.35
52.71
Overall: Javis Score
60.00
55.51
54.96
56.88
Reward models
Qwen3-Omni Captioner (No CoT)
57.00 / 57.00
54.00 / 54.00
56.20 / 56.20
57.22 / 57.22
Qwen3-Omni Captioner (CoT)
52.50 / 57.65
49.60 / 54.87
46.60 / 56.97
50.43 / 58.61
Qwen3-Omni Instruct NoCoT
60.75 / 60.75
54.00 / 54.00
54.60 / 54.60
56.61 / 56.61
Qwen3-Omni Instruct CoT
63.25 / 64.71
54.80 / 55.92
55.00 / 55.33
57.83 / 58.69
+ Easy Cold Start
72.00 / 72.00
59.20 / 59.20
56.20 / 56.20
62.35 / 62.35
+ Hard SFT
74.50 / 74.50
63.60 / 63.60
60.20 / 60.20
65.91 / 65.91
+ GRPO (VA-Judger)
76.25 / 76.25
66.00 / 66.00
63.40 / 63.40
68.43 / 68.43
Figure 2: Overview of the VA-Judger training pipeline. Stage 1 cold-starts the model on easy preference pairs. Stage 2 aligns the model on harder pairs through human rejection sampling. Stage 3 performs GRPO using answer-level and dimension-level rewards grounded in human reasons.
Table 2: Results on the randomly sampled 200-prompt JavisBench subset. Best results are highlighted in green and second-best results are underlined. Higher is better unless marked with ↓. The two Δ rows report the raw change produced by VA-Judger post-training.
Video Quality
AudioBox Quality
Cross-Modal Alignment
Model
VQ ↑
MQ ↑
AQ ↑
AB-CE ↑
AB-CU ↑
AB-PC ↑
AB-PQ ↑
TV-Align ↑
TA-Align ↑
ViCLIP ↑
LTX-2
2.248
0.697
4.767
3.957
5.904
3.041
6.166
0.303
0.105
0.232
LTX-2 + OmniNFT
3.727
0.947
5.399
4.745
6.797
3.150
6.905
0.279
0.133
0.219
LTX-2 + VA-Judger
3.942
1.183
5.610
5.136
6.766
3.606
6.932
0.310
0.180
0.242
Δ vs. LTX-2
+1.693
+0.486
+0.843
+1.179
+0.862
+0.564
+0.766
+0.007
+0.075
+0.009
Δ vs. OmniNFT
+0.214
+0.236
+0.211
+0.391
-0.031
+0.456
+0.026
+0.031
+0.047
+0.023
Figure 3: Qualitative comparisons of LTX-2, OmniNFT, and our VA-Judger post-trained model. VA-Judger better follows the requested temporal actions in both examples.
Table 3: Category distributions of the full JavisBench pool and the random evaluation subset. Each cell reports the number and percentage of examples.
Dimension
Category
Full pool (10,140)
Random subset (200)
Event Scenario
Living Scenario
4,029 (39.7%)
73 (36.5%)
Natural Scenario
2,452 (24.2%)
37 (18.5%)
Urban Scenario
2,140 (21.1%)
66 (33.0%)
Industrial Scenario
1,090 (10.7%)
20 (10.0%)
Virtual Scenario
429 (4.2%)
4 (2.0%)
Visual Style
Camera Shooting
8,586 (84.7%)
174 (87.0%)
2D Animation
915 (9.0%)
17 (8.5%)
3D Animation
639 (6.3%)
9 (4.5%)
Sound Type
Musical Sounds
5,065 (50.0%)
73 (36.5%)
Ambient Sounds
4,183 (41.3%)
126 (63.0%)
Biological Sounds
2,215 (21.8%)
59 (29.5%)
Mechanical Sounds
1,940 (19.1%)
49 (24.5%)
Speech Sounds
1,090 (10.7%)
29 (14.5%)
Spatial Composition
Multiple Subjects
7,588 (74.8%)
167 (83.5%)
Single Subject
2,502 (24.7%)
33 (16.5%)
Off-screen Sound
2,165 (21.4%)
59 (29.5%)
Temporal Composition
Sequential Events
3,698 (36.5%)
46 (23.0%)
Simultaneous Events
3,306 (32.6%)
99 (49.5%)
Single Event
3,136 (30.9%)
55 (27.5%)
Figure 4: Human preference rates over 200 three-way comparisons. For each prompt, participants select the best video-audio output among LTX-2, LTX-2 post-trained with OmniNFT, and LTX-2 post-trained with VA-Judger.
Table 4: Human preference accuracy (%) of video-quality metrics.
Metric
Easy
In-domain
Out-of-domain
Overall
VideoAlign Overall∗
62.39
52.97
48.35
54.29
VideoPhy2
52.86
54.39
50.22
51.66
Motion Quality
59.91
51.06
49.74
53.66
Visual Quality
62.42
52.34
42.93
51.70
HPSv3
56.30
52.12
50.61
52.95
Video Aesthetic
57.52
50.21
47.21
51.50
Video Technical/Aesthetic
57.39
51.69
42.78
49.72
Video Technical/Overall
59.13
52.97
42.26
50.35
Video Technical/Technical
54.13
54.24
47.13
50.98
Figure 6: Overview of VAPref-10K construction and the raw data distribution. Real video-audio clips are captioned by Qwen3.5-Omni, rewritten into text prompts by an LLM, and passed to LTX-2, OVI, and DaVinci-MagiHuman to generate candidate clips.
Table 5: Human preference accuracy (%) of audio-quality metrics.
Metric
Easy
In-domain
Out-of-domain
Overall
AudioBox∗
58.70
50.00
44.00
50.43
AudioBox CE
61.52
49.58
42.09
50.51
AudioBox CU
62.39
48.31
38.61
49.02
AudioBox PC
45.22
54.24
53.74
50.75
AudioBox PQ
64.13
45.34
36.17
47.99
NISQA
53.48
51.69
43.48
48.62
DNSMOS
47.17
52.12
40.52
45.08
Figure 7: Distribution of the source videos across the six categories in VAPref-10K.
Table 6: Human preference accuracy (%) of text-video and text-audio consistency metrics.
Metric
Easy
In-domain
Out-of-domain
Overall
Text-video consistency
CLIP Score∗
55.56
54.66
53.58
54.50
ImageBind T-V
47.17
53.39
50.96
50.04
ViCLIP
53.91
43.83
49.74
50.16
Text-audio consistency
ImageBind T-A∗
58.48
52.97
51.48
54.29
CLAP Score
51.25
51.96
66.11
56.70
Table 7: Human preference accuracy (%) of audio-video semantic-consistency metrics. Audio-Video Alignment, AVHScore, and ImageBind A-V are all ImageBind-based metrics for audio-video semantic consistency, but differ in their preprocessing and feature aggregation procedures.
Metric
Easy
In-domain
Out-of-domain
Overall
Audio-Video Alignment∗
61.30
55.51
51.83
55.94
AVH Score
61.74
58.90
54.26
57.83
ImageBind A-V
61.09
58.47
54.96
57.83
Table 8: Human preference accuracy (%) of synchronization and aggregate audio-video metrics. Values in parentheses are score coverage (%).
Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing methods construct this reward by combining metrics for individual quality dimensions, including audio quality, visual fidelity, and synchronization. However, these metrics evaluate perceptual dimensions separately and fail to capture the overall semantic and temporal coherence among the text prompt, video, and audio that shapes human preferences. Optimizing models against these metrics encourages reward hacking, generating video-audio content that achieves high scores on these metrics yet appears incoherent or unfaithful to human viewers. To address this problem, we first construct a large-scale human-preference dataset VAPref-10K for joint video-audio generation, comprising 9K prompts and 10.3K fine-grained paired comparisons from open-source generation models. We also introduce the VA-Judger-Bench benchmark with both in-domain and out-of-domain model comparisons to evaluate whether reward models truly align with human preferences. We further propose VA-Judger, a chain-of-thought omni-reward model for joint video-audio generation. In particular, VA-Judger first learns from pairs with clear quality gaps to establish structured output and coarse preference discrimination, then distills reliable preference explanations for harder near-quality comparisons via rejection sampling verified against human annotations, and finally performs dimension-wise reinforcement learning that decomposes human feedback into individual quality dimensions for denser reward signals than a single binary preference label. Experiments show that VA-Judger outperforms metric baselines in predicting human preferences on both in-domain and out-of-domain evaluations. Using its human-aligned rewards for post-training audio-video generation model also yields significant improvements in generation quality.