METAL LAB

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling

arXiv:2608.218392026-08-21

A checklist-based data pipeline teaches AI reward models to actually check facts in a video before scoring it, instead of just guessing a number

FIRM-Video builds training data for video-scoring AI by first breaking each video into checkable yes/no questions about instructions, real-world plausibility, and visual quality, then verifying each one against the footage before computing a score. This produced a 88,044-instance dataset (FIRM-Video-90K) and a 750-annotation human benchmark (FIRM-Video-Bench). The resulting 8-billion-parameter model, FIRM-Video-8B, matched human judgments better than existing models and helped pick better videos out of multiple AI-generated candidates.

METAL LAB explanatory visual

How FIRM-Video builds trustworthy training data before scoring a video

Evidence statusMeasured results reported

  1. 1. Checklist constructionFor each video-prompt pair, three separate checklists are built: atomic prompt requirements for Instruction Following, entity/action-based checks for World Coherence, and a fixed list of common visual defects for Perceptual Quality.
  2. 2. Evidence verificationA multimodal evaluator checks each checklist item against the actual video frames, marking it satisfied or violated with brief visual evidence; unverifiable items count as unsatisfied.
  3. 3. Score aggregationOnly verified checklist answers, weighted by importance, are combined into a single numeric score per dimension, avoiding double-counting of the same issue across dimensions.
  4. 4. Dataset + benchmarkRunning this pipeline on 29,348 videos produced 88,044 training instances (FIRM-Video-90K); 250 separate videos were scored by humans to form the 750-annotation test set FIRM-Video-Bench.
  5. 5. End-to-end reward modelQwen3-VL-8B is fine-tuned on this data to become FIRM-Video-8B, which predicts a score and explanation directly from a prompt and video frames in one pass, without running the checklist pipeline at inference time.
An explanatory diagram made by METAL LAB, not a figure supplied by the paper's authors.

What they did

  1. Problem addressed: existing AI judges for text-to-video models tend to look only at obvious features, invent justifications after deciding a score, and mix up different types of errors (e.g. treating a physics glitch as if it were a visual-quality problem).
  2. Method: for each video-prompt pair, FIRM-Video first generates a checklist specific to three aspects (does it follow the instructions, is it physically/logically coherent, does it look visually clean), verifies each checklist item against the actual video frames, then only aggregates the verified answers into a final score and a written explanation.
  3. Data built: this pipeline was run offline on about 29,348 AI-generated videos to produce FIRM-Video-90K (88,044 labeled training instances), and 250 separate videos were scored by human experts to create FIRM-Video-Bench (750 annotations) for testing.
  4. Result: after fine-tuning Qwen3-VL-8B on this data, the resulting FIRM-Video-8B model had the lowest average scoring error (MAE 0.78, down from 1.33 for the un-tuned base model) among all tested models, including several proprietary systems, on FIRM-Video-Bench.
  5. Result: when used to pick the best video out of 8 AI-generated candidates ('Best-of-8'), FIRM-Video-8B consistently picked higher-quality videos than competing selection methods across three different video generators, as measured by the standard VBench metric.
Figure 1: Qualitative evaluation case of FIRM-Video-8B, illustrating its improved alignment with human judgments.
Figure 1: Qualitative evaluation case of FIRM-Video-8B, illustrating its improved alignment with human judgments.
Table 1: Performance comparison on FIRM-Video-Bench across three evaluation dimensions: Instruction Following (IF), World Coherence (WC), and Perceptual Quality (PQ). The two values under Acc./Relaxed Acc. denote strict and relaxed accuracy, respectively. The row marked with † reports the performance of direct data construction pipeline and is excluded from ranking.
ModelInstruction FollowingWorld CoherencePerceptual QualityOverall
MAE↓STD↓Acc./Relaxed Acc.↑SRCC↑MAE↓STD↓Acc./Relaxed Acc.↑SRCC↑MAE↓STD↓Acc./Relaxed Acc.↑SRCC↑MAE↓STD↓Acc./Relaxed Acc.↑
Closed-source models
GPT-50.620.710.50/0.880.741.661.190.20/0.480.490.950.860.34/0.760.511.081.030.35/0.71
Gemini-3.1-Pro0.770.760.40/0.860.711.271.080.28/0.620.541.441.120.21/0.590.411.161.040.30/0.69
Doubao-Seed-2.0-Lite0.800.830.41/0.840.671.561.240.24/0.530.440.800.750.38/0.840.551.051.030.35/0.73
Open-source models
InternVL3-8B1.260.950.24/0.600.572.101.320.15/0.360.221.291.000.24/0.610.221.561.170.21/0.52
InternVL3-38B0.820.830.40/0.820.652.111.320.15/0.350.151.331.080.27/0.580.201.421.210.27/0.58
Qwen3-VL-8B0.930.890.34/0.800.601.791.350.22/0.460.291.281.040.28/0.600.341.331.160.28/0.62
Qwen3-VL-30B-A3B0.950.870.32/0.800.621.901.310.19/0.400.221.361.090.28/0.550.371.401.170.27/0.58
Qwen3-VL-235B-A22B0.930.870.33/0.800.651.631.300.24/0.520.311.271.040.28/0.600.401.271.120.29/0.64
Our methods
FIRM-Video Data Pipeline†0.680.680.43/0.900.690.800.830.41/0.840.630.740.700.40/0.860.610.730.740.41/0.87
FIRM-Video-8B (InternVL3-8B)0.690.730.45/0.870.670.960.950.37/0.760.490.900.770.32/0.800.350.850.830.38/0.81
FIRM-Video-8B (Qwen3-VL-8B)0.650.770.50/0.880.690.860.880.40/0.800.530.820.740.36/0.830.510.780.800.42/0.84
Figure 2: Statistics of the FIRM-Video-90K dataset. Left: distributions across evaluation dimensions and score ranges. Right: distributions of video duration and prompt length.
Figure 2: Statistics of the FIRM-Video-90K dataset. Left: distributions across evaluation dimensions and score ranges. Right: distributions of video duration and prompt length.
Table 3: Best-of-N performance comparison of FIRM-Video-8B and other sampling strategies on VBench. The best and second-best results under each T2V generator are highlighted in bold and underlined, respectively.
Sampling StrategyTotal ScoreQuality ScoreSemantic ScoreVideo QualityVideo–Condition Consistency
Subject ConsistencyBackground ConsistencyTemporal FlickeringImaging QualityMultiple ObjectsColorSpatial RelationshipAppearance Style
T2V model: LaVie-Base
Random79.4281.3471.7792.0997.6497.6365.9237.7389.6543.3023.94
By Qwen3-VL-8B80.0381.6173.7392.6497.5697.6066.8850.9183.1147.3224.27
By InternVL3-8B80.4582.1073.8292.9497.5697.5667.0250.3085.8545.8124.10
By VideoScore279.9581.6173.3392.9597.8097.7666.5453.7386.8645.7023.87
By FIRM-Video-8B80.7282.1475.0693.6697.9198.1267.4658.9988.2051.2024.40
T2V model: CogVideoX-2B
Random78.5780.9069.2593.4796.6696.9059.7249.1688.4361.3022.95
By Qwen3-VL-8B78.4980.2271.5994.4896.8496.9059.2055.5688.4065.8023.30
By InternVL3-8B78.3580.2770.6594.5396.8997.0459.9051.3089.4961.0923.24
By VideoScore278.6780.3072.1794.6396.7896.9860.0861.0587.5963.8222.54
By FIRM-Video-8B79.6881.0374.2894.9297.2897.3360.0062.8892.3766.8623.02
T2V model: Wan2.1-T2V-1.3B
Random80.4084.1365.4994.5798.0898.9868.1857.0185.6376.1319.73
By Qwen3-VL-8B81.4384.3069.9694.5797.9098.8767.5767.9185.9474.3920.67
By InternVL3-8B81.7084.4370.7895.4698.3298.8467.7167.6186.1577.0820.30
By VideoScore281.6484.8668.7495.4598.1398.9967.5772.2683.9174.5619.98
By FIRM-Video-8B82.3684.9172.1996.5498.6899.2268.6273.4092.2381.1020.64
Figure 3: Overview of the FIRM-Video data construction framework and reward model training. Given a prompt and its generated video, FIRM-Video applies a unified check-before-score pipeline to construct fine-grained supervision for instruction following, world coherence, and perceptual quality. The resulting dimension-specific analyses and scores form FIRM-Video-90K, which is then used to train the FIRM-Video-8B reward model.
Figure 3: Overview of the FIRM-Video data construction framework and reward model training. Given a prompt and its generated video, FIRM-Video applies a unified check-before-score pipeline to construct fine-grained supervision for instruction following, world coherence, and perceptual quality. The resulting dimension-specific analyses and scores form FIRM-Video-90K, which is then used to train the FIRM-Video-8B reward model.
Table 4: Aggregation ablations on FIRM-Video-Bench. STD is computed over absolute errors; relaxed accuracy allows a one-point deviation from the expert rating. Best results within each dimension are bold.
DimensionAggregationMAE ↓STD ↓Acc. ↑Relaxed Acc. ↑SRCC ↑
IFmean0.680.700.440.880.69
Importance weighted (ours)0.680.680.430.900.69
WCmean0.860.830.370.820.61
Importance weighted (ours)0.800.830.410.840.63
Figure 4: Best-of-N scaling results on VBench using LaVie-Base as the video generation model. Performance is measured by the VBench total score as N increases.
Figure 4: Best-of-N scaling results on VBench using LaVie-Base as the video generation model. Performance is measured by the VBench total score as N increases.
Table 5: Score distributions of FIRM-Video-90K and FIRM-Video-Bench.
DatasetDimensionScore=1Score=2Score=3Score=4Score 5Total
FIRM-Video-90KIF5,104 (17.4%)8,900 (30.3%)5,551 (18.9%)3,806 (13.0%)5,987 (20.4%)29,348
WC4,795 (16.3%)11,129 (37.9%)5,649 (19.2%)2,510 (8.6%)5,265 (17.9%)29,348
PQ874 (3.0%)3,749 (12.8%)4,257 (14.5%)18,331 (62.5%)2,137 (7.3%)29,348
All10,77323,77815,45724,64713,38988,044
FIRM-Video-BenchIF28 (11.2%)73 (29.2%)63 (25.2%)59 (23.6%)27 (10.8%)250
WC48 (19.2%)76 (30.4%)48 (19.2%)44 (17.6%)34 (13.6%)250
PQ11 (4.4%)54 (21.6%)69 (27.6%)59 (23.6%)57 (22.8%)250
All87203180162118750
Figure 5: Human annotation interface for FIRM-Video-Bench. Annotators rate instruction following, perceptual quality, and world coherence on a five-point scale.
Figure 5: Human annotation interface for FIRM-Video-Bench. Annotators rate instruction following, perceptual quality, and world coherence on a five-point scale.
Table 6: Dimension mapping from VBench dimensions to FIRM-Video reward dimensions. “Quality” and “Cond. Consist.” denote VBench’s Video-Quality and Video–Condition Consistency groups, respectively.
RewardVBench DimensionVBench Group
IFDynamic DegreeQuality
Overall ConsistencyCond. Consist.
Object ClassCond. Consist.
Multiple ObjectsCond. Consist.
Human ActionCond. Consist.
ColorCond. Consist.
Spatial RelationshipCond. Consist.
SceneCond. Consist.
Appearance StyleCond. Consist.
Temporal StyleCond. Consist.
WCSubject ConsistencyQuality
Background ConsistencyQuality
Motion SmoothnessQuality
PQTemporal FlickeringQuality
Aesthetic QualityQuality
Imaging QualityQuality
Figure 6: Qualitative example of FIRM-Video-8B evaluation (1)
Figure 6: Qualitative example of FIRM-Video-8B evaluation (1)
Table 7: Best-of-8 results on the video-quality subdimensions of VBench.
T2V ModelSampling StrategySubject ConsistencyBackground ConsistencyTemporal FlickeringMotion SmoothnessDynamic DegreeAesthetic QualityImaging Quality
LaVie-BaseRandom92.0997.6497.6396.6355.5664.5365.92
Qwen3-VL-8B92.6497.5697.6096.6255.5664.9366.88
InternVL3-8B92.9497.5697.5696.6961.1164.7467.02
VideoScore292.9597.8097.7696.3956.9464.2766.54
FIRM-Video-8B93.6697.9198.1296.5856.9464.1767.46
CogVideoX-2BRandom93.4796.6696.9097.2373.6158.4759.72
Qwen3-VL-8B94.4896.8496.9097.3562.5058.3059.20
InternVL3-8B94.5396.8997.0497.2462.5057.8459.90
VideoScore294.6396.7896.9897.2759.7259.3160.08
FIRM-Video-8B94.9297.2897.3397.3465.2859.1560.00
Wan2.1-T2V-1.3BRandom94.5798.0898.9898.2162.5064.3868.18
Qwen3-VL-8B94.5797.9098.8798.1066.6764.9467.57
InternVL3-8B95.4698.3298.8498.2963.8964.8767.71
VideoScore295.4598.1398.9998.4368.0665.1067.57
FIRM-Video-8B96.5498.6899.2298.4959.7265.6768.62
Figure 7: Qualitative example of FIRM-Video-8B evaluation (2)
Figure 7: Qualitative example of FIRM-Video-8B evaluation (2)
Table 8: Best-of-8 results on the video–condition consistency subdimensions of VBench.
T2V ModelSampling StrategyObject ClassMultiple ObjectsHuman ActionColorSpatial RelationshipSceneAppearance StyleTemporal StyleOverall Consistency
LaVie-BaseRandom93.1237.7392.0089.6543.3052.3323.9424.7327.20
Qwen3-VL-8B91.6150.9196.0083.1147.3252.6924.2725.2527.75
InternVL3-8B95.8150.3096.0085.8545.8151.0224.1025.0327.46
VideoScore291.0653.7393.0086.8645.7050.8023.8725.1627.33
FIRM-Video-8B90.5158.9993.0088.2051.2052.0324.4025.1927.57
CogVideoX-2BRandom76.1149.1687.0088.4361.3039.9022.9523.7724.43
Qwen3-VL-8B83.6255.5691.0088.4065.8035.8323.3023.8425.27
InternVL3-8B82.9151.3089.0089.4961.0937.7923.2423.7425.30
VideoScore284.4161.0591.0087.5963.8240.4822.5423.5425.06
FIRM-Video-8B85.2162.8890.0092.3766.8645.1323.0224.2925.13
Wan2.1-T2V-1.3BRandom74.7657.0168.0085.6376.1324.3519.7323.7423.31
Qwen3-VL-8B82.3667.9183.0085.9474.3925.5120.6723.5324.78
InternVL3-8B84.3467.6182.0086.1577.0830.3820.3023.8624.15
VideoScore279.2772.2679.0083.9174.5623.8419.9823.7323.88
FIRM-Video-8B80.2273.4083.0092.2381.1029.1420.6423.5224.57
Figure 8: Qualitative example of FIRM-Video-8B evaluation (3)
Figure 8: Qualitative example of FIRM-Video-8B evaluation (3)

Findings

  • On FIRM-Video-Bench, FIRM-Video-8B (built on Qwen3-VL-8B) achieved the lowest overall MAE of 0.78 among all tested proprietary and open-source models, and had the lowest error specifically on the World Coherence dimension.
  • On the same benchmark, GPT-5 had the lowest error for Instruction Following (0.62) and Doubao-Seed-2.0-Lite had the lowest error for Perceptual Quality (0.80), showing FIRM-Video-8B did not lead on every single dimension but led overall.
  • In Best-of-8 sampling with VBench, FIRM-Video-8B produced the best Total, Quality, and Semantic Scores across all three tested video generators (LaVie-Base, CogVideoX-2B, Wan2.1-T2V-1.3B), beating the next-best method by 0.27 to 1.01 points on Total Score and 1.24 to 2.11 points on Semantic Score.
  • On the out-of-domain MJ-Bench-Video benchmark, one FIRM-Video-8B variant led on Alignment and Consistency & Coherence, and ranked second on Overall preference accuracy, indicating some generalization beyond the training data.
  • An ablation study showed that weighting checklist items by importance (rather than averaging them equally) improved accuracy, and for the World Coherence dimension specifically reduced the error from 0.86 to 0.80.
Figure 9: Qualitative example of FIRM-Video-8B evaluation (4)
Figure 9: Qualitative example of FIRM-Video-8B evaluation (4)

Where it can be used

  • Selecting the best output among several AI-generated videos for the same prompt (Best-of-N filtering) before showing it to a user.
  • Automated quality control for filtering large video-generation datasets by instruction-following, physical plausibility, or visual defects.
  • As a scoring signal for future reinforcement-learning or preference-optimization training of text-to-video generation models, an application the authors mention but have not yet tested.
Figure 10: Qualitative example of FIRM-Video-8B evaluation (5)
Figure 10: Qualitative example of FIRM-Video-8B evaluation (5)

Limits and open work

  • The authors explicitly state they have not yet used FIRM-Video-8B to directly train or optimize a video generator; it was only tested as an evaluator and selector, so its usefulness for reinforcement learning-based alignment remains unverified.
  • The human benchmark used for testing (FIRM-Video-Bench) covers only 250 videos and 750 annotations, a relatively small test set compared to typical benchmark sizes.
  • Training videos were drawn mainly from two existing preference datasets and over 20 generation models, so performance on video styles or generators outside this distribution is not directly demonstrated.
  • The reward model relies on 8 uniformly sampled frames per video rather than full video, which may miss very short or fine-grained temporal issues.
  • Score thresholds mapping the continuous checklist scores to 5-point ratings were set empirically, and the paper does not report how sensitive results are to this choice.

Why it matters

Video-generation systems are improving fast, but judging whether their output is actually good, faithful to the prompt, and free of physical nonsense is still mostly done by hand or by unreliable automatic scorers. A more trustworthy, checklist-verified reward model gives developers a cheaper way to filter, rank, and eventually train better video generators.

Terms in this paper

  • Reward model · An AI model trained to output a quality score for another AI's output, used for filtering, ranking, or guiding training.
  • Checklist-driven / check-before-score · An approach where specific yes/no questions are verified against evidence first, and only verified answers are combined into a final score.
  • Instruction Following / World Coherence / Perceptual Quality · The three evaluation dimensions this paper uses: does the video match the text prompt, is it physically/logically sensible, and is the image quality clean.
  • Best-of-N sampling · Generating N candidate videos for the same prompt and picking the single best one according to some scoring method.
  • MAE (Mean Absolute Error) · The average size of the gap between a model's predicted score and the human-given score; lower means closer to human judgment.

Original abstract (English)

Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference efficiency places high demands on the quality of training supervision. Existing approaches often rely on holistic judges with fixed rubrics or open-ended reasoning, leading to incomplete inspection, unfaithful justification, and entangled attribution. We introduce FIRM-Video, a unified checklist-driven data construction framework based on a check-before-score principle: construct dimension-specific checklists, verify each criterion against temporal visual evidence, and aggregate only verified decisions. For Instruction Following, FIRM-Video decomposes prompts into weighted atomic requirements; for World Coherence, it constructs prompt-calibrated, target-specific checks grounded in visible entities and actions; and for Perceptual Quality, it applies a generic taxonomy of visual defects. The verified criteria and scores are further transformed into natural-language analyses for end-to-end reward modeling. Subsequently, we construct FIRM-Video-90K with 88,044 dimension-specific instances from 29,348 videos, and introduce FIRM-Video-Bench with 750 point-wise human annotations across 250 videos. The Qwen3-VL-based FIRM-Video-8B achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.

Authors · Peiyuan Zhang

Read on arXiv

Latest papers

All papers →

Latest from METAL LAB

Figures: Peiyuan Zhang et al., arXiv:2608.21839, arxiv-nonexclusive