FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling
A checklist-based data pipeline teaches AI reward models to actually check facts in a video before scoring it, instead of just guessing a number
FIRM-Video builds training data for video-scoring AI by first breaking each video into checkable yes/no questions about instructions, real-world plausibility, and visual quality, then verifying each one against the footage before computing a score. This produced a 88,044-instance dataset (FIRM-Video-90K) and a 750-annotation human benchmark (FIRM-Video-Bench). The resulting 8-billion-parameter model, FIRM-Video-8B, matched human judgments better than existing models and helped pick better videos out of multiple AI-generated candidates.
METAL LAB explanatory visual
How FIRM-Video builds trustworthy training data before scoring a video
Evidence statusMeasured results reported
- 1. Checklist constructionFor each video-prompt pair, three separate checklists are built: atomic prompt requirements for Instruction Following, entity/action-based checks for World Coherence, and a fixed list of common visual defects for Perceptual Quality.
- 2. Evidence verificationA multimodal evaluator checks each checklist item against the actual video frames, marking it satisfied or violated with brief visual evidence; unverifiable items count as unsatisfied.
- 3. Score aggregationOnly verified checklist answers, weighted by importance, are combined into a single numeric score per dimension, avoiding double-counting of the same issue across dimensions.
- 4. Dataset + benchmarkRunning this pipeline on 29,348 videos produced 88,044 training instances (FIRM-Video-90K); 250 separate videos were scored by humans to form the 750-annotation test set FIRM-Video-Bench.
- 5. End-to-end reward modelQwen3-VL-8B is fine-tuned on this data to become FIRM-Video-8B, which predicts a score and explanation directly from a prompt and video frames in one pass, without running the checklist pipeline at inference time.
What they did
- Problem addressed: existing AI judges for text-to-video models tend to look only at obvious features, invent justifications after deciding a score, and mix up different types of errors (e.g. treating a physics glitch as if it were a visual-quality problem).
- Method: for each video-prompt pair, FIRM-Video first generates a checklist specific to three aspects (does it follow the instructions, is it physically/logically coherent, does it look visually clean), verifies each checklist item against the actual video frames, then only aggregates the verified answers into a final score and a written explanation.
- Data built: this pipeline was run offline on about 29,348 AI-generated videos to produce FIRM-Video-90K (88,044 labeled training instances), and 250 separate videos were scored by human experts to create FIRM-Video-Bench (750 annotations) for testing.
- Result: after fine-tuning Qwen3-VL-8B on this data, the resulting FIRM-Video-8B model had the lowest average scoring error (MAE 0.78, down from 1.33 for the un-tuned base model) among all tested models, including several proprietary systems, on FIRM-Video-Bench.
- Result: when used to pick the best video out of 8 AI-generated candidates ('Best-of-8'), FIRM-Video-8B consistently picked higher-quality videos than competing selection methods across three different video generators, as measured by the standard VBench metric.

| Model | Instruction Following | World Coherence | Perceptual Quality | Overall | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MAE↓ | STD↓ | Acc./Relaxed Acc.↑ | SRCC↑ | MAE↓ | STD↓ | Acc./Relaxed Acc.↑ | SRCC↑ | MAE↓ | STD↓ | Acc./Relaxed Acc.↑ | SRCC↑ | MAE↓ | STD↓ | Acc./Relaxed Acc.↑ | |
| Closed-source models | |||||||||||||||
| GPT-5 | 0.62 | 0.71 | 0.50/0.88 | 0.74 | 1.66 | 1.19 | 0.20/0.48 | 0.49 | 0.95 | 0.86 | 0.34/0.76 | 0.51 | 1.08 | 1.03 | 0.35/0.71 |
| Gemini-3.1-Pro | 0.77 | 0.76 | 0.40/0.86 | 0.71 | 1.27 | 1.08 | 0.28/0.62 | 0.54 | 1.44 | 1.12 | 0.21/0.59 | 0.41 | 1.16 | 1.04 | 0.30/0.69 |
| Doubao-Seed-2.0-Lite | 0.80 | 0.83 | 0.41/0.84 | 0.67 | 1.56 | 1.24 | 0.24/0.53 | 0.44 | 0.80 | 0.75 | 0.38/0.84 | 0.55 | 1.05 | 1.03 | 0.35/0.73 |
| Open-source models | |||||||||||||||
| InternVL3-8B | 1.26 | 0.95 | 0.24/0.60 | 0.57 | 2.10 | 1.32 | 0.15/0.36 | 0.22 | 1.29 | 1.00 | 0.24/0.61 | 0.22 | 1.56 | 1.17 | 0.21/0.52 |
| InternVL3-38B | 0.82 | 0.83 | 0.40/0.82 | 0.65 | 2.11 | 1.32 | 0.15/0.35 | 0.15 | 1.33 | 1.08 | 0.27/0.58 | 0.20 | 1.42 | 1.21 | 0.27/0.58 |
| Qwen3-VL-8B | 0.93 | 0.89 | 0.34/0.80 | 0.60 | 1.79 | 1.35 | 0.22/0.46 | 0.29 | 1.28 | 1.04 | 0.28/0.60 | 0.34 | 1.33 | 1.16 | 0.28/0.62 |
| Qwen3-VL-30B-A3B | 0.95 | 0.87 | 0.32/0.80 | 0.62 | 1.90 | 1.31 | 0.19/0.40 | 0.22 | 1.36 | 1.09 | 0.28/0.55 | 0.37 | 1.40 | 1.17 | 0.27/0.58 |
| Qwen3-VL-235B-A22B | 0.93 | 0.87 | 0.33/0.80 | 0.65 | 1.63 | 1.30 | 0.24/0.52 | 0.31 | 1.27 | 1.04 | 0.28/0.60 | 0.40 | 1.27 | 1.12 | 0.29/0.64 |
| Our methods | |||||||||||||||
| FIRM-Video Data Pipeline† | 0.68 | 0.68 | 0.43/0.90 | 0.69 | 0.80 | 0.83 | 0.41/0.84 | 0.63 | 0.74 | 0.70 | 0.40/0.86 | 0.61 | 0.73 | 0.74 | 0.41/0.87 |
| FIRM-Video-8B (InternVL3-8B) | 0.69 | 0.73 | 0.45/0.87 | 0.67 | 0.96 | 0.95 | 0.37/0.76 | 0.49 | 0.90 | 0.77 | 0.32/0.80 | 0.35 | 0.85 | 0.83 | 0.38/0.81 |
| FIRM-Video-8B (Qwen3-VL-8B) | 0.65 | 0.77 | 0.50/0.88 | 0.69 | 0.86 | 0.88 | 0.40/0.80 | 0.53 | 0.82 | 0.74 | 0.36/0.83 | 0.51 | 0.78 | 0.80 | 0.42/0.84 |
| Sampling Strategy | Total Score | Quality Score | Semantic Score | Video Quality | Video–Condition Consistency | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Subject Consistency | Background Consistency | Temporal Flickering | Imaging Quality | Multiple Objects | Color | Spatial Relationship | Appearance Style | ||||
| T2V model: LaVie-Base | |||||||||||
| Random | 79.42 | 81.34 | 71.77 | 92.09 | 97.64 | 97.63 | 65.92 | 37.73 | 89.65 | 43.30 | 23.94 |
| By Qwen3-VL-8B | 80.03 | 81.61 | 73.73 | 92.64 | 97.56 | 97.60 | 66.88 | 50.91 | 83.11 | 47.32 | 24.27 |
| By InternVL3-8B | 80.45 | 82.10 | 73.82 | 92.94 | 97.56 | 97.56 | 67.02 | 50.30 | 85.85 | 45.81 | 24.10 |
| By VideoScore2 | 79.95 | 81.61 | 73.33 | 92.95 | 97.80 | 97.76 | 66.54 | 53.73 | 86.86 | 45.70 | 23.87 |
| By FIRM-Video-8B | 80.72 | 82.14 | 75.06 | 93.66 | 97.91 | 98.12 | 67.46 | 58.99 | 88.20 | 51.20 | 24.40 |
| T2V model: CogVideoX-2B | |||||||||||
| Random | 78.57 | 80.90 | 69.25 | 93.47 | 96.66 | 96.90 | 59.72 | 49.16 | 88.43 | 61.30 | 22.95 |
| By Qwen3-VL-8B | 78.49 | 80.22 | 71.59 | 94.48 | 96.84 | 96.90 | 59.20 | 55.56 | 88.40 | 65.80 | 23.30 |
| By InternVL3-8B | 78.35 | 80.27 | 70.65 | 94.53 | 96.89 | 97.04 | 59.90 | 51.30 | 89.49 | 61.09 | 23.24 |
| By VideoScore2 | 78.67 | 80.30 | 72.17 | 94.63 | 96.78 | 96.98 | 60.08 | 61.05 | 87.59 | 63.82 | 22.54 |
| By FIRM-Video-8B | 79.68 | 81.03 | 74.28 | 94.92 | 97.28 | 97.33 | 60.00 | 62.88 | 92.37 | 66.86 | 23.02 |
| T2V model: Wan2.1-T2V-1.3B | |||||||||||
| Random | 80.40 | 84.13 | 65.49 | 94.57 | 98.08 | 98.98 | 68.18 | 57.01 | 85.63 | 76.13 | 19.73 |
| By Qwen3-VL-8B | 81.43 | 84.30 | 69.96 | 94.57 | 97.90 | 98.87 | 67.57 | 67.91 | 85.94 | 74.39 | 20.67 |
| By InternVL3-8B | 81.70 | 84.43 | 70.78 | 95.46 | 98.32 | 98.84 | 67.71 | 67.61 | 86.15 | 77.08 | 20.30 |
| By VideoScore2 | 81.64 | 84.86 | 68.74 | 95.45 | 98.13 | 98.99 | 67.57 | 72.26 | 83.91 | 74.56 | 19.98 |
| By FIRM-Video-8B | 82.36 | 84.91 | 72.19 | 96.54 | 98.68 | 99.22 | 68.62 | 73.40 | 92.23 | 81.10 | 20.64 |

| Dimension | Aggregation | MAE ↓ | STD ↓ | Acc. ↑ | Relaxed Acc. ↑ | SRCC ↑ |
|---|---|---|---|---|---|---|
| IF | mean | 0.68 | 0.70 | 0.44 | 0.88 | 0.69 |
| Importance weighted (ours) | 0.68 | 0.68 | 0.43 | 0.90 | 0.69 | |
| WC | mean | 0.86 | 0.83 | 0.37 | 0.82 | 0.61 |
| Importance weighted (ours) | 0.80 | 0.83 | 0.41 | 0.84 | 0.63 |
| Dataset | Dimension | Score=1 | Score=2 | Score=3 | Score=4 | Score 5 | Total |
|---|---|---|---|---|---|---|---|
| FIRM-Video-90K | IF | 5,104 (17.4%) | 8,900 (30.3%) | 5,551 (18.9%) | 3,806 (13.0%) | 5,987 (20.4%) | 29,348 |
| WC | 4,795 (16.3%) | 11,129 (37.9%) | 5,649 (19.2%) | 2,510 (8.6%) | 5,265 (17.9%) | 29,348 | |
| PQ | 874 (3.0%) | 3,749 (12.8%) | 4,257 (14.5%) | 18,331 (62.5%) | 2,137 (7.3%) | 29,348 | |
| All | 10,773 | 23,778 | 15,457 | 24,647 | 13,389 | 88,044 | |
| FIRM-Video-Bench | IF | 28 (11.2%) | 73 (29.2%) | 63 (25.2%) | 59 (23.6%) | 27 (10.8%) | 250 |
| WC | 48 (19.2%) | 76 (30.4%) | 48 (19.2%) | 44 (17.6%) | 34 (13.6%) | 250 | |
| PQ | 11 (4.4%) | 54 (21.6%) | 69 (27.6%) | 59 (23.6%) | 57 (22.8%) | 250 | |
| All | 87 | 203 | 180 | 162 | 118 | 750 |

| Reward | VBench Dimension | VBench Group |
|---|---|---|
| IF | Dynamic Degree | Quality |
| Overall Consistency | Cond. Consist. | |
| Object Class | Cond. Consist. | |
| Multiple Objects | Cond. Consist. | |
| Human Action | Cond. Consist. | |
| Color | Cond. Consist. | |
| Spatial Relationship | Cond. Consist. | |
| Scene | Cond. Consist. | |
| Appearance Style | Cond. Consist. | |
| Temporal Style | Cond. Consist. | |
| WC | Subject Consistency | Quality |
| Background Consistency | Quality | |
| Motion Smoothness | Quality | |
| PQ | Temporal Flickering | Quality |
| Aesthetic Quality | Quality | |
| Imaging Quality | Quality |

| T2V Model | Sampling Strategy | Subject Consistency | Background Consistency | Temporal Flickering | Motion Smoothness | Dynamic Degree | Aesthetic Quality | Imaging Quality |
|---|---|---|---|---|---|---|---|---|
| LaVie-Base | Random | 92.09 | 97.64 | 97.63 | 96.63 | 55.56 | 64.53 | 65.92 |
| Qwen3-VL-8B | 92.64 | 97.56 | 97.60 | 96.62 | 55.56 | 64.93 | 66.88 | |
| InternVL3-8B | 92.94 | 97.56 | 97.56 | 96.69 | 61.11 | 64.74 | 67.02 | |
| VideoScore2 | 92.95 | 97.80 | 97.76 | 96.39 | 56.94 | 64.27 | 66.54 | |
| FIRM-Video-8B | 93.66 | 97.91 | 98.12 | 96.58 | 56.94 | 64.17 | 67.46 | |
| CogVideoX-2B | Random | 93.47 | 96.66 | 96.90 | 97.23 | 73.61 | 58.47 | 59.72 |
| Qwen3-VL-8B | 94.48 | 96.84 | 96.90 | 97.35 | 62.50 | 58.30 | 59.20 | |
| InternVL3-8B | 94.53 | 96.89 | 97.04 | 97.24 | 62.50 | 57.84 | 59.90 | |
| VideoScore2 | 94.63 | 96.78 | 96.98 | 97.27 | 59.72 | 59.31 | 60.08 | |
| FIRM-Video-8B | 94.92 | 97.28 | 97.33 | 97.34 | 65.28 | 59.15 | 60.00 | |
| Wan2.1-T2V-1.3B | Random | 94.57 | 98.08 | 98.98 | 98.21 | 62.50 | 64.38 | 68.18 |
| Qwen3-VL-8B | 94.57 | 97.90 | 98.87 | 98.10 | 66.67 | 64.94 | 67.57 | |
| InternVL3-8B | 95.46 | 98.32 | 98.84 | 98.29 | 63.89 | 64.87 | 67.71 | |
| VideoScore2 | 95.45 | 98.13 | 98.99 | 98.43 | 68.06 | 65.10 | 67.57 | |
| FIRM-Video-8B | 96.54 | 98.68 | 99.22 | 98.49 | 59.72 | 65.67 | 68.62 |

| T2V Model | Sampling Strategy | Object Class | Multiple Objects | Human Action | Color | Spatial Relationship | Scene | Appearance Style | Temporal Style | Overall Consistency |
|---|---|---|---|---|---|---|---|---|---|---|
| LaVie-Base | Random | 93.12 | 37.73 | 92.00 | 89.65 | 43.30 | 52.33 | 23.94 | 24.73 | 27.20 |
| Qwen3-VL-8B | 91.61 | 50.91 | 96.00 | 83.11 | 47.32 | 52.69 | 24.27 | 25.25 | 27.75 | |
| InternVL3-8B | 95.81 | 50.30 | 96.00 | 85.85 | 45.81 | 51.02 | 24.10 | 25.03 | 27.46 | |
| VideoScore2 | 91.06 | 53.73 | 93.00 | 86.86 | 45.70 | 50.80 | 23.87 | 25.16 | 27.33 | |
| FIRM-Video-8B | 90.51 | 58.99 | 93.00 | 88.20 | 51.20 | 52.03 | 24.40 | 25.19 | 27.57 | |
| CogVideoX-2B | Random | 76.11 | 49.16 | 87.00 | 88.43 | 61.30 | 39.90 | 22.95 | 23.77 | 24.43 |
| Qwen3-VL-8B | 83.62 | 55.56 | 91.00 | 88.40 | 65.80 | 35.83 | 23.30 | 23.84 | 25.27 | |
| InternVL3-8B | 82.91 | 51.30 | 89.00 | 89.49 | 61.09 | 37.79 | 23.24 | 23.74 | 25.30 | |
| VideoScore2 | 84.41 | 61.05 | 91.00 | 87.59 | 63.82 | 40.48 | 22.54 | 23.54 | 25.06 | |
| FIRM-Video-8B | 85.21 | 62.88 | 90.00 | 92.37 | 66.86 | 45.13 | 23.02 | 24.29 | 25.13 | |
| Wan2.1-T2V-1.3B | Random | 74.76 | 57.01 | 68.00 | 85.63 | 76.13 | 24.35 | 19.73 | 23.74 | 23.31 |
| Qwen3-VL-8B | 82.36 | 67.91 | 83.00 | 85.94 | 74.39 | 25.51 | 20.67 | 23.53 | 24.78 | |
| InternVL3-8B | 84.34 | 67.61 | 82.00 | 86.15 | 77.08 | 30.38 | 20.30 | 23.86 | 24.15 | |
| VideoScore2 | 79.27 | 72.26 | 79.00 | 83.91 | 74.56 | 23.84 | 19.98 | 23.73 | 23.88 | |
| FIRM-Video-8B | 80.22 | 73.40 | 83.00 | 92.23 | 81.10 | 29.14 | 20.64 | 23.52 | 24.57 |

Findings
- On FIRM-Video-Bench, FIRM-Video-8B (built on Qwen3-VL-8B) achieved the lowest overall MAE of 0.78 among all tested proprietary and open-source models, and had the lowest error specifically on the World Coherence dimension.
- On the same benchmark, GPT-5 had the lowest error for Instruction Following (0.62) and Doubao-Seed-2.0-Lite had the lowest error for Perceptual Quality (0.80), showing FIRM-Video-8B did not lead on every single dimension but led overall.
- In Best-of-8 sampling with VBench, FIRM-Video-8B produced the best Total, Quality, and Semantic Scores across all three tested video generators (LaVie-Base, CogVideoX-2B, Wan2.1-T2V-1.3B), beating the next-best method by 0.27 to 1.01 points on Total Score and 1.24 to 2.11 points on Semantic Score.
- On the out-of-domain MJ-Bench-Video benchmark, one FIRM-Video-8B variant led on Alignment and Consistency & Coherence, and ranked second on Overall preference accuracy, indicating some generalization beyond the training data.
- An ablation study showed that weighting checklist items by importance (rather than averaging them equally) improved accuracy, and for the World Coherence dimension specifically reduced the error from 0.86 to 0.80.

Where it can be used
- Selecting the best output among several AI-generated videos for the same prompt (Best-of-N filtering) before showing it to a user.
- Automated quality control for filtering large video-generation datasets by instruction-following, physical plausibility, or visual defects.
- As a scoring signal for future reinforcement-learning or preference-optimization training of text-to-video generation models, an application the authors mention but have not yet tested.

Limits and open work
- The authors explicitly state they have not yet used FIRM-Video-8B to directly train or optimize a video generator; it was only tested as an evaluator and selector, so its usefulness for reinforcement learning-based alignment remains unverified.
- The human benchmark used for testing (FIRM-Video-Bench) covers only 250 videos and 750 annotations, a relatively small test set compared to typical benchmark sizes.
- Training videos were drawn mainly from two existing preference datasets and over 20 generation models, so performance on video styles or generators outside this distribution is not directly demonstrated.
- The reward model relies on 8 uniformly sampled frames per video rather than full video, which may miss very short or fine-grained temporal issues.
- Score thresholds mapping the continuous checklist scores to 5-point ratings were set empirically, and the paper does not report how sensitive results are to this choice.
Why it matters
Video-generation systems are improving fast, but judging whether their output is actually good, faithful to the prompt, and free of physical nonsense is still mostly done by hand or by unreliable automatic scorers. A more trustworthy, checklist-verified reward model gives developers a cheaper way to filter, rank, and eventually train better video generators.
Terms in this paper
- Reward model · An AI model trained to output a quality score for another AI's output, used for filtering, ranking, or guiding training.
- Checklist-driven / check-before-score · An approach where specific yes/no questions are verified against evidence first, and only verified answers are combined into a final score.
- Instruction Following / World Coherence / Perceptual Quality · The three evaluation dimensions this paper uses: does the video match the text prompt, is it physically/logically sensible, and is the image quality clean.
- Best-of-N sampling · Generating N candidate videos for the same prompt and picking the single best one according to some scoring method.
- MAE (Mean Absolute Error) · The average size of the gap between a model's predicted score and the human-given score; lower means closer to human judgment.
Original abstract (English)
Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference efficiency places high demands on the quality of training supervision. Existing approaches often rely on holistic judges with fixed rubrics or open-ended reasoning, leading to incomplete inspection, unfaithful justification, and entangled attribution. We introduce FIRM-Video, a unified checklist-driven data construction framework based on a check-before-score principle: construct dimension-specific checklists, verify each criterion against temporal visual evidence, and aggregate only verified decisions. For Instruction Following, FIRM-Video decomposes prompts into weighted atomic requirements; for World Coherence, it constructs prompt-calibrated, target-specific checks grounded in visible entities and actions; and for Perceptual Quality, it applies a generic taxonomy of visual defects. The verified criteria and scores are further transformed into natural-language analyses for end-to-end reward modeling. Subsequently, we construct FIRM-Video-90K with 88,044 dimension-specific instances from 29,348 videos, and introduce FIRM-Video-Bench with 750 point-wise human annotations across 250 videos. The Qwen3-VL-based FIRM-Video-8B achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.
Read on arXivLatest papers
- FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial OutcomesA dataset that finally teaches AI what biology, chemistry, and physics peer reviewers actually argue about, not just CS reviewers
- JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness EvolutionAn AI system that writes a custom 'operating scaffold' for other AI agents on the spot, for every new task
- The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling PipelineAI language models still charge a hidden 'dialect tax' on AAVE and other non-standard English at every stage, not just tokenization
- Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal BayesiansA math model shows that even a perfectly rational person can be talked into delusion by a chatbot that keeps agreeing with them
- Autonomous Mathematical Discovery in an Open-World Multi-Agent EnvironmentAI agents from different companies self-organized in an open-world simulation and produced new results on five math problems, with no one directing them
- Automata from Agent Traces: Failure and Next-Step PredictionCompressing thousands of LLM agent execution logs into one tiny 7-to-43-state machine that predicts both the next action and eventual failure
- MARS: Multi-Specialist LLM Relay System for Competitive ProgrammingLetting topic-specialist AIs take turns fixing code beats one generalist coder on programming contest problems
- AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared WorkspaceLetting multiple AI coding agents share one workspace and coordinate in real time beats running them one-by-one or in uncoordinated parallel
Latest from METAL LAB
- Tencent open-sources Hy4 preview, edges out GPT-5.6 on coding benchmark
- OpenAI, Kakao Adopt Google's SynthID Watermark
- Google DeepMind Unveils Gemini for Science, an AI Tool for Researchers
- Robot Apollo, running Gemini Robotics 2, found kitchen chores harder than sports
- Sony Music, Warner Chappell Sue Anthropic, Naming Founders Personally
Figures: Peiyuan Zhang et al., arXiv:2608.21839, arxiv-nonexclusive