FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling
让给AI视频打分的评审模型先核实证据再给分,而不是凭印象直接打分
FIRM-Video是一种为视频打分AI构建训练数据的方法,先针对指令遵循、世界常理和画面质量三个维度分别生成可核实的是非清单,再逐条对照视频画面进行验证,最后只用已验证的结果计算分数。该方法产出了含88044条训练样本的数据集FIRM-Video-90K,以及由人工标注的750条评测数据FIRM-Video-Bench。用这些数据训练出的80亿参数模型FIRM-Video-8B,打分结果比现有模型更接近人类判断,在从多个候选视频中挑选最佳结果时也表现更好。
METAL LAB 解读图
FIRM-Video如何在打分前先核实证据
证据状态已报告实测结果
- 1. 生成检查清单针对每个视频与提示词组合,分别生成三类检查清单:指令遵循对应的具体要求条目、世界常理对应的实体与动作检查项、画面质量对应的固定常见缺陷列表。
- 2. 核实证据多模态评审模型对照实际视频画面,逐条判断检查清单中的每一项是否成立,并给出简要视觉依据;无法核实的条目按不满足处理。
- 3. 汇总打分只使用已核实的结果,按重要性加权汇总成每个维度的最终分数,避免同一个问题在不同维度被重复扣分。
- 4. 构建数据集与测试集将上述流程应用于29348个视频,产出88044条训练样本组成FIRM-Video-90K;另外250个视频由人工评分,组成含750条标注的测试集FIRM-Video-Bench。
- 5. 端到端评分模型用这批数据微调Qwen3-VL-8B得到FIRM-Video-8B,实际使用时不再重复检查清单流程,而是直接根据提示词和视频画面一次性预测分数与说明文字。
他们做了什么
- 问题:现有的视频评分AI往往只关注明显特征而忽略细节,常常先定好分数再编造理由,还会把不同类型的问题混在一起重复扣分(比如把物理错误当成画质问题再扣一次分)。
- 方法:给定提示词和生成的视频后,先分别为指令遵循、世界常理、画面质量三个维度生成专属检查清单,再让评审模型对照实际视频画面逐条核实每个条目,最后只将已核实的结果汇总成分数和文字说明。
- 数据构建:该流程被应用于约29348个AI生成的视频,产出88044条带标注的训练样本,组成数据集FIRM-Video-90K;另外250个视频由人类专家单独评分,组成含750条标注的测试集FIRM-Video-Bench。
- 结果:用这批数据微调Qwen3-VL-8B得到的FIRM-Video-8B,在FIRM-Video-Bench上的平均绝对误差为0.78,远低于未微调基础模型的1.33,在所有被测试的商业和开源模型中综合误差最低。
- 结果:在从8个AI生成候选视频中挑选最佳视频的实验中(Best-of-8),按照VBench这一标准评测指标衡量,FIRM-Video-8B在三种不同的视频生成模型上都比其他挑选方法稳定选出质量更高的视频。

| Model | Instruction Following | World Coherence | Perceptual Quality | Overall | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MAE↓ | STD↓ | Acc./Relaxed Acc.↑ | SRCC↑ | MAE↓ | STD↓ | Acc./Relaxed Acc.↑ | SRCC↑ | MAE↓ | STD↓ | Acc./Relaxed Acc.↑ | SRCC↑ | MAE↓ | STD↓ | Acc./Relaxed Acc.↑ | |
| Closed-source models | |||||||||||||||
| GPT-5 | 0.62 | 0.71 | 0.50/0.88 | 0.74 | 1.66 | 1.19 | 0.20/0.48 | 0.49 | 0.95 | 0.86 | 0.34/0.76 | 0.51 | 1.08 | 1.03 | 0.35/0.71 |
| Gemini-3.1-Pro | 0.77 | 0.76 | 0.40/0.86 | 0.71 | 1.27 | 1.08 | 0.28/0.62 | 0.54 | 1.44 | 1.12 | 0.21/0.59 | 0.41 | 1.16 | 1.04 | 0.30/0.69 |
| Doubao-Seed-2.0-Lite | 0.80 | 0.83 | 0.41/0.84 | 0.67 | 1.56 | 1.24 | 0.24/0.53 | 0.44 | 0.80 | 0.75 | 0.38/0.84 | 0.55 | 1.05 | 1.03 | 0.35/0.73 |
| Open-source models | |||||||||||||||
| InternVL3-8B | 1.26 | 0.95 | 0.24/0.60 | 0.57 | 2.10 | 1.32 | 0.15/0.36 | 0.22 | 1.29 | 1.00 | 0.24/0.61 | 0.22 | 1.56 | 1.17 | 0.21/0.52 |
| InternVL3-38B | 0.82 | 0.83 | 0.40/0.82 | 0.65 | 2.11 | 1.32 | 0.15/0.35 | 0.15 | 1.33 | 1.08 | 0.27/0.58 | 0.20 | 1.42 | 1.21 | 0.27/0.58 |
| Qwen3-VL-8B | 0.93 | 0.89 | 0.34/0.80 | 0.60 | 1.79 | 1.35 | 0.22/0.46 | 0.29 | 1.28 | 1.04 | 0.28/0.60 | 0.34 | 1.33 | 1.16 | 0.28/0.62 |
| Qwen3-VL-30B-A3B | 0.95 | 0.87 | 0.32/0.80 | 0.62 | 1.90 | 1.31 | 0.19/0.40 | 0.22 | 1.36 | 1.09 | 0.28/0.55 | 0.37 | 1.40 | 1.17 | 0.27/0.58 |
| Qwen3-VL-235B-A22B | 0.93 | 0.87 | 0.33/0.80 | 0.65 | 1.63 | 1.30 | 0.24/0.52 | 0.31 | 1.27 | 1.04 | 0.28/0.60 | 0.40 | 1.27 | 1.12 | 0.29/0.64 |
| Our methods | |||||||||||||||
| FIRM-Video Data Pipeline† | 0.68 | 0.68 | 0.43/0.90 | 0.69 | 0.80 | 0.83 | 0.41/0.84 | 0.63 | 0.74 | 0.70 | 0.40/0.86 | 0.61 | 0.73 | 0.74 | 0.41/0.87 |
| FIRM-Video-8B (InternVL3-8B) | 0.69 | 0.73 | 0.45/0.87 | 0.67 | 0.96 | 0.95 | 0.37/0.76 | 0.49 | 0.90 | 0.77 | 0.32/0.80 | 0.35 | 0.85 | 0.83 | 0.38/0.81 |
| FIRM-Video-8B (Qwen3-VL-8B) | 0.65 | 0.77 | 0.50/0.88 | 0.69 | 0.86 | 0.88 | 0.40/0.80 | 0.53 | 0.82 | 0.74 | 0.36/0.83 | 0.51 | 0.78 | 0.80 | 0.42/0.84 |
| Sampling Strategy | Total Score | Quality Score | Semantic Score | Video Quality | Video–Condition Consistency | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Subject Consistency | Background Consistency | Temporal Flickering | Imaging Quality | Multiple Objects | Color | Spatial Relationship | Appearance Style | ||||
| T2V model: LaVie-Base | |||||||||||
| Random | 79.42 | 81.34 | 71.77 | 92.09 | 97.64 | 97.63 | 65.92 | 37.73 | 89.65 | 43.30 | 23.94 |
| By Qwen3-VL-8B | 80.03 | 81.61 | 73.73 | 92.64 | 97.56 | 97.60 | 66.88 | 50.91 | 83.11 | 47.32 | 24.27 |
| By InternVL3-8B | 80.45 | 82.10 | 73.82 | 92.94 | 97.56 | 97.56 | 67.02 | 50.30 | 85.85 | 45.81 | 24.10 |
| By VideoScore2 | 79.95 | 81.61 | 73.33 | 92.95 | 97.80 | 97.76 | 66.54 | 53.73 | 86.86 | 45.70 | 23.87 |
| By FIRM-Video-8B | 80.72 | 82.14 | 75.06 | 93.66 | 97.91 | 98.12 | 67.46 | 58.99 | 88.20 | 51.20 | 24.40 |
| T2V model: CogVideoX-2B | |||||||||||
| Random | 78.57 | 80.90 | 69.25 | 93.47 | 96.66 | 96.90 | 59.72 | 49.16 | 88.43 | 61.30 | 22.95 |
| By Qwen3-VL-8B | 78.49 | 80.22 | 71.59 | 94.48 | 96.84 | 96.90 | 59.20 | 55.56 | 88.40 | 65.80 | 23.30 |
| By InternVL3-8B | 78.35 | 80.27 | 70.65 | 94.53 | 96.89 | 97.04 | 59.90 | 51.30 | 89.49 | 61.09 | 23.24 |
| By VideoScore2 | 78.67 | 80.30 | 72.17 | 94.63 | 96.78 | 96.98 | 60.08 | 61.05 | 87.59 | 63.82 | 22.54 |
| By FIRM-Video-8B | 79.68 | 81.03 | 74.28 | 94.92 | 97.28 | 97.33 | 60.00 | 62.88 | 92.37 | 66.86 | 23.02 |
| T2V model: Wan2.1-T2V-1.3B | |||||||||||
| Random | 80.40 | 84.13 | 65.49 | 94.57 | 98.08 | 98.98 | 68.18 | 57.01 | 85.63 | 76.13 | 19.73 |
| By Qwen3-VL-8B | 81.43 | 84.30 | 69.96 | 94.57 | 97.90 | 98.87 | 67.57 | 67.91 | 85.94 | 74.39 | 20.67 |
| By InternVL3-8B | 81.70 | 84.43 | 70.78 | 95.46 | 98.32 | 98.84 | 67.71 | 67.61 | 86.15 | 77.08 | 20.30 |
| By VideoScore2 | 81.64 | 84.86 | 68.74 | 95.45 | 98.13 | 98.99 | 67.57 | 72.26 | 83.91 | 74.56 | 19.98 |
| By FIRM-Video-8B | 82.36 | 84.91 | 72.19 | 96.54 | 98.68 | 99.22 | 68.62 | 73.40 | 92.23 | 81.10 | 20.64 |

| Dimension | Aggregation | MAE ↓ | STD ↓ | Acc. ↑ | Relaxed Acc. ↑ | SRCC ↑ |
|---|---|---|---|---|---|---|
| IF | mean | 0.68 | 0.70 | 0.44 | 0.88 | 0.69 |
| Importance weighted (ours) | 0.68 | 0.68 | 0.43 | 0.90 | 0.69 | |
| WC | mean | 0.86 | 0.83 | 0.37 | 0.82 | 0.61 |
| Importance weighted (ours) | 0.80 | 0.83 | 0.41 | 0.84 | 0.63 |
| Dataset | Dimension | Score=1 | Score=2 | Score=3 | Score=4 | Score 5 | Total |
|---|---|---|---|---|---|---|---|
| FIRM-Video-90K | IF | 5,104 (17.4%) | 8,900 (30.3%) | 5,551 (18.9%) | 3,806 (13.0%) | 5,987 (20.4%) | 29,348 |
| WC | 4,795 (16.3%) | 11,129 (37.9%) | 5,649 (19.2%) | 2,510 (8.6%) | 5,265 (17.9%) | 29,348 | |
| PQ | 874 (3.0%) | 3,749 (12.8%) | 4,257 (14.5%) | 18,331 (62.5%) | 2,137 (7.3%) | 29,348 | |
| All | 10,773 | 23,778 | 15,457 | 24,647 | 13,389 | 88,044 | |
| FIRM-Video-Bench | IF | 28 (11.2%) | 73 (29.2%) | 63 (25.2%) | 59 (23.6%) | 27 (10.8%) | 250 |
| WC | 48 (19.2%) | 76 (30.4%) | 48 (19.2%) | 44 (17.6%) | 34 (13.6%) | 250 | |
| PQ | 11 (4.4%) | 54 (21.6%) | 69 (27.6%) | 59 (23.6%) | 57 (22.8%) | 250 | |
| All | 87 | 203 | 180 | 162 | 118 | 750 |

| Reward | VBench Dimension | VBench Group |
|---|---|---|
| IF | Dynamic Degree | Quality |
| Overall Consistency | Cond. Consist. | |
| Object Class | Cond. Consist. | |
| Multiple Objects | Cond. Consist. | |
| Human Action | Cond. Consist. | |
| Color | Cond. Consist. | |
| Spatial Relationship | Cond. Consist. | |
| Scene | Cond. Consist. | |
| Appearance Style | Cond. Consist. | |
| Temporal Style | Cond. Consist. | |
| WC | Subject Consistency | Quality |
| Background Consistency | Quality | |
| Motion Smoothness | Quality | |
| PQ | Temporal Flickering | Quality |
| Aesthetic Quality | Quality | |
| Imaging Quality | Quality |

| T2V Model | Sampling Strategy | Subject Consistency | Background Consistency | Temporal Flickering | Motion Smoothness | Dynamic Degree | Aesthetic Quality | Imaging Quality |
|---|---|---|---|---|---|---|---|---|
| LaVie-Base | Random | 92.09 | 97.64 | 97.63 | 96.63 | 55.56 | 64.53 | 65.92 |
| Qwen3-VL-8B | 92.64 | 97.56 | 97.60 | 96.62 | 55.56 | 64.93 | 66.88 | |
| InternVL3-8B | 92.94 | 97.56 | 97.56 | 96.69 | 61.11 | 64.74 | 67.02 | |
| VideoScore2 | 92.95 | 97.80 | 97.76 | 96.39 | 56.94 | 64.27 | 66.54 | |
| FIRM-Video-8B | 93.66 | 97.91 | 98.12 | 96.58 | 56.94 | 64.17 | 67.46 | |
| CogVideoX-2B | Random | 93.47 | 96.66 | 96.90 | 97.23 | 73.61 | 58.47 | 59.72 |
| Qwen3-VL-8B | 94.48 | 96.84 | 96.90 | 97.35 | 62.50 | 58.30 | 59.20 | |
| InternVL3-8B | 94.53 | 96.89 | 97.04 | 97.24 | 62.50 | 57.84 | 59.90 | |
| VideoScore2 | 94.63 | 96.78 | 96.98 | 97.27 | 59.72 | 59.31 | 60.08 | |
| FIRM-Video-8B | 94.92 | 97.28 | 97.33 | 97.34 | 65.28 | 59.15 | 60.00 | |
| Wan2.1-T2V-1.3B | Random | 94.57 | 98.08 | 98.98 | 98.21 | 62.50 | 64.38 | 68.18 |
| Qwen3-VL-8B | 94.57 | 97.90 | 98.87 | 98.10 | 66.67 | 64.94 | 67.57 | |
| InternVL3-8B | 95.46 | 98.32 | 98.84 | 98.29 | 63.89 | 64.87 | 67.71 | |
| VideoScore2 | 95.45 | 98.13 | 98.99 | 98.43 | 68.06 | 65.10 | 67.57 | |
| FIRM-Video-8B | 96.54 | 98.68 | 99.22 | 98.49 | 59.72 | 65.67 | 68.62 |

| T2V Model | Sampling Strategy | Object Class | Multiple Objects | Human Action | Color | Spatial Relationship | Scene | Appearance Style | Temporal Style | Overall Consistency |
|---|---|---|---|---|---|---|---|---|---|---|
| LaVie-Base | Random | 93.12 | 37.73 | 92.00 | 89.65 | 43.30 | 52.33 | 23.94 | 24.73 | 27.20 |
| Qwen3-VL-8B | 91.61 | 50.91 | 96.00 | 83.11 | 47.32 | 52.69 | 24.27 | 25.25 | 27.75 | |
| InternVL3-8B | 95.81 | 50.30 | 96.00 | 85.85 | 45.81 | 51.02 | 24.10 | 25.03 | 27.46 | |
| VideoScore2 | 91.06 | 53.73 | 93.00 | 86.86 | 45.70 | 50.80 | 23.87 | 25.16 | 27.33 | |
| FIRM-Video-8B | 90.51 | 58.99 | 93.00 | 88.20 | 51.20 | 52.03 | 24.40 | 25.19 | 27.57 | |
| CogVideoX-2B | Random | 76.11 | 49.16 | 87.00 | 88.43 | 61.30 | 39.90 | 22.95 | 23.77 | 24.43 |
| Qwen3-VL-8B | 83.62 | 55.56 | 91.00 | 88.40 | 65.80 | 35.83 | 23.30 | 23.84 | 25.27 | |
| InternVL3-8B | 82.91 | 51.30 | 89.00 | 89.49 | 61.09 | 37.79 | 23.24 | 23.74 | 25.30 | |
| VideoScore2 | 84.41 | 61.05 | 91.00 | 87.59 | 63.82 | 40.48 | 22.54 | 23.54 | 25.06 | |
| FIRM-Video-8B | 85.21 | 62.88 | 90.00 | 92.37 | 66.86 | 45.13 | 23.02 | 24.29 | 25.13 | |
| Wan2.1-T2V-1.3B | Random | 74.76 | 57.01 | 68.00 | 85.63 | 76.13 | 24.35 | 19.73 | 23.74 | 23.31 |
| Qwen3-VL-8B | 82.36 | 67.91 | 83.00 | 85.94 | 74.39 | 25.51 | 20.67 | 23.53 | 24.78 | |
| InternVL3-8B | 84.34 | 67.61 | 82.00 | 86.15 | 77.08 | 30.38 | 20.30 | 23.86 | 24.15 | |
| VideoScore2 | 79.27 | 72.26 | 79.00 | 83.91 | 74.56 | 23.84 | 19.98 | 23.73 | 23.88 | |
| FIRM-Video-8B | 80.22 | 73.40 | 83.00 | 92.23 | 81.10 | 29.14 | 20.64 | 23.52 | 24.57 |

研究结果
- 在FIRM-Video-Bench测试中,基于Qwen3-VL-8B构建的FIRM-Video-8B综合平均绝对误差为0.78,在所有测试的商业和开源模型中最低,并且在“世界常理”这一维度上误差单独最低。
- 同一测试中,GPT-5在“指令遵循”维度误差最低为0.62,Doubao-Seed-2.0-Lite在“画面质量”维度误差最低为0.80,说明FIRM-Video-8B并非每个细分维度都排第一,但综合表现最优。
- 在基于VBench的Best-of-8实验中,FIRM-Video-8B在三种不同视频生成模型(LaVie-Base、CogVideoX-2B、Wan2.1-T2V-1.3B)上综合分、画质分和语义一致性分均排名第一,比次优方法综合分高出0.27到1.01分,语义一致性分高出1.24到2.11分。
- 在训练数据之外的MJ-Bench-Video测试中,FIRM-Video-8B的一个版本在“一致性”和“对齐度”维度排名第一,在整体偏好准确率上排名第二,显示出一定的泛化能力。
- 一项对比实验显示,按重要性加权汇总检查清单结果比简单平均更准确,尤其在“世界常理”维度上,误差从0.86降到了0.80。

可应用场景
- 从同一提示词生成的多个候选视频中挑出质量最好的一个再展示给用户(Best-of-N筛选)。
- 按指令遵循度、物理合理性或画面缺陷对大批量生成的视频数据集进行自动质量筛查。
- 作者提及但尚未实际测试的用途:将该评分模型用作强化学习或偏好优化的信号,直接用于训练和改进视频生成模型本身。

局限与待验证事项
- 作者明确说明尚未将FIRM-Video-8B用于直接训练或优化视频生成模型,目前只验证了其作为评审和筛选工具的效果,其在强化学习式对齐训练中的作用还未被验证。
- 用于测试的人工标注基准FIRM-Video-Bench仅包含250个视频、750条标注,规模相对较小。
- 训练数据主要来自两个已有的偏好数据集及二十多个生成模型的输出,超出这一范围的视频风格或生成模型上的表现尚未直接验证。
- 评审模型只看每个视频均匀抽取的8帧画面,而不是完整视频,可能会漏掉非常短暂或细微的时间性问题。
- 将连续的检查清单分数映射为五档评分所用的阈值是凭经验设定的,论文没有报告这一设定对结果的敏感程度。
为什么重要
视频生成AI进步很快,但判断生成结果是否真正贴合提示词、是否符合常理、画质是否过关,目前主要靠人工或不太可靠的自动打分。一个先核实证据再打分的评审模型,能让筛选和排序生成结果的成本更低,也为将来用它来直接训练、优化视频生成模型打下基础。
本文术语
- 奖励模型(Reward model) · 专门训练出来给另一个AI的输出结果打分的模型,用于筛选、排序或指导后续训练。
- 先核实再打分(check-before-score) · 先针对具体条目逐一核实是或否,再只用已核实的答案汇总出最终分数的做法。
- 指令遵循 / 世界常理 / 画面质量(IF/WC/PQ) · 本文使用的三个评测维度:视频是否符合提示词要求、是否符合物理常识、画面是否干净清晰。
- Best-of-N采样 · 针对同一提示词生成N个候选视频,再挑出其中打分最高的一个的方法。
- 平均绝对误差(MAE) · 模型预测分数与人类给出分数之间差距的平均值,数值越低说明越接近人类判断。
论文原文摘要(英文)
Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference efficiency places high demands on the quality of training supervision. Existing approaches often rely on holistic judges with fixed rubrics or open-ended reasoning, leading to incomplete inspection, unfaithful justification, and entangled attribution. We introduce FIRM-Video, a unified checklist-driven data construction framework based on a check-before-score principle: construct dimension-specific checklists, verify each criterion against temporal visual evidence, and aggregate only verified decisions. For Instruction Following, FIRM-Video decomposes prompts into weighted atomic requirements; for World Coherence, it constructs prompt-calibrated, target-specific checks grounded in visible entities and actions; and for Perceptual Quality, it applies a generic taxonomy of visual defects. The verified criteria and scores are further transformed into natural-language analyses for end-to-end reward modeling. Subsequently, we construct FIRM-Video-90K with 88,044 dimension-specific instances from 29,348 videos, and introduce FIRM-Video-Bench with 750 point-wise human annotations across 250 videos. The Qwen3-VL-based FIRM-Video-8B achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.
在 arXiv 阅读最新论文
- FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes一个让AI学会生物学、化学和物理学审稿人真正在意什么的数据集,而不只是计算机科学审稿人
- JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness EvolutionAI智能体的实力不只取决于模型本身,还取决于包裹模型的'执行框架',这项研究训练了一个能为每个新任务即时生成该框架的AI
- The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling PipelineAI语言模型对AAVE等非标准英语方言征收的隐性'方言税',不只出现在分词环节,而是贯穿训练与推理全流程
- Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians一个数学模型证明,哪怕是完全理性的人,也会被一味顺着自己说话的聊天机器人带入妄想
- Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment没有中央指挥,来自不同公司的AI智能体在开放世界环境中自行协作,在五个数学难题上做出了新发现
- Automata from Agent Traces: Failure and Next-Step Prediction把成千上万条LLM智能体的执行记录压缩成一个只有7到43个状态的小型状态机,同时预测下一步动作和最终是否失败
- MARS: Multi-Specialist LLM Relay System for Competitive Programming让不同算法专长的AI依次接力改代码,比单一全能AI更能解出编程竞赛题
- AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace让多个AI编程智能体在同一工作区实时协作,比按顺序执行或无协调地并行执行效果更好
METAL LAB 最新报道
图片来源: Peiyuan Zhang et al., arXiv:2608.21839, arxiv-nonexclusive