METAL LAB

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling

arXiv:2608.218392026-08-21

让给AI视频打分的评审模型先核实证据再给分,而不是凭印象直接打分

FIRM-Video是一种为视频打分AI构建训练数据的方法,先针对指令遵循、世界常理和画面质量三个维度分别生成可核实的是非清单,再逐条对照视频画面进行验证,最后只用已验证的结果计算分数。该方法产出了含88044条训练样本的数据集FIRM-Video-90K,以及由人工标注的750条评测数据FIRM-Video-Bench。用这些数据训练出的80亿参数模型FIRM-Video-8B,打分结果比现有模型更接近人类判断,在从多个候选视频中挑选最佳结果时也表现更好。

METAL LAB 解读图

FIRM-Video如何在打分前先核实证据

证据状态已报告实测结果

  1. 1. 生成检查清单针对每个视频与提示词组合,分别生成三类检查清单:指令遵循对应的具体要求条目、世界常理对应的实体与动作检查项、画面质量对应的固定常见缺陷列表。
  2. 2. 核实证据多模态评审模型对照实际视频画面,逐条判断检查清单中的每一项是否成立,并给出简要视觉依据;无法核实的条目按不满足处理。
  3. 3. 汇总打分只使用已核实的结果,按重要性加权汇总成每个维度的最终分数,避免同一个问题在不同维度被重复扣分。
  4. 4. 构建数据集与测试集将上述流程应用于29348个视频,产出88044条训练样本组成FIRM-Video-90K;另外250个视频由人工评分,组成含750条标注的测试集FIRM-Video-Bench。
  5. 5. 端到端评分模型用这批数据微调Qwen3-VL-8B得到FIRM-Video-8B,实际使用时不再重复检查清单流程,而是直接根据提示词和视频画面一次性预测分数与说明文字。
这是 METAL LAB 制作的解读图,并非论文作者提供的原图。

他们做了什么

  1. 问题:现有的视频评分AI往往只关注明显特征而忽略细节,常常先定好分数再编造理由,还会把不同类型的问题混在一起重复扣分(比如把物理错误当成画质问题再扣一次分)。
  2. 方法:给定提示词和生成的视频后,先分别为指令遵循、世界常理、画面质量三个维度生成专属检查清单,再让评审模型对照实际视频画面逐条核实每个条目,最后只将已核实的结果汇总成分数和文字说明。
  3. 数据构建:该流程被应用于约29348个AI生成的视频,产出88044条带标注的训练样本,组成数据集FIRM-Video-90K;另外250个视频由人类专家单独评分,组成含750条标注的测试集FIRM-Video-Bench。
  4. 结果:用这批数据微调Qwen3-VL-8B得到的FIRM-Video-8B,在FIRM-Video-Bench上的平均绝对误差为0.78,远低于未微调基础模型的1.33,在所有被测试的商业和开源模型中综合误差最低。
  5. 结果:在从8个AI生成候选视频中挑选最佳视频的实验中(Best-of-8),按照VBench这一标准评测指标衡量,FIRM-Video-8B在三种不同的视频生成模型上都比其他挑选方法稳定选出质量更高的视频。
Figure 1: Qualitative evaluation case of FIRM-Video-8B, illustrating its improved alignment with human judgments.
Figure 1: Qualitative evaluation case of FIRM-Video-8B, illustrating its improved alignment with human judgments.
Table 1: Performance comparison on FIRM-Video-Bench across three evaluation dimensions: Instruction Following (IF), World Coherence (WC), and Perceptual Quality (PQ). The two values under Acc./Relaxed Acc. denote strict and relaxed accuracy, respectively. The row marked with † reports the performance of direct data construction pipeline and is excluded from ranking.
ModelInstruction FollowingWorld CoherencePerceptual QualityOverall
MAE↓STD↓Acc./Relaxed Acc.↑SRCC↑MAE↓STD↓Acc./Relaxed Acc.↑SRCC↑MAE↓STD↓Acc./Relaxed Acc.↑SRCC↑MAE↓STD↓Acc./Relaxed Acc.↑
Closed-source models
GPT-50.620.710.50/0.880.741.661.190.20/0.480.490.950.860.34/0.760.511.081.030.35/0.71
Gemini-3.1-Pro0.770.760.40/0.860.711.271.080.28/0.620.541.441.120.21/0.590.411.161.040.30/0.69
Doubao-Seed-2.0-Lite0.800.830.41/0.840.671.561.240.24/0.530.440.800.750.38/0.840.551.051.030.35/0.73
Open-source models
InternVL3-8B1.260.950.24/0.600.572.101.320.15/0.360.221.291.000.24/0.610.221.561.170.21/0.52
InternVL3-38B0.820.830.40/0.820.652.111.320.15/0.350.151.331.080.27/0.580.201.421.210.27/0.58
Qwen3-VL-8B0.930.890.34/0.800.601.791.350.22/0.460.291.281.040.28/0.600.341.331.160.28/0.62
Qwen3-VL-30B-A3B0.950.870.32/0.800.621.901.310.19/0.400.221.361.090.28/0.550.371.401.170.27/0.58
Qwen3-VL-235B-A22B0.930.870.33/0.800.651.631.300.24/0.520.311.271.040.28/0.600.401.271.120.29/0.64
Our methods
FIRM-Video Data Pipeline†0.680.680.43/0.900.690.800.830.41/0.840.630.740.700.40/0.860.610.730.740.41/0.87
FIRM-Video-8B (InternVL3-8B)0.690.730.45/0.870.670.960.950.37/0.760.490.900.770.32/0.800.350.850.830.38/0.81
FIRM-Video-8B (Qwen3-VL-8B)0.650.770.50/0.880.690.860.880.40/0.800.530.820.740.36/0.830.510.780.800.42/0.84
Figure 2: Statistics of the FIRM-Video-90K dataset. Left: distributions across evaluation dimensions and score ranges. Right: distributions of video duration and prompt length.
Figure 2: Statistics of the FIRM-Video-90K dataset. Left: distributions across evaluation dimensions and score ranges. Right: distributions of video duration and prompt length.
Table 3: Best-of-N performance comparison of FIRM-Video-8B and other sampling strategies on VBench. The best and second-best results under each T2V generator are highlighted in bold and underlined, respectively.
Sampling StrategyTotal ScoreQuality ScoreSemantic ScoreVideo QualityVideo–Condition Consistency
Subject ConsistencyBackground ConsistencyTemporal FlickeringImaging QualityMultiple ObjectsColorSpatial RelationshipAppearance Style
T2V model: LaVie-Base
Random79.4281.3471.7792.0997.6497.6365.9237.7389.6543.3023.94
By Qwen3-VL-8B80.0381.6173.7392.6497.5697.6066.8850.9183.1147.3224.27
By InternVL3-8B80.4582.1073.8292.9497.5697.5667.0250.3085.8545.8124.10
By VideoScore279.9581.6173.3392.9597.8097.7666.5453.7386.8645.7023.87
By FIRM-Video-8B80.7282.1475.0693.6697.9198.1267.4658.9988.2051.2024.40
T2V model: CogVideoX-2B
Random78.5780.9069.2593.4796.6696.9059.7249.1688.4361.3022.95
By Qwen3-VL-8B78.4980.2271.5994.4896.8496.9059.2055.5688.4065.8023.30
By InternVL3-8B78.3580.2770.6594.5396.8997.0459.9051.3089.4961.0923.24
By VideoScore278.6780.3072.1794.6396.7896.9860.0861.0587.5963.8222.54
By FIRM-Video-8B79.6881.0374.2894.9297.2897.3360.0062.8892.3766.8623.02
T2V model: Wan2.1-T2V-1.3B
Random80.4084.1365.4994.5798.0898.9868.1857.0185.6376.1319.73
By Qwen3-VL-8B81.4384.3069.9694.5797.9098.8767.5767.9185.9474.3920.67
By InternVL3-8B81.7084.4370.7895.4698.3298.8467.7167.6186.1577.0820.30
By VideoScore281.6484.8668.7495.4598.1398.9967.5772.2683.9174.5619.98
By FIRM-Video-8B82.3684.9172.1996.5498.6899.2268.6273.4092.2381.1020.64
Figure 3: Overview of the FIRM-Video data construction framework and reward model training. Given a prompt and its generated video, FIRM-Video applies a unified check-before-score pipeline to construct fine-grained supervision for instruction following, world coherence, and perceptual quality. The resulting dimension-specific analyses and scores form FIRM-Video-90K, which is then used to train the FIRM-Video-8B reward model.
Figure 3: Overview of the FIRM-Video data construction framework and reward model training. Given a prompt and its generated video, FIRM-Video applies a unified check-before-score pipeline to construct fine-grained supervision for instruction following, world coherence, and perceptual quality. The resulting dimension-specific analyses and scores form FIRM-Video-90K, which is then used to train the FIRM-Video-8B reward model.
Table 4: Aggregation ablations on FIRM-Video-Bench. STD is computed over absolute errors; relaxed accuracy allows a one-point deviation from the expert rating. Best results within each dimension are bold.
DimensionAggregationMAE ↓STD ↓Acc. ↑Relaxed Acc. ↑SRCC ↑
IFmean0.680.700.440.880.69
Importance weighted (ours)0.680.680.430.900.69
WCmean0.860.830.370.820.61
Importance weighted (ours)0.800.830.410.840.63
Figure 4: Best-of-N scaling results on VBench using LaVie-Base as the video generation model. Performance is measured by the VBench total score as N increases.
Figure 4: Best-of-N scaling results on VBench using LaVie-Base as the video generation model. Performance is measured by the VBench total score as N increases.
Table 5: Score distributions of FIRM-Video-90K and FIRM-Video-Bench.
DatasetDimensionScore=1Score=2Score=3Score=4Score 5Total
FIRM-Video-90KIF5,104 (17.4%)8,900 (30.3%)5,551 (18.9%)3,806 (13.0%)5,987 (20.4%)29,348
WC4,795 (16.3%)11,129 (37.9%)5,649 (19.2%)2,510 (8.6%)5,265 (17.9%)29,348
PQ874 (3.0%)3,749 (12.8%)4,257 (14.5%)18,331 (62.5%)2,137 (7.3%)29,348
All10,77323,77815,45724,64713,38988,044
FIRM-Video-BenchIF28 (11.2%)73 (29.2%)63 (25.2%)59 (23.6%)27 (10.8%)250
WC48 (19.2%)76 (30.4%)48 (19.2%)44 (17.6%)34 (13.6%)250
PQ11 (4.4%)54 (21.6%)69 (27.6%)59 (23.6%)57 (22.8%)250
All87203180162118750
Figure 5: Human annotation interface for FIRM-Video-Bench. Annotators rate instruction following, perceptual quality, and world coherence on a five-point scale.
Figure 5: Human annotation interface for FIRM-Video-Bench. Annotators rate instruction following, perceptual quality, and world coherence on a five-point scale.
Table 6: Dimension mapping from VBench dimensions to FIRM-Video reward dimensions. “Quality” and “Cond. Consist.” denote VBench’s Video-Quality and Video–Condition Consistency groups, respectively.
RewardVBench DimensionVBench Group
IFDynamic DegreeQuality
Overall ConsistencyCond. Consist.
Object ClassCond. Consist.
Multiple ObjectsCond. Consist.
Human ActionCond. Consist.
ColorCond. Consist.
Spatial RelationshipCond. Consist.
SceneCond. Consist.
Appearance StyleCond. Consist.
Temporal StyleCond. Consist.
WCSubject ConsistencyQuality
Background ConsistencyQuality
Motion SmoothnessQuality
PQTemporal FlickeringQuality
Aesthetic QualityQuality
Imaging QualityQuality
Figure 6: Qualitative example of FIRM-Video-8B evaluation (1)
Figure 6: Qualitative example of FIRM-Video-8B evaluation (1)
Table 7: Best-of-8 results on the video-quality subdimensions of VBench.
T2V ModelSampling StrategySubject ConsistencyBackground ConsistencyTemporal FlickeringMotion SmoothnessDynamic DegreeAesthetic QualityImaging Quality
LaVie-BaseRandom92.0997.6497.6396.6355.5664.5365.92
Qwen3-VL-8B92.6497.5697.6096.6255.5664.9366.88
InternVL3-8B92.9497.5697.5696.6961.1164.7467.02
VideoScore292.9597.8097.7696.3956.9464.2766.54
FIRM-Video-8B93.6697.9198.1296.5856.9464.1767.46
CogVideoX-2BRandom93.4796.6696.9097.2373.6158.4759.72
Qwen3-VL-8B94.4896.8496.9097.3562.5058.3059.20
InternVL3-8B94.5396.8997.0497.2462.5057.8459.90
VideoScore294.6396.7896.9897.2759.7259.3160.08
FIRM-Video-8B94.9297.2897.3397.3465.2859.1560.00
Wan2.1-T2V-1.3BRandom94.5798.0898.9898.2162.5064.3868.18
Qwen3-VL-8B94.5797.9098.8798.1066.6764.9467.57
InternVL3-8B95.4698.3298.8498.2963.8964.8767.71
VideoScore295.4598.1398.9998.4368.0665.1067.57
FIRM-Video-8B96.5498.6899.2298.4959.7265.6768.62
Figure 7: Qualitative example of FIRM-Video-8B evaluation (2)
Figure 7: Qualitative example of FIRM-Video-8B evaluation (2)
Table 8: Best-of-8 results on the video–condition consistency subdimensions of VBench.
T2V ModelSampling StrategyObject ClassMultiple ObjectsHuman ActionColorSpatial RelationshipSceneAppearance StyleTemporal StyleOverall Consistency
LaVie-BaseRandom93.1237.7392.0089.6543.3052.3323.9424.7327.20
Qwen3-VL-8B91.6150.9196.0083.1147.3252.6924.2725.2527.75
InternVL3-8B95.8150.3096.0085.8545.8151.0224.1025.0327.46
VideoScore291.0653.7393.0086.8645.7050.8023.8725.1627.33
FIRM-Video-8B90.5158.9993.0088.2051.2052.0324.4025.1927.57
CogVideoX-2BRandom76.1149.1687.0088.4361.3039.9022.9523.7724.43
Qwen3-VL-8B83.6255.5691.0088.4065.8035.8323.3023.8425.27
InternVL3-8B82.9151.3089.0089.4961.0937.7923.2423.7425.30
VideoScore284.4161.0591.0087.5963.8240.4822.5423.5425.06
FIRM-Video-8B85.2162.8890.0092.3766.8645.1323.0224.2925.13
Wan2.1-T2V-1.3BRandom74.7657.0168.0085.6376.1324.3519.7323.7423.31
Qwen3-VL-8B82.3667.9183.0085.9474.3925.5120.6723.5324.78
InternVL3-8B84.3467.6182.0086.1577.0830.3820.3023.8624.15
VideoScore279.2772.2679.0083.9174.5623.8419.9823.7323.88
FIRM-Video-8B80.2273.4083.0092.2381.1029.1420.6423.5224.57
Figure 8: Qualitative example of FIRM-Video-8B evaluation (3)
Figure 8: Qualitative example of FIRM-Video-8B evaluation (3)

研究结果

  • 在FIRM-Video-Bench测试中,基于Qwen3-VL-8B构建的FIRM-Video-8B综合平均绝对误差为0.78,在所有测试的商业和开源模型中最低,并且在“世界常理”这一维度上误差单独最低。
  • 同一测试中,GPT-5在“指令遵循”维度误差最低为0.62,Doubao-Seed-2.0-Lite在“画面质量”维度误差最低为0.80,说明FIRM-Video-8B并非每个细分维度都排第一,但综合表现最优。
  • 在基于VBench的Best-of-8实验中,FIRM-Video-8B在三种不同视频生成模型(LaVie-Base、CogVideoX-2B、Wan2.1-T2V-1.3B)上综合分、画质分和语义一致性分均排名第一,比次优方法综合分高出0.27到1.01分,语义一致性分高出1.24到2.11分。
  • 在训练数据之外的MJ-Bench-Video测试中,FIRM-Video-8B的一个版本在“一致性”和“对齐度”维度排名第一,在整体偏好准确率上排名第二,显示出一定的泛化能力。
  • 一项对比实验显示,按重要性加权汇总检查清单结果比简单平均更准确,尤其在“世界常理”维度上,误差从0.86降到了0.80。
Figure 9: Qualitative example of FIRM-Video-8B evaluation (4)
Figure 9: Qualitative example of FIRM-Video-8B evaluation (4)

可应用场景

  • 从同一提示词生成的多个候选视频中挑出质量最好的一个再展示给用户(Best-of-N筛选)。
  • 按指令遵循度、物理合理性或画面缺陷对大批量生成的视频数据集进行自动质量筛查。
  • 作者提及但尚未实际测试的用途:将该评分模型用作强化学习或偏好优化的信号,直接用于训练和改进视频生成模型本身。
Figure 10: Qualitative example of FIRM-Video-8B evaluation (5)
Figure 10: Qualitative example of FIRM-Video-8B evaluation (5)

局限与待验证事项

  • 作者明确说明尚未将FIRM-Video-8B用于直接训练或优化视频生成模型,目前只验证了其作为评审和筛选工具的效果,其在强化学习式对齐训练中的作用还未被验证。
  • 用于测试的人工标注基准FIRM-Video-Bench仅包含250个视频、750条标注,规模相对较小。
  • 训练数据主要来自两个已有的偏好数据集及二十多个生成模型的输出,超出这一范围的视频风格或生成模型上的表现尚未直接验证。
  • 评审模型只看每个视频均匀抽取的8帧画面,而不是完整视频,可能会漏掉非常短暂或细微的时间性问题。
  • 将连续的检查清单分数映射为五档评分所用的阈值是凭经验设定的,论文没有报告这一设定对结果的敏感程度。

为什么重要

视频生成AI进步很快,但判断生成结果是否真正贴合提示词、是否符合常理、画质是否过关,目前主要靠人工或不太可靠的自动打分。一个先核实证据再打分的评审模型,能让筛选和排序生成结果的成本更低,也为将来用它来直接训练、优化视频生成模型打下基础。

本文术语

  • 奖励模型(Reward model) · 专门训练出来给另一个AI的输出结果打分的模型,用于筛选、排序或指导后续训练。
  • 先核实再打分(check-before-score) · 先针对具体条目逐一核实是或否,再只用已核实的答案汇总出最终分数的做法。
  • 指令遵循 / 世界常理 / 画面质量(IF/WC/PQ) · 本文使用的三个评测维度:视频是否符合提示词要求、是否符合物理常识、画面是否干净清晰。
  • Best-of-N采样 · 针对同一提示词生成N个候选视频,再挑出其中打分最高的一个的方法。
  • 平均绝对误差(MAE) · 模型预测分数与人类给出分数之间差距的平均值,数值越低说明越接近人类判断。

论文原文摘要(英文)

Reliable reward models are essential for text-to-video evaluation and alignment. However, the trade-off between evaluation accuracy and inference efficiency places high demands on the quality of training supervision. Existing approaches often rely on holistic judges with fixed rubrics or open-ended reasoning, leading to incomplete inspection, unfaithful justification, and entangled attribution. We introduce FIRM-Video, a unified checklist-driven data construction framework based on a check-before-score principle: construct dimension-specific checklists, verify each criterion against temporal visual evidence, and aggregate only verified decisions. For Instruction Following, FIRM-Video decomposes prompts into weighted atomic requirements; for World Coherence, it constructs prompt-calibrated, target-specific checks grounded in visible entities and actions; and for Perceptual Quality, it applies a generic taxonomy of visual defects. The verified criteria and scores are further transformed into natural-language analyses for end-to-end reward modeling. Subsequently, we construct FIRM-Video-90K with 88,044 dimension-specific instances from 29,348 videos, and introduce FIRM-Video-Bench with 750 point-wise human annotations across 250 videos. The Qwen3-VL-based FIRM-Video-8B achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.

作者 · Peiyuan Zhang

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道

图片来源: Peiyuan Zhang et al., arXiv:2608.21839, arxiv-nonexclusive