METAL LAB

영상 생성 AI를 평가하는 채점 AI가 근거를 먼저 확인하고 나서 점수를 매기게 만드는 데이터 구축 방법

arXiv:2608.218392026-08-21

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling

영상 생성 AI를 평가하는 채점 AI가 근거를 먼저 확인하고 나서 점수를 매기게 만드는 데이터 구축 방법

FIRM-Video는 영상 채점 AI를 훈련시키기 위한 데이터를 만들 때, 먼저 지시 이행·현실 타당성·화질이라는 세 항목별로 확인 가능한 예/아니오 체크리스트를 만들고, 실제 영상 장면을 대조해 각 항목을 검증한 뒤에야 점수를 계산한다. 이 과정으로 88,044개 학습 데이터(FIRM-Video-90K)와 사람이 직접 평가한 750개 검증용 데이터(FIRM-Video-Bench)가 만들어졌다. 이 데이터로 학습한 80억 파라미터 모델 FIRM-Video-8B는 기존 모델보다 사람 판단과 더 가까웠고, 여러 영상 후보 중 더 좋은 것을 골라내는 데도 더 나은 성능을 보였다.

METAL LAB 해설 도표

영상을 채점하기 전, 근거를 먼저 확인하는 FIRM-Video 데이터 구축 과정

증거 상태측정 결과가 보고됨

  1. 1. 체크리스트 만들기영상-프롬프트 쌍마다 세 종류의 체크리스트를 만든다: 지시 이행을 위한 세부 요구사항 목록, 현실 타당성을 위한 등장 요소·행동 기반 점검 목록, 화질을 위한 고정된 시각 결함 목록.
  2. 2. 근거 검증멀티모달 평가 AI가 실제 영상 프레임을 보고 체크리스트 각 항목이 충족됐는지 아닌지를 간단한 근거와 함께 판정하며, 확인할 수 없는 항목은 미충족으로 처리한다.
  3. 3. 점수 집계검증된 답변만 중요도에 따라 가중치를 주어 항목별 최종 점수 하나로 합치며, 같은 문제가 여러 항목에서 중복으로 벌점을 받지 않도록 한다.
  4. 4. 데이터셋과 벤치마크 구축이 과정을 영상 29,348개에 적용해 88,044개의 학습 데이터(FIRM-Video-90K)를 만들고, 별도의 영상 250개는 사람이 직접 평가해 750개짜리 시험용 데이터(FIRM-Video-Bench)를 만들었다.
  5. 5. 종단형 채점 모델Qwen3-VL-8B를 이 데이터로 미세조정해 FIRM-Video-8B를 만들었으며, 실제 사용할 때는 체크리스트 과정을 다시 거치지 않고 프롬프트와 영상 프레임만 보고 한 번에 점수와 설명을 예측한다.
METAL LAB이 원문을 바탕으로 재구성한 해설 도표이며, 논문 저자의 원문 figure가 아닙니다.

무엇을 했나

  1. 문제 제기: 기존 영상 평가 AI들은 눈에 띄는 특징만 보고 판단하거나, 이미 정한 점수에 맞춰 사후적으로 이유를 지어내거나, 서로 다른 종류의 오류(예: 물리적 오류와 화질 문제)를 뒤섞어 이중으로 벌점을 주는 문제가 있었다.
  2. 방법: 영상과 프롬프트가 주어지면, 세 가지 평가 항목(프롬프트를 잘 따랐는지, 현실 상식·물리법칙에 맞는지, 시각적으로 깨끗한지)별로 각각 체크리스트를 만들고, 실제 영상 프레임과 대조해 각 항목을 검증한 다음, 검증된 답변만 모아 최종 점수와 설명 문장을 만든다.
  3. 데이터 구축: 이 과정을 약 29,348개의 AI 생성 영상에 적용해 88,044개의 학습 데이터(FIRM-Video-90K)를 만들었고, 별도의 영상 250개는 사람 전문가가 직접 평가해 750개의 검증용 데이터(FIRM-Video-Bench)를 만들었다.
  4. 결과: 이 데이터로 Qwen3-VL-8B를 미세조정한 FIRM-Video-8B는 FIRM-Video-Bench에서 평균 오차(MAE) 0.78을 기록해, 미세조정 전 기본 모델의 1.33보다 크게 낮췄고, 여러 상용·오픈소스 모델 중 전체 평균 오차가 가장 낮았다.
  5. 결과: AI 생성 영상 8개 중 가장 좋은 것을 고르는 실험(Best-of-8)에서, FIRM-Video-8B는 VBench라는 표준 평가 지표 기준으로 세 가지 영상 생성 모델 모두에서 다른 선택 방법보다 일관되게 더 나은 영상을 골라냈다.
Figure 1: Qualitative evaluation case of FIRM-Video-8B, illustrating its improved alignment with human judgments.
Figure 1: Qualitative evaluation case of FIRM-Video-8B, illustrating its improved alignment with human judgments.
Table 1: Performance comparison on FIRM-Video-Bench across three evaluation dimensions: Instruction Following (IF), World Coherence (WC), and Perceptual Quality (PQ). The two values under Acc./Relaxed Acc. denote strict and relaxed accuracy, respectively. The row marked with † reports the performance of direct data construction pipeline and is excluded from ranking.
ModelInstruction FollowingWorld CoherencePerceptual QualityOverall
MAE↓STD↓Acc./Relaxed Acc.↑SRCC↑MAE↓STD↓Acc./Relaxed Acc.↑SRCC↑MAE↓STD↓Acc./Relaxed Acc.↑SRCC↑MAE↓STD↓Acc./Relaxed Acc.↑
Closed-source models
GPT-50.620.710.50/0.880.741.661.190.20/0.480.490.950.860.34/0.760.511.081.030.35/0.71
Gemini-3.1-Pro0.770.760.40/0.860.711.271.080.28/0.620.541.441.120.21/0.590.411.161.040.30/0.69
Doubao-Seed-2.0-Lite0.800.830.41/0.840.671.561.240.24/0.530.440.800.750.38/0.840.551.051.030.35/0.73
Open-source models
InternVL3-8B1.260.950.24/0.600.572.101.320.15/0.360.221.291.000.24/0.610.221.561.170.21/0.52
InternVL3-38B0.820.830.40/0.820.652.111.320.15/0.350.151.331.080.27/0.580.201.421.210.27/0.58
Qwen3-VL-8B0.930.890.34/0.800.601.791.350.22/0.460.291.281.040.28/0.600.341.331.160.28/0.62
Qwen3-VL-30B-A3B0.950.870.32/0.800.621.901.310.19/0.400.221.361.090.28/0.550.371.401.170.27/0.58
Qwen3-VL-235B-A22B0.930.870.33/0.800.651.631.300.24/0.520.311.271.040.28/0.600.401.271.120.29/0.64
Our methods
FIRM-Video Data Pipeline†0.680.680.43/0.900.690.800.830.41/0.840.630.740.700.40/0.860.610.730.740.41/0.87
FIRM-Video-8B (InternVL3-8B)0.690.730.45/0.870.670.960.950.37/0.760.490.900.770.32/0.800.350.850.830.38/0.81
FIRM-Video-8B (Qwen3-VL-8B)0.650.770.50/0.880.690.860.880.40/0.800.530.820.740.36/0.830.510.780.800.42/0.84
Figure 2: Statistics of the FIRM-Video-90K dataset. Left: distributions across evaluation dimensions and score ranges. Right: distributions of video duration and prompt length.
Figure 2: Statistics of the FIRM-Video-90K dataset. Left: distributions across evaluation dimensions and score ranges. Right: distributions of video duration and prompt length.
Table 3: Best-of-N performance comparison of FIRM-Video-8B and other sampling strategies on VBench. The best and second-best results under each T2V generator are highlighted in bold and underlined, respectively.
Sampling StrategyTotal ScoreQuality ScoreSemantic ScoreVideo QualityVideo–Condition Consistency
Subject ConsistencyBackground ConsistencyTemporal FlickeringImaging QualityMultiple ObjectsColorSpatial RelationshipAppearance Style
T2V model: LaVie-Base
Random79.4281.3471.7792.0997.6497.6365.9237.7389.6543.3023.94
By Qwen3-VL-8B80.0381.6173.7392.6497.5697.6066.8850.9183.1147.3224.27
By InternVL3-8B80.4582.1073.8292.9497.5697.5667.0250.3085.8545.8124.10
By VideoScore279.9581.6173.3392.9597.8097.7666.5453.7386.8645.7023.87
By FIRM-Video-8B80.7282.1475.0693.6697.9198.1267.4658.9988.2051.2024.40
T2V model: CogVideoX-2B
Random78.5780.9069.2593.4796.6696.9059.7249.1688.4361.3022.95
By Qwen3-VL-8B78.4980.2271.5994.4896.8496.9059.2055.5688.4065.8023.30
By InternVL3-8B78.3580.2770.6594.5396.8997.0459.9051.3089.4961.0923.24
By VideoScore278.6780.3072.1794.6396.7896.9860.0861.0587.5963.8222.54
By FIRM-Video-8B79.6881.0374.2894.9297.2897.3360.0062.8892.3766.8623.02
T2V model: Wan2.1-T2V-1.3B
Random80.4084.1365.4994.5798.0898.9868.1857.0185.6376.1319.73
By Qwen3-VL-8B81.4384.3069.9694.5797.9098.8767.5767.9185.9474.3920.67
By InternVL3-8B81.7084.4370.7895.4698.3298.8467.7167.6186.1577.0820.30
By VideoScore281.6484.8668.7495.4598.1398.9967.5772.2683.9174.5619.98
By FIRM-Video-8B82.3684.9172.1996.5498.6899.2268.6273.4092.2381.1020.64
Figure 3: Overview of the FIRM-Video data construction framework and reward model training. Given a prompt and its generated video, FIRM-Video applies a unified check-before-score pipeline to construct fine-grained supervision for instruction following, world coherence, and perceptual quality. The resulting dimension-specific analyses and scores form FIRM-Video-90K, which is then used to train the FIRM-Video-8B reward model.
Figure 3: Overview of the FIRM-Video data construction framework and reward model training. Given a prompt and its generated video, FIRM-Video applies a unified check-before-score pipeline to construct fine-grained supervision for instruction following, world coherence, and perceptual quality. The resulting dimension-specific analyses and scores form FIRM-Video-90K, which is then used to train the FIRM-Video-8B reward model.
Table 4: Aggregation ablations on FIRM-Video-Bench. STD is computed over absolute errors; relaxed accuracy allows a one-point deviation from the expert rating. Best results within each dimension are bold.
DimensionAggregationMAE ↓STD ↓Acc. ↑Relaxed Acc. ↑SRCC ↑
IFmean0.680.700.440.880.69
Importance weighted (ours)0.680.680.430.900.69
WCmean0.860.830.370.820.61
Importance weighted (ours)0.800.830.410.840.63
Figure 4: Best-of-N scaling results on VBench using LaVie-Base as the video generation model. Performance is measured by the VBench total score as N increases.
Figure 4: Best-of-N scaling results on VBench using LaVie-Base as the video generation model. Performance is measured by the VBench total score as N increases.
Table 5: Score distributions of FIRM-Video-90K and FIRM-Video-Bench.
DatasetDimensionScore=1Score=2Score=3Score=4Score 5Total
FIRM-Video-90KIF5,104 (17.4%)8,900 (30.3%)5,551 (18.9%)3,806 (13.0%)5,987 (20.4%)29,348
WC4,795 (16.3%)11,129 (37.9%)5,649 (19.2%)2,510 (8.6%)5,265 (17.9%)29,348
PQ874 (3.0%)3,749 (12.8%)4,257 (14.5%)18,331 (62.5%)2,137 (7.3%)29,348
All10,77323,77815,45724,64713,38988,044
FIRM-Video-BenchIF28 (11.2%)73 (29.2%)63 (25.2%)59 (23.6%)27 (10.8%)250
WC48 (19.2%)76 (30.4%)48 (19.2%)44 (17.6%)34 (13.6%)250
PQ11 (4.4%)54 (21.6%)69 (27.6%)59 (23.6%)57 (22.8%)250
All87203180162118750
Figure 5: Human annotation interface for FIRM-Video-Bench. Annotators rate instruction following, perceptual quality, and world coherence on a five-point scale.
Figure 5: Human annotation interface for FIRM-Video-Bench. Annotators rate instruction following, perceptual quality, and world coherence on a five-point scale.
Table 6: Dimension mapping from VBench dimensions to FIRM-Video reward dimensions. “Quality” and “Cond. Consist.” denote VBench’s Video-Quality and Video–Condition Consistency groups, respectively.
RewardVBench DimensionVBench Group
IFDynamic DegreeQuality
Overall ConsistencyCond. Consist.
Object ClassCond. Consist.
Multiple ObjectsCond. Consist.
Human ActionCond. Consist.
ColorCond. Consist.
Spatial RelationshipCond. Consist.
SceneCond. Consist.
Appearance StyleCond. Consist.
Temporal StyleCond. Consist.
WCSubject ConsistencyQuality
Background ConsistencyQuality
Motion SmoothnessQuality
PQTemporal FlickeringQuality
Aesthetic QualityQuality
Imaging QualityQuality
Figure 6: Qualitative example of FIRM-Video-8B evaluation (1)
Figure 6: Qualitative example of FIRM-Video-8B evaluation (1)
Table 7: Best-of-8 results on the video-quality subdimensions of VBench.
T2V ModelSampling StrategySubject ConsistencyBackground ConsistencyTemporal FlickeringMotion SmoothnessDynamic DegreeAesthetic QualityImaging Quality
LaVie-BaseRandom92.0997.6497.6396.6355.5664.5365.92
Qwen3-VL-8B92.6497.5697.6096.6255.5664.9366.88
InternVL3-8B92.9497.5697.5696.6961.1164.7467.02
VideoScore292.9597.8097.7696.3956.9464.2766.54
FIRM-Video-8B93.6697.9198.1296.5856.9464.1767.46
CogVideoX-2BRandom93.4796.6696.9097.2373.6158.4759.72
Qwen3-VL-8B94.4896.8496.9097.3562.5058.3059.20
InternVL3-8B94.5396.8997.0497.2462.5057.8459.90
VideoScore294.6396.7896.9897.2759.7259.3160.08
FIRM-Video-8B94.9297.2897.3397.3465.2859.1560.00
Wan2.1-T2V-1.3BRandom94.5798.0898.9898.2162.5064.3868.18
Qwen3-VL-8B94.5797.9098.8798.1066.6764.9467.57
InternVL3-8B95.4698.3298.8498.2963.8964.8767.71
VideoScore295.4598.1398.9998.4368.0665.1067.57
FIRM-Video-8B96.5498.6899.2298.4959.7265.6768.62
Figure 7: Qualitative example of FIRM-Video-8B evaluation (2)
Figure 7: Qualitative example of FIRM-Video-8B evaluation (2)
Table 8: Best-of-8 results on the video–condition consistency subdimensions of VBench.
T2V ModelSampling StrategyObject ClassMultiple ObjectsHuman ActionColorSpatial RelationshipSceneAppearance StyleTemporal StyleOverall Consistency
LaVie-BaseRandom93.1237.7392.0089.6543.3052.3323.9424.7327.20
Qwen3-VL-8B91.6150.9196.0083.1147.3252.6924.2725.2527.75
InternVL3-8B95.8150.3096.0085.8545.8151.0224.1025.0327.46
VideoScore291.0653.7393.0086.8645.7050.8023.8725.1627.33
FIRM-Video-8B90.5158.9993.0088.2051.2052.0324.4025.1927.57
CogVideoX-2BRandom76.1149.1687.0088.4361.3039.9022.9523.7724.43
Qwen3-VL-8B83.6255.5691.0088.4065.8035.8323.3023.8425.27
InternVL3-8B82.9151.3089.0089.4961.0937.7923.2423.7425.30
VideoScore284.4161.0591.0087.5963.8240.4822.5423.5425.06
FIRM-Video-8B85.2162.8890.0092.3766.8645.1323.0224.2925.13
Wan2.1-T2V-1.3BRandom74.7657.0168.0085.6376.1324.3519.7323.7423.31
Qwen3-VL-8B82.3667.9183.0085.9474.3925.5120.6723.5324.78
InternVL3-8B84.3467.6182.0086.1577.0830.3820.3023.8624.15
VideoScore279.2772.2679.0083.9174.5623.8419.9823.7323.88
FIRM-Video-8B80.2273.4083.0092.2381.1029.1420.6423.5224.57
Figure 8: Qualitative example of FIRM-Video-8B evaluation (3)
Figure 8: Qualitative example of FIRM-Video-8B evaluation (3)

실제로 확인된 결과

  • FIRM-Video-Bench에서 Qwen3-VL-8B를 기반으로 만든 FIRM-Video-8B가 전체 평균 오차 0.78로, 시험한 상용·오픈소스 모델 전체 중 가장 낮았고, 현실 타당성(World Coherence) 항목에서는 단독으로 가장 낮은 오차를 기록했다.
  • 같은 시험에서 지시 이행 항목은 GPT-5가 오차 0.62로 가장 낮았고, 화질 항목은 Doubao-Seed-2.0-Lite가 오차 0.80으로 가장 낮아, FIRM-Video-8B가 모든 세부 항목에서 1등은 아니었지만 종합 오차는 가장 낮았다.
  • VBench를 이용한 Best-of-8 실험에서 FIRM-Video-8B는 시험한 세 가지 영상 생성 모델(LaVie-Base, CogVideoX-2B, Wan2.1-T2V-1.3B) 전부에서 종합점수·화질점수·의미일치점수가 가장 높았으며, 그다음으로 좋은 방법보다 종합점수는 0.27~1.01점, 의미일치점수는 1.24~2.11점 더 높았다.
  • 훈련 데이터와 다른 종류의 데이터인 MJ-Bench-Video 시험에서도 FIRM-Video-8B의 한 변형이 지시 일치·일관성 항목에서 1위, 전체 선호도 정확도에서 2위를 기록해 어느 정도 일반화 가능성을 보였다.
  • 체크리스트 항목의 중요도에 가중치를 주는 방식이 단순 평균보다 정확했고, 특히 현실 타당성 항목에서는 오차를 0.86에서 0.80으로 줄였다는 비교 실험 결과가 나왔다.
Figure 9: Qualitative example of FIRM-Video-8B evaluation (4)
Figure 9: Qualitative example of FIRM-Video-8B evaluation (4)

어디에 쓸 수 있나

  • 같은 프롬프트로 여러 개 생성된 영상 중 가장 좋은 하나를 골라 사용자에게 보여주는 Best-of-N 필터링 작업.
  • 대량으로 생성된 영상 데이터셋을 지시 이행도·물리적 타당성·시각적 결함 기준으로 자동 걸러내는 품질 관리 작업.
  • 저자들이 언급하지만 아직 시험하지 않은 용도로, 영상 생성 모델을 강화학습이나 선호도 최적화로 직접 훈련시킬 때 점수 신호로 활용하는 것.
Figure 10: Qualitative example of FIRM-Video-8B evaluation (5)
Figure 10: Qualitative example of FIRM-Video-8B evaluation (5)

한계와 남은 검증

  • 저자들은 FIRM-Video-8B를 영상 생성 모델 훈련에 직접 사용해 본 적이 없다고 명시했으며, 지금까지는 평가와 선별 용도로만 검증되었기 때문에 강화학습 기반 정렬에서의 효과는 아직 확인되지 않았다.
  • 검증에 쓰인 사람 평가 벤치마크(FIRM-Video-Bench)는 영상 250개, 주석 750개로 규모가 비교적 작다.
  • 훈련 데이터는 주로 두 개의 기존 선호도 데이터셋과 20여 개의 생성 모델에서 가져온 것이라, 이 범위를 벗어난 영상 스타일이나 생성 모델에 대한 성능은 직접 확인되지 않았다.
  • 모델은 영상 전체가 아니라 균일하게 뽑은 8개 프레임만 보고 판단하기 때문에, 매우 짧거나 세밀한 시간적 문제는 놓칠 수 있다.
  • 체크리스트 점수를 5단계 등급으로 바꾸는 기준값은 경험적으로 정한 것이며, 이 기준을 바꿨을 때 결과가 얼마나 민감하게 달라지는지는 논문에 나와 있지 않다.

왜 중요한가

영상 생성 AI는 빠르게 발전하고 있지만, 그 결과물이 실제로 좋은지, 지시를 잘 따랐는지, 물리적으로 말이 되는지를 신뢰성 있게 판단하는 일은 여전히 사람이 직접 하거나 부정확한 자동 채점에 맡겨져 있다. 근거를 먼저 검증하는 채점 AI가 있으면 영상 데이터를 걸러내고 순위를 매기는 작업이 더 저렴해지고, 나중에는 영상 생성 모델 자체를 개선하는 훈련 신호로도 쓸 수 있다.

이 논문의 용어

  • 리워드 모델(Reward model) · 다른 AI가 만든 결과물에 점수를 매기도록 훈련된 AI로, 걸러내기·순위 매기기·추가 훈련에 쓰인다.
  • 체크리스트 기반 검증(check-before-score) · 세부 항목을 하나씩 예/아니오로 먼저 확인하고, 확인된 답만 모아서 최종 점수를 내는 방식.
  • 지시 이행 / 현실 타당성 / 화질 (IF/WC/PQ) · 이 논문이 쓰는 세 가지 평가 항목: 영상이 프롬프트를 따랐는지, 물리·상식적으로 말이 되는지, 화면이 깨끗한지.
  • Best-of-N 샘플링 · 같은 프롬프트로 N개의 영상을 만들고, 그중 점수가 가장 높은 하나를 골라내는 방식.
  • 평균절대오차(MAE) · 모델이 예측한 점수와 사람이 준 점수 사이의 차이를 평균낸 값으로, 낮을수록 사람 판단과 가깝다는 뜻.

저자 · Peiyuan Zhang

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사

그림 출처: Peiyuan Zhang et al., arXiv:2608.21839, arxiv-nonexclusive