NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
AI가 몇 시간짜리 일본 영상을 보고 '분위기 파악'까지 할 수 있을까 테스트하는 벤치마크
NARU는 155개의 일본어 장편 영상(총 146.8시간)에 기반한 1,481개의 질문으로, AI 모델이 긴 이야기의 흐름을 추적하는 능력과 일본 특유의 암묵적 문화 코드를 읽어내는 능력을 함께 평가한다. 사람이 처음부터 끝까지 수작업으로 주석을 달기엔 너무 길고 어려운 영상을, AI가 먼저 구간별로 분석하고 사람 전문가 68명이 검증하는 방식으로 벤치마크를 만들었다. 평가 결과 상용 최상위 모델도 문화적 함의를 읽는 문제에서 크게 흔들렸고, 오픈소스 모델은 전반적으로 훨씬 뒤처졌다.
무엇을 했나
- 30분에서 4시간에 달하는 일본어 장편 영상 155개(총 146.8시간)를 모아, 등장인물 변화나 줄거리 전개를 묻는 '서사 이해' 문제와 손님에게 커피를 다시 권하는 것이 사실은 '이제 그만 돌아가라'는 뜻일 수 있다는 식의 '문화적 함의 이해' 문제를 합쳐 1,481개 만들었다.
- 영상을 5분 단위로 잘라 AI가 순차적으로 요약·인물·사건 정보를 뽑아내고, 이전 구간 내용을 참고해 이어지는 흐름을 놓치지 않도록 하는 계층적 주석 파이프라인을 사용했다. 이렇게 만든 초안 질문은 정답을 유추할 수 있는 '지름길'이 없는지 반복 점검한 뒤, 일본어 원어민 검증자 68명이 두 단계에 걸쳐 최종 확인했다.
- Gemini-3-Flash가 76.2%로 가장 높은 정답률을 보였고 Gemini-3-Pro(70.0%), Gemini-2.5-Flash(51.4%)가 뒤를 이었으며, 오픈소스 모델들은 29.6~39.8%에 머물렀다.
- 상위 모델일수록 인물 파악보다 영상 전체를 관통하는 주제 파악(N.4)에서 더 어려움을 겪었고, 반대로 하위 모델은 인물이 누구인지조차 놓치는 경우가 많았다.
- 화면에 보이는 정보량을 8프레임에서 128프레임으로 늘리면 서사 이해 정확도는 뚜렷이 좋아졌지만(최대 20.5%포인트 상승), 문화적 이해 정확도는 거의 개선되지 않거나 오히려 떨어지기도 해, 문화 이해는 단순히 많이 보는 것보다 배경지식과 추론력에 더 좌우된다는 점을 보여줬다.

| Level | Task | Type of Evidence | Code | # |
|---|---|---|---|---|
| Narrative (N, 745) | Character/Entity Evolution | A character/entity across segments | N.1 | 185 |
| Sequential/Topical Flow | Events or topics over time | N.2 | 187 | |
| Plot/Conflict Progression | A causal thread or conflict | N.3 | 186 | |
| Idea/Thematic Development | Motifs, claims, or narrative cues | N.4 | 187 | |
| Cultural (C, 736) | Aizuchi (Conversational Mechanics) | Backchannels and response timing | C.1 | 143 |
| Kuuki wo Yomu (Situational Awareness) | Social atmosphere or implicit norms | C.2 | 147 | |
| Subtext Interpretation | Surface utterance plus context | C.3 | 148 | |
| Cultural Context Recognition | Culturally specific references | C.4 | 149 | |
| Sentiment Analysis | Verbal, visual, and social cues | C.5 | 149 |

| Model | Sampling Rate | N.1 | N.2 | N.3 | N.4 | Narr. Avg | C.1 | C.2 | C.3 | C.4 | C.5 | Cult. Avg | Overall |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Gemini-3-Flash | 0.25 FPS | 83.8 | 90.9 | 83.9 | 78.1 | 84.2 | 71.3 | 69.4 | 57.4 | 73.2 | 69.8 | 68.2 | 76.2 |
| Gemini-3-Pro | 0.25 FPS | 74.6 | 78.6 | 76.3 | 66.8 | 74.1 | 60.1 | 68.7 | 64.2 | 69.8 | 67.1 | 66.0 | 70.0 |
| Gemini-2.5-Flash | 0.25 FPS | 51.3 | 63.1 | 58.1 | 50.8 | 55.8 | 37.8 | 49.7 | 48.6 | 50.3 | 48.3 | 46.9 | 51.4 |
| Qwen3.5-9B | 128 Frames | 34.6 | 48.7 | 34.9 | 34.2 | 38.1 | 41.3 | 49.0 | 34.5 | 40.9 | 41.6 | 41.4 | 39.8 |
| Qwen3VL-8B | 0.25 FPS | 35.7 | 46.5 | 39.8 | 35.8 | 39.5 | 32.2 | 40.1 | 33.1 | 38.9 | 32.9 | 35.4 | 37.4 |
| Qwen2.5VL-7B | 128 Frames | 25.4 | 42.2 | 30.6 | 26.2 | 31.1 | 22.4 | 40.8 | 25.7 | 24.2 | 28.2 | 28.2 | 29.7 |
| MiniCPM-o-2.6 | 128 Frames | 23.2 | 41.2 | 27.4 | 31.6 | 30.8 | 25.9 | 37.4 | 27.7 | 22.1 | 28.2 | 28.3 | 29.6 |
| InternVL3.5 | 64 Frames | 20.0 | 36.4 | 25.8 | 31.0 | 28.3 | 32.2 | 41.5 | 31.8 | 35.6 | 32.2 | 34.6 | 31.5 |

| Model | N.1 | N.2 | N.3 | N.4 | C.1 | C.2 | C.3 | C.4 | C.5 | Avg |
|---|---|---|---|---|---|---|---|---|---|---|
| Gemini-3-Flash | 0.78 | 0.66 | 0.72 | 0.75 | 0.69 | 0.87 | 0.93 | 0.85 | 0.80 | 0.78 |
| Gemini-3-Pro | 0.77 | 0.61 | 0.71 | 0.69 | 0.65 | 0.87 | 0.88 | 0.80 | 0.72 | 0.75 |
| Gemini-2.5-Flash | 0.62 | 0.50 | 0.58 | 0.65 | 0.56 | 0.77 | 0.84 | 0.68 | 0.73 | 0.66 |
| Qwen3.5-9B | 0.54 | 0.49 | 0.47 | 0.50 | 0.52 | 0.66 | 0.65 | 0.60 | 0.59 | 0.56 |
| Qwen3-VL-8B | 0.39 | 0.23 | 0.34 | 0.36 | 0.47 | 0.64 | 0.62 | 0.39 | 0.53 | 0.44 |
| Qwen2.5-VL-7B | 0.35 | 0.24 | 0.27 | 0.28 | 0.43 | 0.47 | 0.40 | 0.28 | 0.40 | 0.35 |
| MiniCPM-o-2.6 | 0.21 | 0.12 | 0.15 | 0.16 | 0.23 | 0.26 | 0.30 | 0.15 | 0.31 | 0.21 |
| InternVL3.5 | 0.40 | 0.14 | 0.27 | 0.35 | 0.43 | 0.49 | 0.47 | 0.39 | 0.51 | 0.38 |
왜 중요한가
긴 영상을 요약하거나 문화적으로 민감한 콘텐츠를 자동으로 검토하는 서비스를 만들려면, AI가 단순히 장면을 나열하는 게 아니라 이야기의 맥락과 사회적 뉘앙스를 함께 읽어야 한다. NARU는 지금의 AI가 이런 '고맥락' 이해에서 어디가 약한지 구체적으로 드러내, 다음 세대 영상 이해 AI가 무엇을 개선해야 하는지 보여준다.
이 논문의 용어
- MLLM · 이미지·영상·텍스트를 함께 이해하도록 학습된 멀티모달 대형언어모델
- 쿠우키오요무(空気を読む) · 말로 하지 않은 사회적 분위기나 눈치를 읽는 일본 특유의 소통 방식
- 아이즈치(相槌) · '응', '그렇구나' 같은 짧은 맞장구로, 동의보다는 듣고 있다는 신호에 가까움
- 타테마에·혼네(建前·本音) · 겉으로 드러내는 공식적 태도(타테마에)와 속으로 품은 진짜 의도(혼네)의 구분
- FActScore 리콜 · 정답에 포함된 세부 사실들 중 모델 답변이 실제로 맞춘 비율로 채점하는 방식
논문 원문 초록 (영문)
Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks rarely evaluate these capabilities jointly, particularly in high-context, non-English media. To address this gap, we introduce NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video. NARU consists of 1,481 questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. To construct the benchmark at this scale, we propose a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal. The construction process includes two native-speaker verification stages involving 68 annotators. Evaluations across eight model configurations reveal substantial limitations in both long-range narrative integration and culturally grounded reasoning. By exposing these persistent gaps, NARU offers a systematic testing ground for developing MLLMs capable of reliably interpreting long-form, high-context video.
arXiv에서 원문 보기최신 논문
- Optimal Skill Selection for LLM Agents with Provable Bicriteria GuaranteesAI 에이전트에게 어떤 '스킬 문서'를 몇 개나 줘야 잘 작동하는지, 수학적으로 최적해를 보장하며 골라주는 방법
- Specification-delta-driven data governance: an empirical study of the {\guillemotleft}spec-delta{\guillemotright} as the unit of change in lakehouse data platforms데이터 플랫폼 변경도 코드처럼 '설계도 조각'을 붙여서 검토하면 어떨까: 실험 설계 논문
- Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction텍스트가 빠지거나 망가져도, AI가 그 자리를 대신할 '가짜 텍스트'를 한 번에 만들지 않고 여러 번 고쳐가며 감정을 더 정확히 읽어낸다
- LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment모델을 통째로 학습하지 않고도, 학습 초반 몇 걸음의 기울기를 미리 훔쳐봐서 LoRA를 더 똑똑하게 초기화하는 법
- Reliable Financial Named Entity Recognition under Domain Shift금융 AI가 서류체 문장에서 배운 자신감은 트위터로 가면 거짓말이 된다
- Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM AgentsAI 비서가 뭘 기억할지 결정할 때, '물어봐야 할 순간'에 되레 세상에 확인하고 넘어간다
- Stopping and Routing LLM Judge PanelsAI 채점관을 몇 명 불러야 하는지, 언제 멈춰야 하는지 정하는 방법
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving논문 속 시연이 아니라 실제 서비스에 넣을 수 있는 희소 어텐션 만들기
METAL LAB 최신 기사
그림 출처: Yuheng Huang et al., arXiv:2608.13210, arxiv-nonexclusive