NepOOC-M: Bilingual Nepali-English Benchmark and Comparative Analysis of Multimodal Architectures for OOC Detection
arXiv:2608.192122026-08-21
네팔어 가짜뉴스 탐지, 사진 안 보고 문장만 읽어도 정답률 94%
연구팀이 진짜 사진에 거짓 설명을 붙이는 '탈맥락(OOC)' 가짜뉴스를 네팔어로 처음 벤치마크로 만들었다. 이미지-텍스트 조합 모델 5종과 텍스트만 보는 모델, 이미지만 보는 모델을 비교한 결과, 텍스트만 읽는 모델이 이미지까지 보는 최고 모델과 통계적으로 똑같은 성능을 냈다. 반대로 이미지만 보는 모델은 거의 찍는 수준에 그쳤다.
무엇을 했나
네팔에서 실제로 퍼진 가짜뉴스 사건들을 모아 진짜 사진 545장과 그에 딸린 정상 설명, 거짓 설명 쌍 545개씩 총 1,090개 이미지-캡션 쌍을 만들고 5가지 유형(조작 서술, 오설명, 시간 불일치, 지역 불일치, 인물 불일치)으로 사람이 직접 분류했다
두 명 이상의 라벨링 담당자 간 일치도가 카파 0.84로 매우 높아 데이터 품질을 신뢰할 수 있음을 보였다
이미지 인식 모델(ResNet-50 등)과 다국어 문장 이해 모델(mBERT 등)을 다섯 가지 방식으로 조합한 모델, 그리고 텍스트만 쓰는 모델과 이미지만 쓰는 모델을 학습시켜 비교했다
가장 성능이 좋았던 이미지+텍스트 조합 모델(ResNet-50+mBERT)과 텍스트만 쓰는 mBERT 모델이 똑같이 94.65% 성능을 냈고, 통계 검정(McNemar 검정)으로도 두 모델이 사실상 같은 실수를 한다는 것이 확인됐다
이미지만 보는 모델은 33~50% 정확도로 거의 무작위 추측 수준이었고, 학습 데이터를 늘렸을 때 성능이 꾸준히 오르는 것을 보면 모델 구조를 복잡하게 만드는 것보다 데이터를 더 모으는 것이 발전에 더 중요하다는 결론을 냈다
Figure 1: Representative pristine and OOC pairs across five typologies. Red boxes highlight manipulated phrases.
Figure 2: Data collection and annotation pipeline: acquisition, verification, annotation, and benchmark output.
TABLE II: Source distribution (n=545 unique images).
Source Type
Count
%
Fact-checking organisations
469
86.1%
Social media archives
49
9.0%
Online news portals
27
5.0%
Total
545
100.0%
Figure 3: Five multimodal architectures organised by visual encoder, text encoder, and fusion mechanism. Double border indicates the best-performing multimodal model (ResNet-50+mBERT; 94.65% Macro-F1).
TABLE III: Typology distribution (n=545 OOC samples).
Typology
Count
% of OOC
Fabricated
299
54.9%
Miscaptioned
136
25.0%
Temporal mismatch
56
10.3%
Geographic mismatch
44
8.1%
Identity mismatch
10
1.8%
Total OOC
545
100.0%
Figure 4: Macro-F1 vs. training fraction with standard deviation bands over 5 seeds.
TABLE IV: Dataset statistics.
Attribute
Category
Count (%)
Label
Pristine
545 (50.0%)
OOC
545 (50.0%)
Language
Nepali
856 (78.5%)
English
158 (14.5%)
Code-switched
76 0(7.0%)
Split
Train
754 (69.2%)
Validation
108 0(9.9%)
Test
228 (20.9%)
Typology IAA (Cohen’s κ)
0.84
Binary IAA, non-FC (Cohen’s κ)
0.81
Figure 5: Precision–Recall and ROC curves for all five models on test split (n=228), best seed per model.
TABLE V: Leakage validation: Δ=Accrandom−Acccluster (mean over 5 seeds). All |Δ|<1% indicates no leakage.
TABLE VII: Training configuration. LR=learning rate, WD=weight decay, BS=batch size, EP=epochs, Pat.=early-stopping patience. †Adaptive LR: 1e−4 at ≤50%, 5e−5 at >50%. ‡StepLR: step 10, γ=0.5.
Model
Opt.
LR
WD
BS
EP
Pat.
Sched.
CNN+LSTM
Adam
1e−4
1e−5
32
80
10
StepLR‡
ViT+TCN
AdamW
5e−5
1e−4
32
100
10
Cos+WU
ResNet +mBERT
AdamW
vis: 1e−4 txt: 2e−5
1e−4
32
50
10
Cos+WU
CLIP
AdamW
†
0.05
8
100
12
—
ViT+MuRIL
AdamW
†
0.05
8
100
12
Cos+WU
TABLE VIII: Main results on test split (n=228, mean±std over 5 seeds).
Model
Type
Acc. (%)
Macro-F1 (%)
AUC
mBERT
Text-only
94.65±0.20
94.65±0.20
0.9697±0.0126
MuRIL
Text-only
94.38±0.35
94.38±0.35
0.9567±0.0158
ResNet-50+mBERT
Multimodal
94.65±0.20
94.65±0.20
0.9662±0.0142
ViT+MuRIL
Multimodal
93.33±0.37
93.33±0.37
0.9505±0.0057
ViT+TCN
Multimodal
92.11±1.35
92.10±1.36
0.9616±0.0064
CNN+LSTM
Multimodal
78.16±10.97
78.15±10.97
0.8548±0.1203
CLIP
Multimodal
69.39±0.72
69.00±0.72
0.7127±0.0028
TABLE IX: McNemar’s test with Yates correction (α=0.05). See text (§5.1) for comparison pairs and discordant counts.
Model A
Model B
Seed
Disc.
b
c
χ2
p
Sig.
text-only mBERT
ResNet-50+mBERT
042
1
0
1
0.000
1.000
×
123
0
0
0
0.000
1.000
×
456
0
0
0
0.000
1.000
×
789
2
2
0
0.500
0.480
×
2024
0
0
0
0.000
1.000
×
Summary
0.6 mean
median p=1.000; 0/5 sig.
ViT+MuRIL
ResNet-50+mBERT
042
1
0
1
0.000
1.000
×
123
0
0
0
0.000
1.000
×
456
3
1
2
0.000
1.000
×
789
1
0
1
0.000
1.000
×
2024
5
0
5
3.200
0.074
×
Summary
2.0 mean
median p=1.000; 0/5 sig.
text-only MuRIL
text-only mBERT
042
2
1
1
0.500
0.480
×
123
3
1
2
0.000
1.000
×
456
0
0
0
0.000
1.000
×
789
5
0
5
3.200
0.074
×
2024
1
0
1
0.000
1.000
×
Summary
2.2 mean
median p=1.000; 0/5 sig.
TABLE X: OOC-class metrics at 100% training data (mean±std over 5 seeds).
Model
Macro-F1 (%)
OOC Prec. (%)
OOC Rec. (%)
OOC F1 (%)
CNN+LSTM
78.15±10.97
79.23±11.27
76.32±11.42
77.72±11.19
ViT+TCN
92.10±1.36
92.20±2.81
92.11±1.96
92.11±1.24
CLIP
69.00±0.72
74.97±1.34
58.25±1.33
65.54±0.87
ResNet-50+mBERT
94.65±0.20
95.37±0.38
93.86±0.00†
94.61±0.19
ViT+MuRIL
93.33±0.37
94.14±1.22
92.46±1.71
93.27±0.43
TABLE XI: Scaling results: Macro-F1 (mean±std over 5 seeds per fraction).
Model
25%
50%
75%
100%
CNN+LSTM
55.22±4.80
61.65±4.99
67.67±9.51
78.15±10.97
ViT+TCN
69.25±11.31
89.12±1.05
90.87±2.57
92.10±1.36
CLIP
62.01±1.25
63.45±2.27
66.14±1.42
69.00±0.72
ResNet-50+mBERT
93.15±0.91
93.77±1.10
94.38±0.79
94.65±0.20
ViT+MuRIL
90.61±0.67
92.72±0.67
93.33±0.72
93.33±0.37
TABLE XII: Modality ablation: Macro-F1 at 100% training data (mean±std over 5 seeds).
Model
Text-only (%)
Image-only (%)
Multimodal (%)
CNN+LSTM
74.91±16.80
49.80±0.27
78.15±10.97
ViT+TCN
90.87±1.80
33.33±0.00
92.10±1.36
CLIP
67.48±1.06
49.20±0.48
69.00±0.72
ResNet+mBERT
94.65±0.20
49.00±0.96
94.65±0.20
ViT+MuRIL
93.95±0.84
49.83±0.18
93.33±0.37
TABLE XIII: Per-seed mean confusion matrices on test split (n=228; 114 Pristine, 114 OOC).
Model
Actual
Pred: Pristine
Pred: OOC
CNN+LSTM
Pristine
91.2
22.8
OOC
27.0
87.0
ViT+TCN
Pristine
105.0
9.0
OOC
9.0
105.0
CLIP
Pristine
91.8
22.2
OOC
47.6
66.4
ResNet-50+mBERT
Pristine
108.8
5.2
OOC
7.0
107.0
ViT+MuRIL
Pristine
107.4
6.6
OOC
8.6
105.4
TABLE XIV: Per-typology OOC Macro-F1 at 100% training data (mean±std over 5 seeds). n=OOC test instances per typology.
Typology
n
CNN+LSTM
ViT+TCN
CLIP
ResNet-50+mBERT
ViT+MuRIL
Fabricated
62
83.93±7.57
96.14±0.99
74.49±1.06
95.80±1.00
95.97±0.39
Miscaptioned
28
87.92±8.92
95.09±2.23
68.53±1.70
98.18±0.00†
95.89±1.37
Other Mismatches‡
24
89.01±9.50
91.43±4.21
74.81±4.42
97.79±0.00†
96.38±3.16
왜 중요한가
지금까지 가짜뉴스 탐지 연구는 영어 위주였고 네팔어처럼 자원이 부족한 언어에는 공개된 평가 데이터셋조차 없었는데, 이 연구가 그 공백을 처음으로 메웠다. 또한 '이미지와 텍스트를 같이 보면 무조건 더 잘 탐지한다'는 통념이 적어도 이 규모의 데이터에서는 사실이 아니라는 것을 보여줘, 앞으로 저자원 언어 연구에서 텍스트 단독 모델을 기본 비교 대상으로 반드시 넣어야 한다는 실무적 시사점을 준다.
이 논문의 용어
탈맥락(Out-of-Context, OOC) 가짜뉴스 · 사진 자체는 조작하지 않고 진짜 사진에 엉뚱하거나 거짓된 설명(캡션)을 붙여 사실을 왜곡하는 방식의 허위정보
mBERT · 104개 언어를 함께 학습한 다국어 문장 이해 인공지능 모델
ResNet-50 · 사진 속 특징을 추출하는 데 널리 쓰이는 이미지 인식 신경망 구조
매크로 F1 점수(Macro-F1) · 각 분류(정상/가짜)를 동등하게 취급해 정밀도와 재현율을 종합한 성능 지표
McNemar 검정 · 두 분류 모델의 정답·오답 패턴이 통계적으로 다른지 비교하는 방법
코헨의 카파(Cohen's kappa) · 여러 사람이 매긴 라벨이 우연이 아니라 실제로 얼마나 일치하는지 나타내는 지표
논문 원문 초록 (영문)
Out-of-context (OOC) misinformation pairs authentic images with misleading captions to construct false narratives without image manipulation, making detection a problem of multimodal alignment rather than image forensics. Despite the prevalence and consequences of OOC misinformation in Nepal, no public benchmark exists for Nepali. We introduce NepOOC, the first publicly available Nepali-dominant multilingual OOC benchmark, comprising 1,090 image-caption pairs (545 pristine, 545 OOC) annotated across five typologies (fabricated, miscaptioned, temporal mismatch, geographic mismatch, identity mismatch) with inter-annotator agreement kappa = 0.84. Systematic evaluation of five multimodal architectures alongside text-only and image-only baselines reveals that caption semantics appear sufficient for strong performance at the current dataset scale. A text-only mBERT model achieves 94.65+/-0.20% Macro-F1, statistically equivalent to the best multimodal system (ResNet-50+mBERT, 94.65+/-0.20%; McNemar median p = 1.000, 0/5 seeds significant at alpha = 0.05). Image-only models perform near chance (33-50%), while training-size scaling suggests that dataset expansion is a more direct path to progress than architectural sophistication or regional specialisation.