Figure 2: Data collection and annotation pipeline: acquisition, verification, annotation, and benchmark output.
TABLE II: Source distribution (n=545 unique images).
Source Type
Count
%
Fact-checking organisations
469
86.1%
Social media archives
49
9.0%
Online news portals
27
5.0%
Total
545
100.0%
Figure 3: Five multimodal architectures organised by visual encoder, text encoder, and fusion mechanism. Double border indicates the best-performing multimodal model (ResNet-50+mBERT; 94.65% Macro-F1).
TABLE III: Typology distribution (n=545 OOC samples).
Typology
Count
% of OOC
Fabricated
299
54.9%
Miscaptioned
136
25.0%
Temporal mismatch
56
10.3%
Geographic mismatch
44
8.1%
Identity mismatch
10
1.8%
Total OOC
545
100.0%
Figure 4: Macro-F1 vs. training fraction with standard deviation bands over 5 seeds.
TABLE IV: Dataset statistics.
Attribute
Category
Count (%)
Label
Pristine
545 (50.0%)
OOC
545 (50.0%)
Language
Nepali
856 (78.5%)
English
158 (14.5%)
Code-switched
76 0(7.0%)
Split
Train
754 (69.2%)
Validation
108 0(9.9%)
Test
228 (20.9%)
Typology IAA (Cohen’s κ)
0.84
Binary IAA, non-FC (Cohen’s κ)
0.81
Figure 5: Precision–Recall and ROC curves for all five models on test split (n=228), best seed per model.
TABLE V: Leakage validation: Δ=Accrandom−Acccluster (mean over 5 seeds). All |Δ|<1% indicates no leakage.
Out-of-context (OOC) misinformation pairs authentic images with misleading captions to construct false narratives without image manipulation, making detection a problem of multimodal alignment rather than image forensics. Despite the prevalence and consequences of OOC misinformation in Nepal, no public benchmark exists for Nepali. We introduce NepOOC, the first publicly available Nepali-dominant multilingual OOC benchmark, comprising 1,090 image-caption pairs (545 pristine, 545 OOC) annotated across five typologies (fabricated, miscaptioned, temporal mismatch, geographic mismatch, identity mismatch) with inter-annotator agreement kappa = 0.84. Systematic evaluation of five multimodal architectures alongside text-only and image-only baselines reveals that caption semantics appear sufficient for strong performance at the current dataset scale. A text-only mBERT model achieves 94.65+/-0.20% Macro-F1, statistically equivalent to the best multimodal system (ResNet-50+mBERT, 94.65+/-0.20%; McNemar median p = 1.000, 0/5 seeds significant at alpha = 0.05). Image-only models perform near chance (33-50%), while training-size scaling suggests that dataset expansion is a more direct path to progress than architectural sophistication or regional specialisation.