AI news and explainers at 7 AM weekdays, plus a Sunday weekly at 8Get it in your inbox›
MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
arXiv:2607.276372026-07-31
MMOOC is a 41K-question benchmark testing whether multimodal AI refuses truly unanswerable questions while still answering ones that just look confusing
Multimodal AI models (MLLMs) that look at an image and a question should refuse when the image genuinely can't support an answer, but should still answer when the question is merely misleadingly worded yet answerable. MMOOC is a 41K-plus image-question benchmark that separately defines five truly unanswerable (Out-of-Context) categories and three answerable-but-tricky (Shifted In-Context) categories to test both behaviors at once. Testing 18 open and closed models showed most still fail to balance refusal and answering under these shifted contexts.
METAL LAB explanatory visual
MMOOC structure: truly unanswerable vs. answerable-but-tricky questions
Evidence statusMeasured results reported
Step 1: Generate image-question pairsQuestions generated via Qwen3.5-122B-A10B, GPT-4o, and o1, combined with manually authored questions and Auto-Shuffle samples from MME, MMStar, OK-VQA
Step 2: Triple-model filteringGPT-4o, o1, and o3 each independently judge answerability; only samples with unanimous agreement are kept, followed by human verification
Step 4: Model evaluation18 open- and closed-source MLLMs answer yes/no, multiple-choice, and open-ended questions; accuracy/refusal rate plus reasoning quality scored by GPT-5.6, Claude Opus 5, and DeepSeek-V4-Pro as judges
Step 5: Improvement experimentsPost-training (SFT, DPO) and prompting strategies (refusal prompts, chain-of-thought) tested to see how refusal ability trades off against general performance
An explanatory diagram made by METAL LAB, not a figure supplied by the paper's authors.
What they did
Motivation: prior benchmarks mostly checked only whether models refuse unanswerable questions, or covered limited question formats and visual scenarios, overlooking cases where the context is shifted but the core question remains answerable (Shifted In-Context).
Method: each sample is first split into 'answerable given the image' vs. 'not', then subdivided into 5 Out-of-Context categories (multimodal ambiguity, visual false premises, uncertain spatial/physical context, unclear logical/symbolic, missing knowledge) and 3 Shifted In-Context categories (misleading premise, partial answerability, image-question mismatch). Questions were generated using Qwen3.5-122B-A10B, GPT-4o, and o1, filtered to keep only cases where GPT-4o, o1, and o3 agreed on answerability, then manually verified.
Scale: MMOOC contains over 41K image-question pairs spanning three question formats (yes/no, multiple-choice, open-ended VQA), eight shift types, and six visual scenarios; response correctness and reasoning quality were scored via an LLM-as-a-Judge protocol using GPT-5.6, Claude Opus 5, and DeepSeek-V4-Pro as independent judges.
Models tested: 18 MLLMs in total, including 13 open-source models (Qwen3-VL, InternVL3, Gemma-4, Llama-4-Maverick, etc.) and 5 closed-source models (GPT-4o, o1, o3, Gemini-3.1-Pro, Claude-Opus-4.6).
Key result: nearly all models scored low and inconsistently on truly unanswerable (OOC) questions -- for instance Qwen3-VL-2B scored only 5.75 (yes/no) and 8.25 (open-ended VQA) on the uncertain spatial/physical context category -- and larger model size did not reliably improve this refusal ability.
Figure 1: Comparison of conventional, refusal, and our MMOOC benchmarks. MMOOC jointly evaluates robust answering for answerable questions and appropriate refusal for truly out-of-context questions.
Table 1: Comparison of existing out-of-context evaluation benchmarks. Data Scale denotes the total number of question-answer pairs. QA Format indicates the supported question formats. Shift Types denotes the number of defined shift types in each benchmark. Visual Scenarios indicates the range of visual understanding and reasoning settings covered by each benchmark. QA Construction indicates how the question-answer pairs are constructed. Distractor Robustness indicates whether the benchmark evaluates correct answering under distracting contexts.
Figure 2: Examples of the eight MMOOC scenarios. The top row presents five Out-of-Context categories requiring refusal: Multimodal Ambiguity (MA), Visual False Premises (VFP), Uncertain Spatial & Physical Context (USPC), Unclear Logical & Symbolic (ULS), and Missing Knowledge & Background (MKB). The bottom row presents three answerable Shifted In-Context scenarios: Misleading Premise, Partial Answerability, and Image–Question Mismatch.
Table 3: Average performance on Out-of-Context tasks across various models, computed from Refusal Rate and Refusal Rationality. The OOC category abbreviations are: MA: Multimodal Ambiguity; VFP: Visual False Premises; USPC: Uncertain Spatial & Physical Context; ULS: Unclear Logical & Symbolic; and MKB: Missing Knowledge & Background. Complete detailed results are presented in the Appendix.
YesNo
MCQ
VQA
Model
MA
VFP
USPC
ULS
MKB
MA
VFP
USPC
ULS
MKB
MA
VFP
USPC
ULS
MKB
Open-source LMMs
Qwen3-VL-2B
13.75
15.75
5.75
10.50
19.00
40.00
29.25
45.50
22.75
21.50
26.00
29.50
8.25
12.00
21.75
Qwen3-VL-8B
28.00
58.75
15.75
35.50
30.25
36.50
22.00
71.50
39.50
34.75
55.25
86.00
28.50
35.00
49.00
Qwen3-VL-30B
44.75
58.00
22.50
26.75
35.50
37.25
36.25
76.75
32.75
38.50
42.25
54.75
20.00
20.00
19.25
Qwen3.5-27B
36.25
63.25
22.50
28.00
43.25
32.50
27.50
69.25
16.75
35.50
46.00
84.25
27.75
40.25
39.00
Qwen3.5-122B-A10B
30.50
65.25
26.25
28.25
32.75
28.25
32.25
67.75
35.25
23.00
46.00
85.25
26.50
42.00
41.50
LLaVA-1.5-7B
1.75
8.75
2.25
6.25
3.00
1.75
5.25
1.75
3.00
2.50
3.25
5.75
5.75
14.75
8.00
InternVL3-2B
21.25
39.50
3.75
17.00
28.50
4.50
14.75
4.25
11.25
5.00
28.25
26.50
11.50
15.50
33.75
InternVL3-8B
48.25
55.50
12.00
39.25
47.50
29.75
5.00
46.25
21.25
18.50
21.75
43.00
7.00
15.50
32.00
Gemma-4-26B
39.50
66.25
27.75
49.00
76.25
70.25
28.50
85.00
57.25
61.25
79.25
86.50
39.50
58.00
80.00
Gemma-4-31B
54.50
54.25
35.25
46.75
72.50
52.25
34.25
88.75
47.50
43.50
71.50
83.00
40.50
54.00
70.50
Llama-4-Maverick
40.00
51.25
19.00
27.75
34.00
32.25
26.25
47.25
27.00
36.25
52.25
78.25
20.75
31.75
58.00
Ministral-3-8B
46.75
60.50
26.50
31.00
61.75
41.50
31.25
54.25
27.75
44.50
54.50
59.25
33.00
37.25
58.00
Ministral-3-14B
45.25
52.75
28.50
45.75
61.50
50.25
24.50
75.50
42.00
40.50
52.50
50.75
15.00
36.50
36.50
Closed-source LMMs
Gemini-3.1-Pro
4.75
9.00
21.75
27.75
20.75
4.50
14.25
24.25
12.50
15.00
7.75
6.50
16.50
9.50
20.25
GPT-4o
43.00
59.25
30.00
33.75
53.75
51.25
24.75
64.50
34.00
34.25
43.50
58.25
21.00
54.00
57.75
o1
57.50
71.75
44.75
54.75
76.25
45.50
22.75
68.25
30.00
45.75
75.00
43.50
40.75
47.50
53.75
o3
29.75
35.50
18.00
34.25
43.25
37.75
22.25
37.50
16.00
19.75
47.75
61.00
24.75
34.50
41.25
Claude-Opus-4.6
34.25
67.75
15.25
30.50
30.75
14.25
5.25
31.75
6.75
4.75
14.75
35.00
15.50
23.50
46.50
Figure 4: Performance comparison of different models.
Table 4: Average performance on in-context (IC) shift tasks. MP: Misleading Premise, PA: Partial Answerability, and IQM: Image–Question Mismatch. Detailed results for all IC categories are provided in the Appendix.
YN
MCQ
VQA
Model
MP
PA
IQM
MP
PA
IQM
MP
PA
IQM
Open-source LMMs
Qwen3-VL-2B
86.00
70.25
71.50
78.00
67.50
77.25
63.00
36.25
82.75
Qwen3-VL-8B
90.00
72.25
75.25
84.75
75.25
90.25
83.25
58.00
87.75
Qwen3-VL-30B
88.50
79.50
78.50
85.50
82.00
95.00
76.50
55.00
82.75
Qwen3.5-27B
88.25
75.75
79.25
93.75
85.50
94.50
90.50
63.00
89.00
Qwen3.5-122B-A10B
80.25
78.25
75.75
87.25
90.50
94.50
90.25
72.25
86.75
LLaVA-1.5-7B
55.50
55.00
58.00
31.00
34.00
38.00
28.00
11.00
55.25
InternVL3-2B
76.25
63.75
59.75
64.00
68.00
70.50
52.00
41.50
73.25
InternVL3-8B
71.75
66.75
62.00
70.00
71.50
69.25
62.50
53.00
70.75
Gemma-4-26B
81.75
79.00
74.00
89.25
85.25
83.25
74.25
67.50
78.00
Gemma-4-31B
75.75
77.25
67.75
84.25
84.75
93.25
75.75
70.75
80.50
Llama-4-Maverick
76.00
74.25
71.75
84.50
86.50
85.75
77.50
58.00
80.00
Ministral-3-8B
73.25
72.75
71.00
80.25
71.75
90.00
76.75
64.25
80.50
Ministral-3-14B
75.25
81.00
72.50
83.00
78.50
79.25
65.25
55.50
73.50
Closed-source LMMs
Gemini-3.1-Pro
70.75
62.75
59.50
59.50
58.75
75.75
61.50
22.75
52.00
GPT-4o
82.50
78.25
72.00
70.50
76.25
82.00
76.75
62.50
80.25
o1
79.25
81.25
71.00
79.50
73.25
85.25
74.75
56.50
81.00
o3
78.25
77.25
63.75
75.00
75.25
79.00
71.00
44.25
85.25
Claude-Opus-4.6
88.25
82.75
72.50
58.00
45.25
40.00
21.75
32.25
35.75
Figure 6: Performance under different prompts.
Table A1: Question-only refusal performance, where models receive only the question without the associated image. Ref. denotes the refusal score. Higher values indicate a stronger tendency to identify the question as unanswerable in the absence of visual evidence.
Model
Ref.
Open-source LMMs
Qwen3-VL-2B
42.00
Qwen3-VL-8B
30.00
Qwen3-VL-30B
58.00
Qwen3.5-27B
32.00
Qwen3.5-122B-A10B
20.00
LLaVA-1.5-7B
10.00
InternVL3-2B
8.00
InternVL3-8B
24.00
Gemma-4-26B
92.00
Gemma-4-31B
76.00
Llama-4-Maverick
52.00
Ministral-3-8B
78.00
Ministral-3-14B
84.00
Closed-source LMMs
Gemini-3.1-Pro
4.00
GPT-4o
26.00
o1
78.00
o3
46.00
Claude-Opus-4.6
62.00
Figure 7: Robustness to misleading prompts, evaluated by our core metrics Rref (Refusal Rate) and Rrat (Reasoning Rationality).
Table A2: Detailed performance on the OOC YesNo tasks. Ref., Rat., and Mean denote Refusal Rate, Refusal Rationality, and their average, respectively.
MA
VFP
USPC
ULS
MKB
Model
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Open-source LMMs
Qwen3-VL-2B
10.00
17.50
13.75
8.00
23.50
15.75
2.00
9.50
5.75
6.00
15.00
10.50
14.00
24.00
19.00
Qwen3-VL-8B
16.00
40.00
28.00
48.00
69.50
58.75
8.00
23.50
15.75
26.00
45.00
35.50
16.00
44.50
30.25
Qwen3-VL-30B
36.00
53.50
44.75
46.00
70.00
58.00
12.00
33.00
22.50
18.00
35.50
26.75
28.00
43.00
35.50
Qwen3.5-27B
22.00
50.50
36.25
44.00
82.50
63.25
10.00
35.00
22.50
14.00
42.00
28.00
28.00
58.50
43.25
Qwen3.5-122B-A10B
16.00
45.00
30.50
50.00
80.50
65.25
12.00
40.50
26.25
14.00
42.50
28.25
18.00
47.50
32.75
LLaVA-1.5-7B
0.00
3.50
1.75
2.00
15.50
8.75
0.00
4.50
2.25
2.00
10.50
6.25
0.00
6.00
3.00
InternVL3-2B
18.00
24.50
21.25
30.00
49.00
39.50
0.00
7.50
3.75
12.00
22.00
17.00
22.00
35.00
28.50
InternVL3-8B
46.00
50.50
48.25
46.00
65.00
55.50
6.00
18.00
12.00
34.00
44.50
39.25
42.00
53.00
47.50
Gemma-4-26B
32.00
47.00
39.50
58.00
74.50
66.25
22.00
33.50
27.75
42.00
56.00
49.00
74.00
78.50
76.25
Gemma-4-31B
52.00
57.00
54.50
46.00
62.50
54.25
32.00
38.50
35.25
42.00
51.50
46.75
70.00
75.00
72.50
Llama-4-Maverick
32.00
48.00
40.00
42.00
60.50
51.25
12.00
26.00
19.00
18.00
37.50
27.75
24.00
44.00
34.00
Ministral-3-8B
38.00
55.50
46.75
50.00
71.00
60.50
18.00
35.00
26.50
24.00
38.00
31.00
56.00
67.50
61.75
Ministral-3-14B
34.00
56.50
45.25
42.00
63.50
52.75
22.00
35.00
28.50
38.00
53.50
45.75
58.00
65.00
61.50
Closed-source LMMs
Gemini-3.1-Pro
4.00
5.50
4.75
12.00
6.00
9.00
28.00
15.50
21.75
18.00
37.50
27.75
32.00
9.50
20.75
GPT-4o
38.00
48.00
43.00
46.00
72.50
59.25
26.00
34.00
30.00
26.00
41.50
33.75
44.00
63.50
53.75
o1
52.00
63.00
57.50
60.00
83.50
71.75
38.00
51.50
44.75
48.00
61.50
54.75
72.00
80.50
76.25
o3
26.00
33.50
29.75
27.00
43.75
35.50
12.00
24.00
18.00
26.00
42.50
34.25
36.00
50.50
43.25
Claude-Opus-4.6
26.00
42.50
34.25
60.00
75.50
67.75
4.00
26.50
15.25
22.00
39.00
30.50
20.00
41.50
30.75
Figure 8: Robustness to Gaussian noise, evaluated by accuracy (Acc) and reasoning rationality (Accrat).
Table A3: Detailed performance on the OOC MCQ tasks. Ref., Rat., and Mean denote Refusal Rate, Refusal Rationality, and their average, respectively.
MA
VFP
USPC
ULS
MKB
Model
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Open-source LMMs
Qwen3-VL-2B
36.00
44.00
40.00
22.00
36.50
29.25
44.00
47.00
45.50
16.00
29.50
22.75
16.00
27.00
21.50
Qwen3-VL-8B
32.00
41.00
36.50
14.00
30.00
22.00
70.00
73.00
71.50
34.00
45.00
39.50
28.00
41.50
34.75
Qwen3-VL-30B
32.00
42.50
37.25
28.00
44.50
36.25
76.00
77.50
76.75
26.00
39.50
32.75
26.00
51.00
38.50
Qwen3.5-27B
26.00
39.00
32.50
16.00
39.00
27.50
68.00
70.50
69.25
10.00
23.50
16.75
26.00
45.00
35.50
Qwen3.5-122B-A10B
22.00
34.50
28.25
24.00
40.50
32.25
64.00
71.50
67.75
28.00
42.50
35.25
12.00
34.00
23.00
LLaVA-1.5-7B
2.00
1.50
1.75
4.00
6.50
5.25
2.00
1.50
1.75
2.00
4.00
3.00
2.00
3.00
2.50
InternVL3-2B
4.00
5.00
4.50
12.00
17.50
14.75
4.00
4.50
4.25
10.00
12.50
11.25
4.00
6.00
5.00
InternVL3-8B
28.00
31.50
29.75
2.00
8.00
5.00
48.00
44.50
46.25
18.00
24.50
21.25
16.00
21.00
18.50
Gemma-4-26B
66.00
74.50
70.25
20.00
37.00
28.50
86.00
84.00
85.00
50.00
64.50
57.25
56.00
66.50
61.25
Gemma-4-31B
48.00
56.50
52.25
26.00
42.50
34.25
88.00
89.50
88.75
46.00
49.00
47.50
40.00
47.00
43.50
Llama-4-Maverick
26.00
38.50
32.25
16.00
36.50
26.25
38.00
57.50
47.25
20.00
34.00
27.00
32.00
40.50
36.25
Ministral-3-8B
34.00
49.00
41.50
26.00
36.50
31.25
52.00
56.50
54.25
18.00
37.50
27.75
38.00
51.00
44.50
Ministral-3-14B
44.00
56.50
50.25
14.00
35.00
24.50
76.00
75.00
75.50
34.00
50.00
42.00
30.00
51.00
40.50
Closed-source LMMs
Gemini-3.1-Pro
4.00
5.00
4.50
20.00
8.50
14.25
36.00
12.50
24.25
20.00
5.00
12.50
20.00
10.00
15.00
GPT-4o
50.00
52.50
51.25
18.00
31.50
24.75
64.00
65.00
64.50
30.00
38.00
34.00
30.00
38.50
34.25
o1
44.00
47.00
45.50
18.00
27.50
22.75
70.00
66.50
68.25
26.00
34.00
30.00
40.00
51.50
45.75
o3
38.00
37.50
37.75
20.00
24.50
22.25
38.00
37.00
37.50
14.00
18.00
16.00
18.00
21.50
19.75
Claude-Opus-4.6
10.00
18.50
14.25
2.00
8.50
5.25
20.00
43.50
31.75
4.00
9.50
6.75
0.00
9.50
4.75
Table A4: Detailed performance on the OOC VQA tasks. Ref., Rat., and Mean denote Refusal Rate, Refusal Rationality, and their average, respectively.
MA
VFP
USPC
ULS
MKB
Model
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Ref.
Rat.
Mean
Open-source LMMs
Qwen3-VL-2B
26.00
26.00
26.00
30.00
29.00
29.50
6.00
10.50
8.25
12.00
12.00
12.00
20.00
23.50
21.75
Qwen3-VL-8B
52.00
58.50
55.25
86.00
86.00
86.00
22.00
35.00
28.50
30.00
40.00
35.00
46.00
52.00
49.00
Qwen3-VL-30B
40.00
44.50
42.25
54.00
55.50
54.75
14.00
26.00
20.00
16.00
24.00
20.00
18.00
20.50
19.25
Qwen3.5-27B
38.00
54.00
46.00
84.00
84.50
84.25
14.00
41.50
27.75
38.00
42.50
40.25
34.00
44.00
39.00
Qwen3.5-122B-A10B
40.00
52.00
46.00
84.00
86.50
85.25
14.00
39.00
26.50
36.00
48.00
42.00
36.00
47.00
41.50
LLaVA-1.5-7B
2.00
4.50
3.25
6.00
5.50
5.75
4.00
7.50
5.75
14.00
15.50
14.75
6.00
10.00
8.00
InternVL3-2B
28.00
28.50
28.25
30.00
23.00
26.50
8.00
15.00
11.50
12.00
19.00
15.50
32.00
35.50
33.75
InternVL3-8B
18.00
25.50
21.75
44.00
42.00
43.00
2.00
12.00
7.00
12.00
19.00
15.50
30.00
34.00
32.00
Gemma-4-26B
78.00
80.50
79.25
88.00
85.00
86.50
34.00
45.00
39.50
54.00
62.00
58.00
78.00
82.00
80.00
Gemma-4-31B
72.00
71.00
71.50
82.00
84.00
83.00
34.00
47.00
40.50
52.00
56.00
54.00
70.00
71.00
70.50
Llama-4-Maverick
48.00
56.50
52.25
78.00
78.50
78.25
12.00
29.50
20.75
26.00
37.50
31.75
54.00
62.00
58.00
Ministral-3-8B
46.00
63.00
54.50
56.00
62.50
59.25
30.00
36.00
33.00
28.00
46.50
37.25
54.00
62.00
58.00
Ministral-3-14B
49.25
46.00
52.50
50.00
51.50
50.75
10.00
20.00
15.00
32.00
41.00
36.50
30.00
43.00
36.50
Closed-source LMMs
Gemini-3.1-Pro
8.00
7.50
7.75
10.00
3.00
6.50
22.00
11.00
16.50
14.00
5.00
9.50
30.00
10.50
20.25
GPT-4o
42.00
45.00
43.50
60.00
56.50
58.25
16.00
26.00
21.00
52.00
56.00
54.00
54.00
61.50
57.75
o1
78.00
72.00
75.00
46.00
41.00
43.50
36.00
45.50
40.75
44.00
51.00
47.50
52.00
55.50
53.75
o3
48.00
47.50
47.75
63.00
59.00
61.00
22.00
27.50
24.75
32.00
37.00
34.50
42.00
40.50
41.25
Claude-Opus-4.6
12.00
17.50
14.75
34.00
36.00
35.00
10.00
21.00
15.50
20.00
27.00
23.50
44.00
49.00
46.50
Table A5: Detailed performance on the shifted in-context YesNo tasks. ACC, Acc.rat, and Mean denote answer accuracy, answer rationality, and their average, respectively.
MP
PA
IQM
Model
ACC
Acc.rat
Mean
ACC
Acc.rat
Mean
ACC
Acc.rat
Mean
Open-source LMMs
Qwen3-VL-2B
88.00
84.00
86.00
72.00
68.50
70.25
62.00
81.00
71.50
Qwen3-VL-8B
92.00
88.00
90.00
74.00
70.50
72.25
68.00
82.50
75.25
Qwen3-VL-30B
86.00
91.00
88.50
80.00
79.00
79.50
70.00
87.00
78.50
Qwen3.5-27B
84.00
92.50
88.25
76.00
75.50
75.75
68.00
90.50
79.25
Qwen3.5-122B-A10B
70.00
90.50
80.25
74.00
82.50
78.25
66.00
85.50
75.75
LLaVA-1.5-7B
56.00
55.00
55.50
54.00
56.00
55.00
50.00
66.00
58.00
InternVL3-2B
82.00
70.50
76.25
62.00
65.50
63.75
48.00
71.50
59.75
InternVL3-8B
64.00
79.50
71.75
60.00
73.50
66.75
48.00
76.00
62.00
Gemma-4-26B
86.00
77.50
81.75
82.00
76.00
79.00
72.00
76.00
74.00
Gemma-4-31B
72.00
79.50
75.75
80.00
74.50
77.25
60.00
75.50
67.75
Llama-4-Maverick
66.00
86.00
76.00
66.00
82.50
74.25
68.00
75.50
71.75
Ministral-3-8B
64.00
82.50
73.25
68.00
77.50
72.75
60.00
82.00
71.00
Ministral-3-14B
74.00
76.50
75.25
82.00
80.00
81.00
64.00
81.00
72.50
Closed-source LMMs
Gemini-3.1-Pro
74.00
67.50
70.75
62.00
63.50
62.75
54.00
65.00
59.50
GPT-4o
78.00
87.00
82.50
74.00
82.50
78.25
58.00
86.00
72.00
o1
70.00
88.50
79.25
76.00
86.50
81.25
56.00
86.00
71.00
o3
76.00
80.50
78.25
74.00
80.50
77.25
50.00
77.50
63.75
Claude-Opus-4.6
82.00
94.50
88.25
84.00
81.50
82.75
58.00
87.00
72.50
Table A6: Detailed performance on the shifted in-context MCQ tasks. ACC, Acc.rat, and Mean denote answer accuracy, answer rationality, and their average, respectively.
MP
PA
IQM
Model
ACC
Acc.rat
Mean
ACC
Acc.rat
Mean
ACC
Acc.rat
Mean
Open-source LMMs
Qwen3-VL-2B
90.00
66.00
78.00
66.00
69.00
67.50
84.00
70.50
77.25
Qwen3-VL-8B
86.00
83.50
84.75
74.00
76.50
75.25
92.00
88.50
90.25
Qwen3-VL-30B
92.00
79.00
85.50
84.00
80.00
82.00
94.00
96.00
95.00
Qwen3.5-27B
98.00
89.50
93.75
88.00
83.00
85.50
96.00
93.00
94.50
Qwen3.5-122B-A10B
88.00
86.50
87.25
92.00
89.00
90.50
96.00
93.00
94.50
LLaVA-1.5-7B
36.00
26.00
31.00
30.00
38.00
34.00
42.00
34.00
38.00
InternVL3-2B
78.00
50.00
64.00
70.00
66.00
68.00
84.00
57.00
70.50
InternVL3-8B
84.00
56.00
70.00
74.00
69.00
71.50
78.00
60.50
69.25
Gemma-4-26B
98.00
80.50
89.25
86.00
84.50
85.25
86.00
80.50
83.25
Gemma-4-31B
90.00
78.50
84.25
86.00
83.50
84.75
94.00
92.50
93.25
Llama-4-Maverick
94.00
75.00
84.50
90.00
83.00
86.50
92.00
79.50
85.75
Ministral-3-8B
88.00
72.50
80.25
68.00
75.50
71.75
94.00
86.00
90.00
Ministral-3-14B
90.00
76.00
83.00
76.00
81.00
78.50
80.00
78.50
79.25
Closed-source LMMs
Gemini-3.1-Pro
58.00
61.00
59.50
56.00
61.50
58.75
78.00
73.50
75.75
GPT-4o
76.00
65.00
70.50
80.00
72.50
76.25
90.00
74.00
82.00
o1
90.00
69.00
79.50
74.00
72.50
73.25
90.00
80.50
85.25
o3
88.00
62.00
75.00
78.00
72.50
75.25
88.00
70.00
79.00
Claude-Opus-4.6
62.00
54.00
58.00
48.00
42.50
45.25
42.00
38.00
40.00
Table A7: Detailed performance on the shifted in-context VQA tasks. ACC, Acc.rat, and Mean denote answer accuracy, answer rationality, and their average, respectively.
MP
PA
IQM
Model
ACC
Acc.rat
Mean
ACC
Acc.rat
Mean
ACC
Acc.rat
Mean
Open-source LMMs
Qwen3-VL-2B
68.00
58.00
63.00
22.00
50.50
36.25
88.00
77.50
82.75
Qwen3-VL-8B
86.00
80.50
83.25
62.00
54.00
58.00
90.00
85.50
87.75
Qwen3-VL-30B
78.00
75.00
76.50
44.00
66.00
55.00
82.00
83.50
82.75
Qwen3.5-27B
92.00
89.00
90.50
56.00
70.00
63.00
88.00
90.00
89.00
Qwen3.5-122B-A10B
90.00
90.50
90.25
68.00
76.50
72.25
86.00
87.50
86.75
LLaVA-1.5-7B
28.00
28.00
28.00
1.00
21.00
11.00
58.00
52.50
55.25
InternVL3-2B
56.00
48.00
52.00
30.00
53.00
41.50
78.00
68.50
73.25
InternVL3-8B
70.00
55.00
62.50
44.00
62.00
53.00
76.00
65.50
70.75
Gemma-4-26B
74.00
74.50
74.25
58.00
77.00
67.50
78.00
78.00
78.00
Gemma-4-31B
78.00
73.50
75.75
66.00
75.50
70.75
80.00
81.00
80.50
Llama-4-Maverick
80.00
75.00
77.50
50.00
66.00
58.00
80.00
80.00
80.00
Ministral-3-8B
80.00
73.50
76.75
58.00
70.50
64.25
80.00
81.00
80.50
Ministral-3-14B
60.00
70.50
65.25
46.00
65.00
55.50
72.00
75.00
73.50
Closed-source LMMs
Gemini-3.1-Pro
64.00
59.00
61.50
18.00
27.50
22.75
50.00
54.00
52.00
GPT-4o
80.00
73.50
76.75
52.00
73.00
62.50
82.00
78.50
80.25
o1
80.00
69.50
74.75
52.00
61.00
56.50
84.00
78.00
81.00
o3
80.00
62.00
71.00
30.00
58.50
44.25
92.00
78.50
85.25
Claude-Opus-4.6
22.00
21.50
21.75
26.00
38.50
32.25
34.00
37.50
35.75
Table A8: Refusal performance on image–question mismatch samples derived from MME, MMStar, and OK-VQA. The three subsets correspond to YesNo, multiple-choice, and open-ended VQA formats, respectively. Higher scores indicate stronger capability to identify and appropriately refuse image–question mismatches.
YesNo
MCQ
VQA
Model
MME-Mismatch
MMStar-Mismatch
OK-VQA-Mismatch
Open-source LMMs
Qwen3-VL-2B
1.00
62.00
60.00
Qwen3-VL-8B
2.00
66.00
79.00
Qwen3-VL-30B
4.00
62.00
83.00
Qwen3.5-27B
7.00
66.00
60.00
Qwen3.5-122B-A10B
32.00
73.00
71.00
LLaVA-1.5-7B
0.00
2.00
6.00
InternVL3-2B
58.00
24.00
77.00
InternVL3-8B
32.00
59.00
85.00
Gemma-4-26B
68.00
73.00
88.00
Gemma-4-31B
67.00
76.00
89.00
Llama-4-Maverick
73.00
57.00
73.00
Ministral-3-8B
38.00
69.00
85.00
Ministral-3-14B
39.00
64.00
68.00
Closed-source LMMs
Gemini-3.1-Pro
45.00
73.00
82.00
GPT-4o
94.00
77.00
86.00
o1
58.00
72.00
87.00
o3
54.00
70.00
79.00
Claude-Opus-4.6
32.00
66.00
70.00
Findings
All 18 models scored low and inconsistently on truly unanswerable (OOC) questions, with uncertain spatial/physical context and unclear logical/symbolic categories being hardest (e.g., Qwen3-VL-2B scored only 5.75 on yes/no and 8.25 on open-ended VQA for uncertain spatial/physical context).
Larger model size did not consistently improve OOC performance: Qwen3-vl-30B outperformed Qwen3-vl-2B, but Qwen3.5-122B-A10B did not consistently beat smaller Qwen3-vl variants.
Among closed-source models, o1 achieved the strongest overall OOC performance, ranking first on yes/no and VQA tasks and competitive on MCQ, while Gemini-3.1-Pro and Claude-Opus-4.6 showed clear weaknesses on several OOC categories.
Most models performed substantially better on answerable Shifted In-Context tasks than on OOC tasks, but Partial Answerability in open-ended VQA remained particularly weak, meaning contextual shifts still disrupted reliable answering.
Post-training (SFT and, to a lesser extent, DPO) improved refusal rate and refusal rationality on OOC settings across the tested models, but this often came with a drop in general multimodal performance on MMStar, indicating a trade-off.
Where it can be used
Pre-deployment evaluation of vision-language chatbots or visual Q&A services to check whether they invent answers to unsupported questions.
Checking whether image-based customer support or document/medical image query tools over-refuse rather than answering the parts of a question that are actually supported.
Using the benchmark to measure the trade-off between refusal alignment and general capability when applying post-training methods like SFT or DPO to a model.
Limits and open work
The benchmark currently covers only image-text interaction; extending it to video, audio, and embodied settings is left as future work by the authors.
Question-only experiments (without images) showed some models achieving high refusal scores from linguistic cues alone or conservative refusal strategies, so this baseline should be treated as diagnostic rather than proof of genuine multimodal reasoning.
On image-question mismatch samples drawn from MME, MMStar, and OK-VQA (Table A8), performance varied greatly by format (yes/no, MCQ, VQA), showing limited cross-format generalization.
Analyses of prompt type sensitivity, misleading-prompt robustness, and Gaussian noise robustness were demonstrated on specific models (e.g., Qwen3-VL-8B, o1) rather than across the full model set.
Why it matters
If a multimodal AI confidently fabricates answers to questions its image can't actually support, it spreads misinformation; if it refuses everything that looks tricky, it becomes useless. MMOOC gives developers a way to measure both failure modes together, which matters for anyone deploying vision-language chatbots or image-based Q&A tools where reliability depends on knowing when to answer and when to say no.
Terms in this paper
MLLM (Multimodal Large Language Model) · an AI model that takes both an image and text as input to produce an answer
Out-of-Context (OOC) · a question that the given image genuinely cannot support an answer to
Shifted In-Context · a question phrased in a misleading or confusing way that can nonetheless still be answered from the image
LLM-as-a-Judge · using other AI models as graders to score how sound and well-reasoned a response is
Refusal Rationality · a score measuring how well-grounded and coherent a model's explanation is when it refuses to answer
Original abstract (English)
Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts. A reliable MLLM should refuse truly out-of-context (OOC) questions with subject-level context shifts while still answering shifted in-context (Shifted IC) questions with non-subject context shifts. Existing benchmarks mainly target OOC or visually unanswerable questions, but overlook answerable Shifted IC cases and cover limited OOC shifts. To fill this gap, we present MMOOC, a large-scale benchmark for evaluating refusal and robust answering abilities of MLLMs. MMOOC contains over 41K image-question pairs, including answerable Shifted IC cases and unanswerable OOC cases, spanning three question formats, eight shift types and six visual scenarios, with data quality ensured through MLLM-based filtering and human verification. We evaluate model responses using Accuracy and Refusal Rate, and further introduce an LLM-as-a-Judge metric to assess the correctness of model reasoning. Experiments on diverse MLLMs show that current models still struggle to balance answer-ability and refusal under shifted contexts. We further analyze key failure patterns and show that post-training can improve robustness. MMOOC will be made publicly available.