Figure 1: Overview of HealMed. a. Benchmark scope, comprising MCQA, NLI and open-ended QA across nine languages and nine source datasets. b. Construction pipeline, including English source-example selection, machine translation into eight target languages, two-stage review and revision by bilingual medical experts, and final consistency checking and curation.
Table 1: Language-specific refusal rates in open-ended QA. Values are percentages of 400 zero-shot responses per language. A refusal was counted only when the model explicitly declined to answer and provided no substantive response. Overall rates were calculated across 3,600 responses per model. Gemini-3 denotes Gemini-3-Flash.
Language
GPT-4o
GPT-5.4
o4-mini
Gemini-3
English
0.00
0.00
0.00
0.00
German
0.00
0.00
0.00
0.00
Spanish
0.00
0.00
0.00
0.00
Portuguese
0.00
0.25
0.00
0.00
Japanese
0.00
0.25
0.00
0.00
Chinese
0.25
0.00
0.00
0.00
Thai
0.25
0.00
0.25
0.00
Swahili
1.50
0.00
0.00
0.00
Zulu
2.00
0.00
0.00
0.00
Overall
0.44
0.06
0.03
0.00
Figure 2: Multilingual performance on the MCQA and NLI tasks in HealMed. a. Macro-average accuracy across four MCQA and two NLI datasets for 14 models. Blue bars show the lower-resource mean, and red extensions show the gap to the higher-resource mean. Circles, squares and diamonds denote proprietary, open-source and medically specialized models, respectively. b. Language-level accuracy across models. c. Within-model accuracy shifts relative to English. Grey points represent individual models; colored points and horizontal lines show the mean and interquartile range. Black, blue and red denote English, other higher-resource languages and lower-resource languages, respectively. Higher-resource languages comprise English, German, Spanish, Portuguese, Japanese and Chinese; lower-resource languages comprise Thai, Swahili and Zulu.
Table 2: Expert assessment and revision of machine-translated data by language. Accuracy (Acc.), fluency (Flu.) and completeness (Comp.) are mean expert ratings on five-point scales. Revision is the mean normalized word-level edit distance between the original machine translations and reviewer-submitted revisions. Each language contains 1,000 instances.
Language
Acc.
Flu.
Comp.
Revision (%)
German
4.87
4.75
4.83
12.1
Spanish
4.92
4.92
4.85
1.0
Portuguese
4.74
4.84
4.93
2.7
Japanese
4.70
4.62
4.96
5.7
Chinese
4.73
4.51
4.98
5.2
Thai
4.71
4.79
4.93
9.1
Swahili
4.50
4.57
4.63
1.8
Zulu
4.60
4.68
4.92
1.1
Overall
4.72
4.71
4.88
4.8
Figure 3: Open-ended QA performance on HealMed. a. Mean LLM-as-judge scores for ten models, macro-averaged across the three QA datasets. Blue and red points show higher- and lower-resource means; diamonds show means across all languages. b. Score shifts relative to each model’s English score. Positive values indicate higher scores than the English baseline. Grey lines show individual models and the dark line shows their mean. c. Proportions of responses with an overall score below 3, a safety score of 2 or lower, or specific language and applicability issues. Other higher-resource languages comprise English, German, Spanish, Portuguese, Japanese and Chinese. Bubble size and color indicate the proportion; values are shown when at least 10%.
Table 3: Criterion-level comparison of expert and LLM-based evaluations. Scores range from 1 to 5. Δ denotes the LLM score minus the expert score and was calculated before rounding.
Criterion
Expert
LLM
Δ
Chinese
Completeness
4.22
3.80
−0.42
Reference alignment
4.24
3.62
−0.62
Clinical consensus
4.33
4.04
−0.29
Clinical appropriateness
4.31
3.93
−0.38
Safety
4.69
4.20
−0.49
Overall
4.36
3.92
−0.44
Japanese
Completeness
4.20
3.44
−0.76
Reference alignment
4.76
3.71
−1.04
Clinical consensus
4.82
4.20
−0.62
Clinical appropriateness
4.78
4.02
−0.76
Safety
4.87
4.27
−0.60
Overall
4.68
3.93
−0.76
Thai
Completeness
3.28
3.38
+0.10
Reference alignment
3.64
3.62
−0.02
Clinical consensus
4.12
4.11
−0.01
Clinical appropriateness
3.80
3.91
+0.11
Safety
4.36
4.22
−0.13
Overall
3.84
3.85
+0.01
Figure 4: Evaluation shifts between expert-reviewed HealMed and machine-translated (MT) data. a. English-adjusted mean accuracy shifts (HealMed minus MT) across 14 models and six MCQA and NLI datasets. Cells are labelled when |Δ|≥3 percentage points (pp). b. English-adjusted mean LLM-as-judge score shifts across five models and three open-ended QA datasets. Symbols denote datasets, and horizontal lines span their mean shifts. In both panels, Overall reports the mean absolute model-level shift across the corresponding models and datasets. Positive values indicate higher performance on expert-reviewed data; English is shown unadjusted as a same-source control.
Table 4: Language use in observable reasoning traces generated by HuatuoGPT-o1-72B. The “Traces” column denotes the number of responses containing an explicit reasoning trace among 15 responses examined per language. ”Target” and ”English” report the number and percentage of traces written predominantly in the target language or English, respectively.
Language
Traces
Target
English
Japanese
9/15
7 (77.8%)
2 (22.2%)
Chinese
15/15
15 (100%)
0 (0%)
Thai
7/15
0 (0%)
7 (100%)
Table 5: Benchmark components and sample allocation in HealMed. Sample counts denote the number of aligned examples included in each language.
Task
Component
Samples per language, n
Expected model output
MCQA
HeadQA
75
Correct option
MedQA
75
Correct option
MedExpQA
75
Correct option
MMLU-Pro
75
Correct option
NLI
BioNLI
150
Relation label
MedNLI
150
Relation label
QA
ExpertQA-Bio
40
Free-form answer
ExpertQA-Med
160
Free-form answer
LiveQA
200
Free-form answer
Total
1,000
Table 6: Comparison of selected multilingual medical benchmarks. Reported scales use source-specific units and are not directly comparable. Quality metrics indicates whether aggregate translation-quality scores or revision statistics were reported; score effects indicates whether model performance was compared between machine-translated and expert-reviewed versions. WorldMedQA-V identified four country-level validators and seven contributors to English-translation validation, but did not report the number of unique reviewers because these roles may overlap. EN, English; MT, machine translation; N/A, not applicable; NR, not reported. For BRIDGE, reference standards were inherited from the source datasets, benchmark-wide human-review coverage and reviewer numbers were not reported.
Benchmark scope
Data construction
Human review
Translation audit
Benchmark
Reported scale
Lang.
Tasks
Language design
Review coverage (%)
Reviewers, n
Quality metrics
Score effects
MMedBench
53,566 QA pairs
6
MCQA; rationales
Aggregated; non-aligned
14.1
3
No
No
XMedBench
21,326 records
6
MCQA
Native + MT; non-aligned
10.2
NR
No
No
MedExpQA
2,488 records
4
MCQA
MT + manual revision; aligned
100
NR
No
No
WorldMedQA-V
568 evaluation items
4 + EN
Multimodal MCQA
Local–English pairs
100
4 country 7 EN
No
No
MultiMed-X
2,450 translations
7 + EN
NLI; open QA
MT + expert revision; aligned
100
∼12
No
No
BRIDGE
1,418,042 samples
9
8 task types
Native-source aggregation; non-aligned
Varies by source; NR
NR
N/A
N/A
HealMed
9,000 instances
9
MCQA; NLI; open QA
MT + two-expert medical revision; aligned
100
23
Yes
Yes
Table 7: Models and evaluation configurations used in HealMed.
Model
Parameters
Release date
Model type
Proprietary models
GPT-5.4
Not disclosed
5 Mar 2026
Proprietary
o4-mini
Not disclosed
16 Apr 2025
Proprietary
Gemini-3-Flash
Not disclosed
17 Dec 2025
Proprietary
Claude-Sonnet-5
Not disclosed
30 Jun 2026
Proprietary
Claude-Opus-4.8
Not disclosed
28 May 2026
Proprietary
Open-source models
DeepSeek-V3
671B
26 Dec 2024
open-source, general-purpose
Qwen2.5-72B-Instruct
72.7B
19 Sep 2024
open-source, general-purpose
Qwen3-32B
32.8B
29 Apr 2025
open-source, general-purpose
Qwen3-32B-thinking
32.8B
29 Apr 2025
open-source, general-purpose
LLaMA3.3-70B-Instruct
70B
6 Dec 2024
open-source, general-purpose
Gemma-3-27B-it
27B
12 Mar 2025
open-source, general-purpose
Medically specialized models
HuatuoGPT-o1-72B
72.7B
28 Dec 2024
open-source, medical
MedGemma-27B-text-it
27B
20 May 2025
open-source, medical
MediPhi
3.8B
3 Feb 2025
open-source, medical
Table 8: Five-point rubric used by medical experts to assess machine translations. Each dimension was scored separately.
Score
Accuracy
Fluency
Completeness
5
All concepts and medical terms are translated correctly and precisely. Terminology is professional, contextually appropriate and consistent with established usage in the target language.
The translation is natural and easy to read. Its grammar, wording and sentence structure conform to professional conventions in the target language.
The meaning and all relevant details of the source are retained. There are no omissions or unsupported additions.
4
Most concepts and terms are translated correctly. Minor errors, imprecise wording or simplified terminology may occur but do not affect overall understanding.
The translation is generally natural and clear. Minor stiffness, awkward wording or grammatical errors do not affect comprehension.
The main meaning is retained. Only minor or non-essential details are omitted or expressed unclearly.
3
The main concepts are conveyed, but some errors or imprecise terms may cause partial misunderstanding. The reader may need to infer the intended meaning of some terms.
The translation is understandable but noticeably unnatural in places. Rigid sentence structures, unsuitable word choices or grammatical errors require some effort from the reader.
Most of the source meaning is conveyed, but some information is missing, added or unclear. Important details may require inference.
2
Several important concepts or terms are mistranslated, substantially affecting comprehension. Terminology may be incorrect or inconsistent.
The translation is difficult to read smoothly. It contains awkward transitions, unclear connections or frequent grammatical and structural errors.
Core information is not fully preserved. Noticeable omissions or unnecessary additions reduce correspondence with the source and affect comprehension.
1
Frequent and severe mistranslations prevent the source meaning from being conveyed. Much of the content or terminology does not correspond to the source.
The translation is highly unnatural or difficult to understand. Literal phrasing, disorganized sentence structure and severe grammatical errors may make it unreadable.
Substantial omissions or incorrect additions prevent the translation from reflecting the source. Important passages are missing, and the intended meaning is difficult to recover.
Table 9: Inference settings used for model evaluation.
Setting
MCQA
BioNLI
MedNLI
Open-ended QA
Prompting strategy
zero-shot
zero-shot
zero-shot
zero-shot
Temperature
0.7 where supported; otherwise omitted
Maximum output tokens
32
32
32
2,048
Top-p
Not set; model- or provider-default
Top-k
Not set; model- or provider-default where supported
Random seed
Not set
Stop sequence
Not set
Table 10: Complete zero-shot accuracy (%) on the four MCQA datasets. Avg. denotes the macro-average across the nine languages. Within each dataset and language, the highest value is shown in bold and the second-highest is underlined.
Dataset
Model
EN
DE
ES
PT
JA
ZH
TH
SW
ZU
Avg.
HeadQA
Proprietary Models
Gemini-3-Flash
94.67
94.67
96.00
96.00
92.00
92.00
90.67
92.00
88.00
92.89
GPT-5.4
90.67
93.33
92.00
93.33
89.33
89.33
86.67
90.67
84.00
89.93
o4-mini
96.00
94.67
97.33
96.00
92.00
90.67
94.67
85.33
85.33
92.44
Claude-Opus-4.8
96.00
96.00
96.00
94.67
90.67
90.67
89.33
93.33
82.67
92.15
Claude-Sonnet-5
90.67
96.00
94.67
94.67
92.00
90.67
93.33
90.67
73.33
90.67
open-source Models
DeepSeek-V3
82.67
84.00
80.00
84.00
82.67
82.67
72.00
60.00
54.67
75.85
Gemma-3-27B-it
80.00
78.67
82.67
78.67
78.67
73.33
74.67
66.67
54.67
74.22
LLaMA3.3-70B-Instruct
85.33
80.00
84.00
84.00
77.33
69.33
73.33
65.33
46.67
73.92
Qwen2.5-72B-Instruct
80.00
80.00
81.33
82.67
82.67
78.67
76.00
37.33
9.33
67.56
Qwen3-32B
76.00
74.67
73.33
74.67
66.67
72.00
58.67
45.33
8.00
61.04
Qwen3-32B-thinking
92.00
90.67
92.00
93.33
90.67
86.67
88.00
65.33
24.00
80.30
Specialized Models
HuatuoGPT-o1-72B
89.33
84.00
90.67
81.33
86.67
86.67
86.67
56.00
41.33
78.07
MedGemma-27B
90.67
78.67
81.33
78.67
80.00
74.67
72.00
60.00
40.00
72.89
MediPhi
70.67
49.33
60.00
60.00
45.33
41.33
37.33
26.67
16.00
45.18
MedQA
Proprietary Models
Gemini-3-Flash
94.67
92.00
94.67
93.33
90.67
93.33
93.33
93.33
92.00
93.04
GPT-5.4
92.00
96.00
96.00
97.33
93.33
90.67
97.33
97.33
88.00
94.22
o4-mini
98.67
97.33
96.00
97.33
93.33
96.00
98.67
94.67
92.00
96.00
Claude-Opus-4.8
98.67
93.33
98.67
97.33
92.00
92.00
88.00
89.33
78.67
92.00
Claude-Sonnet-5
88.00
89.33
88.00
94.67
86.67
90.67
85.33
90.67
74.67
87.56
open-source Models
DeepSeek-V3
76.00
77.33
72.00
74.67
69.33
72.00
72.00
58.67
42.67
68.30
Gemma-3-27B-it
65.33
61.33
60.00
64.00
58.67
68.00
57.33
57.33
49.33
60.15
LLaMA3.3-70B-Instruct
85.33
82.67
80.00
86.67
72.00
80.00
81.33
69.33
34.67
74.67
Qwen2.5-72B-Instruct
76.00
78.67
74.67
74.67
72.00
73.33
69.33
42.67
22.67
64.89
Qwen3-32B
62.67
57.33
65.33
54.67
53.33
62.67
69.33
46.67
4.00
52.89
Qwen3-32B-thinking
89.33
90.67
89.33
94.67
84.00
85.33
93.33
70.67
36.00
81.48
Table 11: Complete zero-shot accuracy (%) on the two NLI datasets. Avg. denotes the macro-average across the nine languages. Within each dataset and language, the highest value is shown in bold and the second-highest is underlined.
Dataset
Model
EN
DE
ES
PT
JA
ZH
TH
SW
ZU
Avg.
BioNLI
Proprietary Models
Gemini-3-Flash
79.33
71.33
74.00
72.67
71.33
71.33
74.67
73.33
70.00
73.11
GPT-5.4
73.33
68.00
70.00
68.67
68.00
68.67
68.00
67.33
66.67
68.74
o4-mini
74.00
71.33
71.33
71.33
71.33
70.67
70.00
66.67
63.33
70.00
Claude-Opus-4.8
78.00
70.67
74.67
72.00
72.00
76.00
74.00
74.67
70.67
73.63
Claude-Sonnet-5
72.67
68.00
71.33
68.00
68.00
69.33
68.67
67.33
62.00
68.37
open-source Models
DeepSeek-V3
77.33
65.33
64.67
62.67
62.00
70.67
62.67
60.00
49.33
63.85
Gemma-3-27B-it
66.67
64.67
64.00
62.00
63.33
65.33
60.00
58.00
56.00
62.22
LLaMA3.3-70B-Instruct
71.33
56.00
64.00
64.67
63.33
62.00
61.33
59.33
49.33
61.26
Qwen2.5-72B-Instruct
75.33
67.33
72.67
66.67
60.67
65.33
57.33
58.67
58.67
64.74
Qwen3-32B
70.67
61.33
70.67
66.00
58.00
60.67
58.00
48.67
55.33
61.04
Qwen3-32B-thinking
74.00
66.67
70.00
68.00
68.00
66.00
67.33
62.67
35.33
64.22
Specialized Models
HuatuoGPT-o1-72B
66.67
64.00
66.67
64.00
62.00
60.67
63.33
58.67
61.33
63.04
MedGemma-27B
59.33
58.00
64.67
67.33
59.33
50.67
53.33
51.33
54.00
57.55
MediPhi
66.67
63.33
63.33
62.67
56.67
61.33
57.33
47.33
52.67
59.04
MedNLI
Proprietary Models
Gemini-3-Flash
91.33
88.67
88.67
87.33
86.67
82.67
86.67
82.00
76.67
85.63
GPT-5.4
85.33
84.67
82.00
81.33
82.67
80.67
84.00
82.00
79.33
82.44
o4-mini
90.67
90.00
87.33
86.67
84.00
81.33
84.00
84.67
79.33
85.33
Claude-Opus-4.8
89.33
88.00
88.00
86.67
86.00
83.33
89.33
86.67
84.67
86.89
Claude-Sonnet-5
91.33
93.33
90.00
92.67
87.33
87.33
88.00
85.33
81.33
88.52
open-source Models
DeepSeek-V3
84.67
68.67
75.33
72.00
78.67
75.33
76.67
62.67
44.00
70.89
Gemma-3-27B-it
88.67
86.67
86.00
83.33
84.00
76.00
78.00
76.67
57.33
79.63
LLaMA3.3-70B-Instruct
81.33
74.67
76.00
77.33
80.00
70.00
70.67
68.00
30.67
69.85
Qwen2.5-72B-Instruct
86.67
82.00
84.00
84.00
82.67
78.00
80.00
57.33
32.67
74.15
Qwen3-32B
75.33
75.33
82.67
83.33
81.33
74.00
67.33
34.67
30.67
67.18
Qwen3-32B-thinking
83.33
81.33
82.67
80.00
80.00
79.33
80.67
65.33
36.00
74.30
Table 12: Complete LLM-as-judge scores on the three open-ended QA datasets. Scores range from 1 to 5 and are calculated as the mean of completeness, reference alignment, clinical consensus, clinical appropriateness and safety. Avg. denotes the macro-average across the nine languages. Within each dataset and language, the highest value is shown in bold and the second-highest is underlined.
We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn from nine datasets and covering three task formats: MCQA, NLI and open-ended QA. The benchmark was developed over two years by 23 physicians and medical experts based across nine countries and regions. Each translation was evaluated and revised by two experts fluent in English and the corresponding target language. On HealMed, performance declined most in low-resource languages, although the size of the gap varied markedly across languages and models. The strongest proprietary models were the most stable across languages, whereas many open-source and medically specialized models showed larger and less consistent gaps. Medical specialization alone did not ensure multilingual robustness. Furthermore, expert revision could either raise or lower measured performance, indicating that translation quality materially affects cross-language evaluation results.
作者 · Yingjian Chen (Drew), Fan Gao (Drew), Sherry T. Tong (Drew), Haoyu Zhang (Drew), Aosong Feng (Drew), Kevin W. Jin (Drew)