One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

Auditing Cross-Lingual Fairness in Language Model Watermarking

arXiv:2608.200472026-08-21

AI text watermarks that are supposed to catch machine-written content work far less reliably in many non-English languages, and the gap tracks language families, not individual languages

Companies increasingly embed invisible statistical signatures ('watermarks') in AI-generated text so it can later be flagged as machine-written, but almost all testing of these systems has been done in English. This study builds a testing framework and runs it across 11 languages, 6 watermarking methods, and 3 open AI models, finding that detection and text-quality preservation both break down unevenly depending on language family. The unevenness is structural: it comes from shared properties of related languages (like tokenization or morphology), not quirks of one specific language.

What they did

  1. Six watermarking schemes were tested on 11 languages across 4 writing scripts and 8 language families, using 3 open-weight AI models in both 'base' (plain text continuation) and 'instruction-tuned' (chatbot-style) modes.
  2. The team recalibrated detection cutoffs for each language and model rather than using each scheme's default English-tuned setting, and separately measured whether watermark signals were fully absent versus just miscalibrated.
  3. Quality preservation was checked with three independent methods (overall text distribution similarity, per-example semantic similarity, and fluency under a reference model), because a single measure can hide problems the others catch.
  4. Two watermarking schemes marketed as 'distortion-free' (SynthID-Text and EXPEdit) turned out to be the most damaging to text quality in practice, and instruction-tuned models showed much weaker detection than base models across most schemes.
  5. Statistical analysis showed that most of the unfairness across languages came from differences between language families (e.g. Germanic vs. Semitic vs. Sinitic) rather than from any single unlucky language, suggesting the problem is built into how these systems handle language structure.
Figure 1: Per-language quality preservation under three paradigms, base-FLORES (top) and instruct-AYA (bottom); n=500 matched pairs per cell. Left: MAUVE on mean-pooled XLM-R-large embeddings (1.0 identical, 0.0 fully separable). Center: BERTScore F1, rescaled against per-language FLORES non-pair baselines (negative: paired generations less similar than random non-pairs). Right: PPL preservation under XGLM-7.5B, exp⁡(−|log⁡PPLR​(𝒳+)−log⁡PPLR​(𝒳−)|)∈(0,1]. Rows: (scheme, generator) cells; columns: eleven languages grouped by typological family.
Figure 1: Per-language quality preservation under three paradigms, base-FLORES (top) and instruct-AYA (bottom); n=500 matched pairs per cell. Left: MAUVE on mean-pooled XLM-R-large embeddings (1.0 identical, 0.0 fully separable). Center: BERTScore F1, rescaled against per-language FLORES non-pair baselines (negative: paired generations less similar than random non-pairs). Right: PPL preservation under XGLM-7.5B, exp⁡(−|log⁡PPLR​(𝒳+)−log⁡PPLR​(𝒳−)|)∈(0,1]. Rows: (scheme, generator) cells; columns: eleven languages grouped by typological family.
Table 1: The eleven evaluation languages with script, typological family, and Joshi resource tier.
ISO 639-3LanguageScriptTypological familyJoshi et al. tier
engEnglishLatinGermanic5
nldDutchLatinGermanic4
fraFrenchLatinRomance5
spaSpanishLatinRomance5
porPortugueseLatinRomance4
turTurkishLatinTurkic4
vieVietnameseLatinAustroasiatic4
hinHindiDevanagariIndic4
arbArabicArabicSemitic5
zhoChineseHanSinitic5
jpnJapaneseHan + KanaJaponic5
Figure 2: Per-scheme disparity decomposition under three quality paradigms, base-FLORES (top) and instruct-AYA (bottom). Bar height is GE2 over the per-language preservation vector, averaged across the three generators per scheme. Stacking shows the typological-partition decomposition into between-family (dark) and within-family (light) components; the percentage above is the between-family share. Absolute disparity is orders of magnitude smaller than under MAUVE or PPL but the between-family pattern persists. Y-axes differ across panels.
Figure 2: Per-scheme disparity decomposition under three quality paradigms, base-FLORES (top) and instruct-AYA (bottom). Bar height is GE2 over the per-language preservation vector, averaged across the three generators per scheme. Stacking shows the typological-partition decomposition into between-family (dark) and within-family (light) components; the percentage above is the between-family share. Absolute disparity is orders of magnitude smaller than under MAUVE or PPL but the between-family pattern persists. Y-axes differ across panels.
Table 2: Detection fairness by scheme, generator, and regime, subset all, α=0.01. Mean TPR shown with 95% bootstrap CIs (B=1000 resamples of the 11-language vector). Min TPR is the realized minimum. DI<.8 counts languages with TPRl<0.8⋅maxl⁡TPRl. GE2 uses the typological partition (Germanic, Romance, Indic, Semitic, Sinitic, Japonic, Turkic, Austroasiatic); Btw. % is its between-family share. Δ=mean​TPR​(τg)−mean​TPR​(τl). Values below 10−3 reported as <0.001; entries marked are undefined when the per-language mean is zero.
AUCTPR at τg (global)TPR at τl (per-lang)
SchemeGen.MeanMinMean TPRMin TPRDI<.8GE2Btw. %Mean TPRΔ
Base regime (FLORES+ continuation prompts)
KGWMistral0.9990.9970.992 [0.987, 0.996]0.9780<0.00183.60.991+0.001
Gemma0.9970.9910.969 [0.959, 0.979]0.9340<0.00172.40.974−0.005
Qwen0.9960.9870.965 [0.939, 0.986]0.8800<0.00185.10.966−0.001
UnigramMistral0.9960.9810.868 [0.802, 0.927]0.63830.00861.30.971−0.103
Gemma0.9860.9100.000 [0.000, 0.000]0.000110.799−0.799
Qwen0.9940.9720.898 [0.845, 0.942]0.72220.00493.80.949−0.051
SynthIDMistral1.0001.0001.000 [0.999, 1.000]0.9980<0.001100.01.000−0.000
Gemma1.0001.0001.000 [0.999, 1.000]0.9980<0.001100.01.000+0.000
Qwen1.0000.9990.997 [0.995, 0.999]0.9900<0.00197.50.998−0.000
EXPEditMistral1.0000.9980.999 [0.999, 1.000]0.9960<0.00188.01.000−0.000
Gemma1.0000.9981.000 [0.999, 1.000]0.9980<0.001100.01.000+0.000
Qwen0.9980.9930.991 [0.982, 0.997]0.9460<0.00199.20.992−0.001
DIPMistral0.9970.9890.975 [0.966, 0.984]0.9380<0.00192.60.975+0.000
Gemma0.9920.9650.963 [0.947, 0.977]0.9140<0.00199.50.965−0.002
Qwen0.9920.9770.931 [0.891, 0.968]0.79800.00394.70.929+0.002
XSIRMistral0.9810.9620.784 [0.743, 0.821]0.67220.00491.50.833−0.049
Gemma0.9840.9770.766 [0.712, 0.817]0.58840.00696.00.829−0.063
Qwen0.9780.9300.215 [0.118, 0.362]0.050100.49398.90.794−0.579
Instruct regime (AYA native-speaker prompts)
KGWMistral0.9530.9090.626 [0.540, 0.707]0.38670.02695.30.632−0.005
Gemma0.9000.8480.389 [0.325, 0.463]0.222100.05382.40.411−0.022
Qwen0.9660.9370.689 [0.628, 0.753]0.51080.01178.60.723−0.034
UnigramMistral0.9090.8650.311 [0.171, 0.453]0.02260.29799.80.450−0.139
Gemma0.8510.7900.037 [0.023, 0.052]0.010100.22681.20.151−0.113
Qwen0.9430.8800.546 [0.455, 0.621]0.21040.03291.70.636−0.090
SynthIDMistral0.9910.9820.905 [0.865, 0.940]0.78210.00377.70.903+0.002
Gemma0.9460.9020.602 [0.530, 0.675]0.422100.02079.70.597+0.005
Qwen0.9930.9770.941 [0.906, 0.971]0.79210.00291.90.942−0.002
Figure 3: Joint per-language detection-quality landscape, MAUVE paradigm. Markers: per-language (TPR​@​τg,MAUVE) pairs at α=0.01. Circles: base-FLORES. Triangles: instruct-AYA. Color: typological family. Annotations sB, sI: mean pairwise Euclidean distance within each regime. BERTScore and PPL-preservation companion panels in Appendix E.
Figure 3: Joint per-language detection-quality landscape, MAUVE paradigm. Markers: per-language (TPR​@​τg,MAUVE) pairs at α=0.01. Circles: base-FLORES. Triangles: instruct-AYA. Color: typological family. Annotations sB, sI: mean pairwise Euclidean distance within each regime. BERTScore and PPL-preservation companion panels in Appendix E.
Table 3: Cross-paradigm rank agreement: per-cell Spearman ρ between per-language preservation vectors, mean across (model, regime) cells with per-cell minimum in parentheses (subset=all; n=6 cells per scheme, n=36 pooled). Pairs use BERTScore F1 (rescaled) and PPL preservation under XGLM-7.5B; BS-raw gives nearly identical rankings (ρBS-raw↔BS-rescaled=+0.87). Negative minima identify cells where two paradigms produce opposite per-language orderings.
SchemeMAUVE ↔ BSMAUVE ↔ PPLBS ↔ PPL
KGW+0.61 (+0.45)+0.03 (−0.80)+0.03 (−0.76)
Unigram+0.56 (+0.23)+0.14 (−0.39)+0.20 (−0.67)
SynthID+0.76 (+0.47)+0.81 (+0.27)+0.77 (+0.45)
EXPEdit+0.68 (+0.26)+0.75 (+0.39)+0.70 (+0.46)
DIP+0.46 (−0.03)+0.43 (+0.05)+0.50 (−0.05)
XSIR+0.42 (−0.09)+0.10 (−0.71)+0.00 (−0.76)
All schemes+0.58 (−0.09)+0.38 (−0.80)+0.37 (−0.76)
Figure 4: Per-language detection diagnostics, subset all, α=0.01. Within each panel, left: TPR at the global empirical-FPR threshold τg; right: threshold-free AUC. Rows are (scheme, generator) cells; columns are the eleven evaluation languages, grouped by typological family.
Figure 4: Per-language detection diagnostics, subset all, α=0.01. Within each panel, left: TPR at the global empirical-FPR threshold τg; right: threshold-free AUC. Rows are (scheme, generator) cells; columns are the eleven evaluation languages, grouped by typological family.
Table 4: Per-cell language adherence on the eleven evaluation languages, base-FLORES (left) and instruct-AYA (right). NWM% is the GlotLID-v3 on-target rate for nowatermark generations; Δ¯ is the mean shift in percentage points across the six watermarked schemes. mean rows are unweighted means across the eleven languages.
Base regime (FLORES+ continuation prompts)Instruct regime (AYA native-speaker prompts)
MistralGemmaQwenMistralGemmaQwen
LangNWM%Δ¯NWM%Δ¯NWM%Δ¯NWM%Δ¯NWM%Δ¯NWM%Δ¯
eng99.6−0.398.8−2.098.4−0.692.6−0.980.6−1.793.2−2.0
fra99.2−1.197.6−4.089.2−0.291.4+0.586.2−1.794.8−4.9
spa100.0−1.496.8−2.691.2+1.195.2−1.362.6+2.792.0−5.0
por99.2−0.596.8−1.098.2−0.896.2−2.070.6+3.096.4−4.3
nld96.2−7.095.8+1.680.8−4.673.2−3.189.2−0.490.4−5.8
zho97.6−1.892.8−1.296.4−4.289.0−7.995.6−0.498.4−0.6
jpn96.8−3.395.6−1.699.8−3.371.0−10.794.0−1.998.4−7.4
arb89.2−3.087.6−3.887.8+2.967.6−15.684.2−1.281.0−15.0
hin96.2−9.188.2−2.398.8−3.952.2−6.583.8−1.483.2+0.1
tur89.0+1.396.8−4.357.2−1.181.8−8.192.8−0.695.6−13.5
vie98.6−4.777.0+4.799.0−1.079.2−7.398.8−0.695.2−6.4
mean96.5−2.893.1−1.590.6−1.480.9−5.785.3−0.492.6−5.9
(b) Instruct-AYA regime.
(b) Instruct-AYA regime.
Table 5: Length-distribution similarity between LID subsets and the all-subset baseline. Each row is a (regime, subset) pair; ratio=subset-mean length/all-subset mean length, computed per (regime, model, language, watermark) cell and summarized across cells. Cells require n≥100 in both subsets to enter the summary; n skipped counts cells that fail this floor (typically off_target on high-adherence cells).
RegimeSubsetn keptn skippedMeanMedianMinMax
Baseon_target_only23101.0031.0000.9961.081
Baseoff_target_only82230.9290.9080.8591.023
Instructon_target_only23101.0001.0000.9641.057
Instructoff_target_only262051.0251.0170.9471.119
Figure 5: PPL preservation under mGPT-1.3B cross-reference. Format and color scale match Figure 1.
Figure 5: PPL preservation under mGPT-1.3B cross-reference. Format and color scale match Figure 1.
Table 6: LID categorization robustness: per-cell adherence under a 3×3 sweep of the GlotLID per-sentence confidence floor (conf_floor) and the per-text on-target sentence fraction (on_target_ratio). ρ is the Spearman rank-correlation between the per-cell on-target-rate vector under each configuration and the default (†). Mean on-target % is the unweighted mean across the 462 cells of the evaluation grid. Mean/Max |Δ| are the mean and maximum absolute per-cell shift in on-target rate (pp) relative to default.
conf_flooron_target_ratioρ vs defaultMean on-target %Mean |Δ| (pp)Max |Δ| (pp)
0.30.70.934792.02.1913.60
0.30.80.946890.70.9113.00
0.30.90.862786.34.8731.00
0.50.70.991091.11.2611.80
0.5†0.81.000089.80.000.00
0.50.90.931785.44.4331.20
0.70.70.948687.93.7624.60
0.70.80.962786.83.0424.80
0.70.90.938582.57.3631.80
(b) Instruct-AYA regime.
(b) Instruct-AYA regime.
Table 7: Per-language calibration gap Δl=TPRl​(τl)−TPRl​(τg), base-FLORES (top) and instruct-AYA (bottom), subset all, α=0.01. Positive: per-language calibration recovers detection. Negative: per-language calibration worsens detection. Zero: detection insensitive to calibration locality. Row means equal −1× the Δ column of Table 2.
GermanicRomanceIndicSemiticSiniticJaponicTurkicAus.
SchemeGen.engnldfraspaporhinarbzhojpnturvie
Base regime (FLORES+ continuation prompts)
KGWMistral0.000+0.004+0.004+0.0020.000−0.0080.000−0.0060.000−0.004+0.002
Gemma+0.006+0.030+0.008+0.016+0.038−0.020+0.044−0.010−0.0640.000+0.010
Qwen+0.028+0.018−0.002−0.002+0.002−0.0080.000+0.0080.000−0.0300.000
UnigramMistral+0.074+0.354+0.218+0.292+0.142−0.062+0.004+0.064−0.018+0.074−0.012
Gemma+0.996+0.992+0.968+0.990+0.9600.000+0.958+0.976+0.984+0.9600.000
Qwen+0.024+0.034−0.040+0.002+0.022+0.106+0.230+0.134−0.008+0.092−0.034
SynthIDMistral0.0000.0000.0000.0000.0000.0000.000+0.0020.0000.0000.000
Gemma0.0000.0000.0000.0000.0000.0000.0000.0000.0000.0000.000
Qwen0.0000.0000.0000.0000.000+0.0020.000+0.0060.000−0.0040.000
EXPEditMistral0.000+0.0020.0000.0000.0000.0000.0000.0000.0000.0000.000
Gemma0.0000.0000.0000.0000.0000.0000.0000.0000.0000.0000.000
Qwen+0.002−0.0060.0000.0000.000+0.0120.0000.0000.000+0.0020.000
DIPMistral0.0000.000+0.002+0.002+0.002+0.0020.000−0.0140.000+0.0040.000
Gemma0.000−0.004+0.0020.000+0.002+0.010+0.004+0.004+0.018−0.0100.000
Qwen0.000−0.008+0.004+0.008−0.004−0.004+0.002+0.0100.000−0.028+0.002
XSIRMistral+0.024−0.054+0.166+0.106+0.086−0.022−0.066+0.092+0.014+0.014+0.178
Gemma+0.028+0.054+0.054+0.114+0.026+0.146+0.056−0.042+0.168+0.122−0.030
Qwen+0.516+0.774+0.770+0.802+0.912−0.286+0.684+0.638+0.464+0.360+0.732
Instruct regime (AYA native-speaker prompts)
KGWMistral+0.112+0.032−0.250+0.028+0.118−0.016+0.014+0.030−0.018+0.012−0.004
Gemma+0.030−0.070+0.108+0.046+0.112−0.0380.000−0.036−0.082+0.134+0.034
Qwen+0.012+0.014−0.036−0.058+0.110−0.106+0.138+0.016+0.084+0.058+0.140
UnigramMistral+0.340+0.380+0.090+0.282+0.316−0.1060.000+0.070−0.178+0.282+0.052
Gemma+0.024−0.054+0.174+0.178+0.056−0.070−0.036+0.376+0.276+0.272+0.050
Qwen+0.084−0.018−0.108−0.012+0.078+0.128+0.438+0.352+0.094+0.088−0.132
SynthIDMistral+0.056+0.020−0.026−0.034−0.048+0.0160.0000.000−0.0220.000+0.016
Gemma+0.050−0.030+0.046−0.022−0.020−0.022−0.124+0.038−0.102+0.094+0.036
Qwen+0.018−0.002+0.010−0.022−0.002+0.0160.0000.0000.0000.0000.000
Figure 6: On-target minus all-subset preservation, base-FLORES regime (top) and instruct-AYA regime (bottom). Each cell is Δ=preservationon-target−preservationall for the indicated paradigm; rows are (scheme, generator), columns are languages. Positive (red): on-target restriction raises preservation; negative (blue): on-target restriction lowers it. Color scale fixed at [−0.3,+0.3]; values outside this range saturate. The EXPEdit-Mistral, Arabic, BERTScore F1 entry in the instruct panel is omitted (—) because the on-target sample size falls below the n≥100 paired-set floor.
Figure 6: On-target minus all-subset preservation, base-FLORES regime (top) and instruct-AYA regime (bottom). Each cell is Δ=preservationon-target−preservationall for the indicated paradigm; rows are (scheme, generator), columns are languages. Positive (red): on-target restriction raises preservation; negative (blue): on-target restriction lowers it. Color scale fixed at [−0.3,+0.3]; values outside this range saturate. The EXPEdit-Mistral, Arabic, BERTScore F1 entry in the instruct panel is omitted (—) because the on-target sample size falls below the n≥100 paired-set floor.
Table 8: Detection fairness summary at α=0.05, base-FLORES (top) and instruct-AYA (bottom), subset all. Format mirrors Table 2; AUC columns are omitted (threshold-independent, identical to Table 2) and bootstrap CIs on Mean TPR are omitted for compactness. L=11 throughout. Values below 10−3 reported as <0.001.
TPR at τg (global)TPR at τl (per-lang)
SchemeGen.Mean TPRMin TPRDI<.8GE2Btw. %Mean TPRΔ
Base regime (FLORES+ continuation prompts)
KGWMistral0.9970.9920<0.00161.10.997+0.000
Gemma0.9910.9800<0.00193.20.993−0.002
Qwen0.9870.9520<0.00157.80.986+0.000
UnigramMistral0.9770.9200<0.00154.90.988−0.011
Gemma0.9880.9760<0.00155.90.898+0.089
Qwen0.9610.86000.00186.30.979−0.018
SynthIDMistral1.0000.9980<0.001100.01.000+0.000
Gemma1.0000.9980<0.001100.01.000+0.000
Qwen0.9990.9960<0.00174.30.999+0.000
EXPEditMistral1.0000.9960<0.001100.01.000+0.000
Gemma1.0000.9980<0.001100.01.000+0.000
Qwen0.9950.9660<0.00199.20.995−0.000
DIPMistral0.9900.9720<0.00194.70.989+0.001
Gemma0.9810.9420<0.00197.90.981−0.001
Qwen0.9660.89600.00191.60.966+0.000
XSIRMistral0.9240.86800.00183.30.935−0.011
Gemma0.9290.84000.00192.80.944−0.015
Qwen0.5030.27690.07694.60.919−0.416
Instruct regime (AYA native-speaker prompts)
KGWMistral0.8010.62220.00685.00.813−0.012
Gemma0.6160.474100.01579.70.617−0.001
Qwen0.8600.74010.00378.10.869−0.009
UnigramMistral0.5170.09050.14995.80.659−0.143
Gemma0.3800.15450.05080.30.408−0.028
Qwen0.7270.36640.01893.80.790−0.063
SynthIDMistral0.9670.9260<0.00177.80.969−0.002
Gemma0.7910.66240.00575.90.785+0.005
Qwen0.9710.8980<0.00195.70.973−0.002
Figure 7: Joint per-language detection–quality landscape, BERTScore-rescaled (top) and PPL-XGLM (bottom) paradigms. Format matches Figure 3: markers are per-language (TPR​@​τg,quality) pairs at α=0.01, n=33 per regime (11 languages × 3 generators); circles are base-FLORES, triangles are instruct-AYA; color encodes typological family; sB, sI are mean pairwise Euclidean distance within each regime.
Figure 7: Joint per-language detection–quality landscape, BERTScore-rescaled (top) and PPL-XGLM (bottom) paradigms. Format matches Figure 3: markers are per-language (TPR​@​τg,quality) pairs at α=0.01, n=33 per regime (11 languages × 3 generators); circles are base-FLORES, triangles are instruct-AYA; color encodes typological family; sB, sI are mean pairwise Euclidean distance within each regime.
Table 9: Detection summary on the on-target subset, α=0.01, global threshold τg, base-FLORES (left) and instruct-AYA (right). Thresholds re-calibrated on each cell’s on-target unwatermarked pool; L=11 throughout. The off-target subset is omitted: in the base regime no cell meets the nneg≥200 inclusion floor, and in the instruct regime only six (Mistral-generator) cells pass the relaxed α=0.05 floor with L≤3, insufficient for the fairness aggregates this appendix reports. The Δ vs. all columns report change in Mean TPR relative to subset all (Table 2).
Base regime (FLORES+ continuation)Instruct regime (AYA native-speaker)
SchemeGen.Mean TPRMin TPRDI<.8Δ vs. allMean TPRMin TPRDI<.8Δ vs. all
KGWMistral0.9950.9830+0.0030.6810.4256+0.055
Gemma0.9730.9390+0.0040.4260.26410+0.037
Qwen0.9870.9630+0.0220.6970.5208+0.008
UnigramMistral0.8850.6542+0.0170.3880.0117+0.077
Gemma0.9870.9740+0.9870.0200.0009−0.017
Qwen0.9250.7621+0.0270.5470.1764+0.001
SynthIDMistral1.0000.99800.0000.9240.7741+0.019
Gemma1.0000.99800.0000.6100.42610+0.008
Qwen0.9990.9960+0.0020.9410.76410.000
EXPEditMistral1.0001.0000+0.0010.6750.2345+0.098
Gemma1.0000.99700.0000.2320.09110+0.007
Qwen0.9940.9410+0.0030.5080.05910+0.012
DIPMistral0.9800.9520+0.0050.4880.2669+0.033
Gemma0.9770.9540+0.0140.1980.10510+0.002
Qwen0.9570.8480+0.0260.5070.2008−0.006
XSIRMistral0.7990.6882+0.0150.2960.08110+0.030
Gemma0.7930.6003+0.0270.0920.01910+0.011
Qwen0.2070.04710−0.0080.0470.00010+0.004
Table 10: Bottom-quintile-mean Rawlsian floor as a stability companion to Min TPR (Table 2), base-FLORES (left) and instruct-AYA (right), subset all, α=0.01, global threshold τg. With L=11 evaluation languages the bottom quintile is ⌊0.2⋅11⌋=2 languages; the columns report the mean TPR of the two worst-performing languages per cell, with the languages listed minimum first (arbitrary tie-breaking). The Unigram-Gemma base cell is degenerate (all eleven per-language TPRs are zero); its bottom-quintile language identities are reported but uninformative.
Base regime (FLORES+)Instruct regime (AYA)
SchemeGen.Min TPRBot.-q meanBot.-q langsMin TPRBot.-q meanBot.-q langs
KGWMistral0.9780.979zho, hin0.3860.447fra, por
Gemma0.9340.943arb, por0.2220.251por, vie
Qwen0.8800.894tur, nld0.5100.564por, nld
UnigramMistral0.6380.672nld, spa0.0220.029spa, fra
Gemma0.0000.000eng, fra0.0100.010por, tur
Qwen0.7220.744tur, arb0.2100.306arb, hin
SynthIDMistral0.9980.998zho, tur0.7820.788fra, spa
Gemma0.9980.999vie, eng0.4220.437vie, por
Qwen0.9900.992tur, zho0.7920.842hin, por
EXPEditMistral0.9960.997zho, nld0.2380.244eng, spa
Gemma0.9980.998jpn, vie0.0840.101jpn, por
Qwen0.9460.968hin, nld0.1740.226hin, eng
DIPMistral0.9380.950zho, jpn0.2640.265fra, spa
Gemma0.9140.917hin, vie0.1100.115por, jpn
Qwen0.7980.819tur, hin0.2960.345hin, fra
XSIRMistral0.6720.673zho, vie0.0880.116fra, spa
Gemma0.5880.635jpn, hin0.0160.019tur, jpn
Qwen0.0500.064por, fra0.0000.000nld, vie
Table 11: Robustness of the generalized-entropy decomposition to choice of inequality measure: GE2 (half the squared coefficient of variation, more sensitive to inequality at the top of the per-language TPR distribution) versus GE0 (mean log deviation, more sensitive to inequality at the bottom). Base-FLORES (left) and instruct-AYA (right), subset all, α=0.01, global threshold τg; typological partition as in Table 2. Total values below 10−3 reported as <0.001. The Unigram-Gemma base GE0 entry is obtained by replacing zero TPRs with a small ϵ (per the strict-positivity requirement of GE0); the resulting Total=0 and Btw.%=100.0 are regularization artifacts, marked †.
Base regime (FLORES+)Instruct regime (AYA)
GE2GE0GE2GE0
SchemeGen.TotalBtw. %TotalBtw. %TotalBtw. %TotalBtw. %
KGWMistral<0.00183.6<0.00183.70.02695.30.02791.4
Gemma<0.00172.4<0.00172.70.05382.40.04673.3
Qwen0.00185.10.00185.10.01178.60.01172.5
UnigramMistral0.00861.30.00955.70.29799.80.55694.8
Gemma<0.001†100.0†0.22681.20.24373.9
Qwen0.00493.80.00594.60.03291.70.04694.6
SynthIDMistral<0.001100.0<0.001100.00.00377.70.00375.6
Gemma<0.001100.0<0.001100.00.02079.70.01975.6
Qwen<0.00197.5<0.00197.50.00291.90.00292.2
EXPEditMistral<0.00188.0<0.00188.10.06780.00.08768.2
Gemma<0.001100.0<0.001100.00.07184.20.07779.0
Qwen<0.00199.2<0.00199.20.08084.40.09381.4
DIPMistral<0.00192.6<0.00192.80.05899.50.05999.2
Gemma<0.00199.5<0.00199.50.08491.60.06884.4
Qwen0.00394.70.00394.60.03599.60.03599.6
XSIRMistral0.00491.50.00492.10.18591.80.16083.4
Gemma0.00696.00.00796.30.17576.50.22081.1
Qwen0.49398.90.32692.53.78799.93.24479.0
Table 12: Robustness of the GE2 between-share to choice of partition: typological family (8 groups: Germanic, Romance, Indic, Semitic, Sinitic, Japonic, Turkic, Austroasiatic, with 5 of 11 languages in non-singleton groups) vs. script (4 groups: Latin pooling eng, nld, fra, spa, por, tur, vie; Devanagari=hin; Arabic=arb; Han pooling zho, jpn; with 9 of 11 languages in non-singleton groups). Base-FLORES (left) and instruct-AYA (right), subset all, α=0.01, global threshold τg. GE2 Total depends only on the per-language vector and is partition-independent (shown once per regime). Totals below 10−3 reported as <0.001; for these cells the between-share values involve ratios of near-zero variance components and should not be over-interpreted.
Base regime (FLORES+)Instruct regime (AYA)
Btw. %Btw. %
SchemeGen.GE2FamilyScriptGE2FamilyScript
KGWMistral<0.00183.646.20.02695.332.8
Gemma<0.00172.462.20.05382.429.8
Qwen0.00185.126.20.01178.626.7
UnigramMistral0.00861.328.80.29799.846.3
Gemma0.22681.261.7
Qwen0.00493.826.10.03291.775.0
SynthIDMistral<0.001100.017.10.00377.733.0
Gemma<0.001100.05.70.02079.747.1
Qwen<0.00197.57.50.00291.979.3
EXPEditMistral<0.00188.031.70.06780.030.2
Gemma<0.001100.017.10.07184.28.5
Qwen<0.00199.295.80.08084.441.6
DIPMistral<0.00192.669.30.05899.538.9
Gemma<0.00199.548.20.08491.625.2
Qwen0.00394.732.50.03599.653.4
XSIRMistral0.00491.554.80.18591.864.7
Gemma0.00696.012.70.17576.57.6
Qwen0.49398.993.53.78799.999.8
Table 13: Quality fairness summary by scheme, generator, regime, and measurement, subset all. Each row reports mean per-language preservation with the per-cell minimum and floor language (ISO 639-3 superscript) in parentheses. DI<.8 counts languages with per-language preservation below 0.8×maxl. GE2 uses the typological partition (Germanic, Romance, Indic, Semitic, Sinitic, Japonic, Turkic, Austroasiatic); Btw. % is its between-family share. BERTScore mean (min) is reported on rescaled F1 [28]; DI<.8 and GE2 are omitted for BERTScore because neither registers signal on the bounded raw scale and both behave pathologically on the rescaled scale (§5.2).
MAUVEBERTScore F1 (rescaled)PPL preservation (XGLM-7.5B)
SchemeGenmean (min)DI<.8GE2Btw. %mean (min)mean (min)DI<.8GE2Btw. %
Base regime (FLORES+ continuation)
KGWMistral0.947 (0.879)jpn00.00186.0-0.041 (-0.079)vie0.866 (0.793)spa00.00266.4
KGWGemma0.855 (0.435)arb10.01398.9-0.081 (-0.194)vie0.899 (0.812)fra00.00291.0
KGWQwen0.889 (0.471)hin10.01295.6-0.034 (-0.091)tur0.834 (0.686)hin30.00594.0
UnigramMistral0.747 (0.420)arb50.02649.5-0.070 (-0.134)arb0.872 (0.811)spa00.00296.4
UnigramGemma0.638 (0.444)spa60.01767.8-0.098 (-0.191)jpn0.911 (0.815)nld00.00268.3
UnigramQwen0.734 (0.123)hin30.04397.2-0.055 (-0.128)zho0.896 (0.647)hin10.00598.0
SynthIDMistral0.187 (0.020)arb100.43559.0-0.115 (-0.163)vie0.299 (0.145)arb70.04888.4
SynthIDGemma0.095 (0.029)tur100.64143.0-0.146 (-0.254)vie0.269 (0.195)tur70.02173.2
SynthIDQwen0.183 (0.035)vie100.45553.0-0.112 (-0.188)tur0.308 (0.203)arb90.02878.2
EXPEditMistral0.079 (0.006)arb100.53657.1-0.149 (-0.228)vie0.218 (0.079)arb60.08599.1
EXPEditGemma0.013 (0.006)jpn100.57443.4-0.220 (-0.333)vie0.135 (0.076)zho90.03382.7
EXPEditQwen0.067 (0.013)arb100.87447.9-0.145 (-0.245)tur0.224 (0.135)tur90.04274.6
DIPMistral0.973 (0.933)hin00.00095.6-0.015 (-0.040)vie0.809 (0.750)hin00.00177.6
DIPGemma0.932 (0.878)jpn00.00194.7-0.053 (-0.183)vie0.774 (0.702)nld00.00168.0
DIPQwen0.947 (0.864)hin00.00181.6-0.008 (-0.071)tur0.811 (0.714)tur00.00292.7
XSIRMistral0.859 (0.642)hin20.00686.6-0.062 (-0.110)hin0.922 (0.816)nld00.00273.2
XSIRGemma0.683 (0.461)fra60.01880.9-0.123 (-0.226)vie0.925 (0.785)nld10.00275.8
XSIRQwen0.850 (0.603)hin20.00889.9-0.042 (-0.129)tur0.883 (0.745)hin10.00386.4
Instruct regime (AYA native-speaker)
KGWMistral0.989 (0.970)arb00.00058.90.255 (0.088)arb0.892 (0.853)jpn00.00191.0
KGWGemma0.990 (0.959)arb00.00065.30.397 (0.338)tur0.967 (0.941)hin00.00080.0
KGWQwen0.987 (0.954)hin00.00090.80.300 (0.192)tur0.885 (0.749)hin10.00292.5
UnigramMistral0.981 (0.938)nld00.00040.30.240 (0.087)arb0.897 (0.835)hin00.00197.4
UnigramGemma0.996 (0.991)jpn00.00054.10.400 (0.329)tur0.965 (0.941)nld00.00078.0
UnigramQwen0.980 (0.927)hin00.00097.70.297 (0.184)tur0.885 (0.700)hin10.00395.9
SynthIDMistral0.721 (0.066)arb40.10799.60.149 (-0.117)arb0.595 (0.185)arb70.07095.4
SynthIDGemma0.994 (0.979)fra00.00050.00.365 (0.296)tur0.931 (0.885)zho00.00090.0
SynthIDQwen0.833 (0.237)tur20.03397.10.219 (0.047)tur0.645 (0.442)tur50.01490.7
Table 14: Single-covariate R2 for per-language detection TPR at τg, α=0.01, subset all. Each fit regresses the per-language outcome vector on a single covariate over 33 (language, generator) points. Stars from the slope p-value: ∗p<0.05, p∗⁣∗<0.01.
CovariateKGWUnigramSynthIDEXPEditDIPXSIR
Base regime (FLORES+ continuation, n=33)
Fertility (tokens/sent.)0.0160.0520.0080.834∗∗0.238∗∗0.001
NWM surprisal (XGLM)0.0660.0740.0580.261∗∗0.156∗0.003
Latin script0.0290.0010.0050.0510.0090.019
NWM adherence0.600∗∗0.0060.415∗∗0.0010.461∗∗0.091
Instruct regime (AYA native-speaker, n=33)
Fertility (tokens/sent.)0.0170.0460.0090.0010.0000.107
NWM surprisal (XGLM)0.0240.0030.0090.0400.0170.011
Latin script0.0460.0140.0350.0490.0400.058
NWM adherence0.0000.0000.0030.0470.0000.291∗∗
Table 15: Single-covariate R2 for per-language MAUVE preservation. Each fit regresses the per-language outcome vector on a single covariate over 33 (language, generator) points. Stars from the slope p-value: ∗p<0.05, p∗⁣∗<0.01. BERTScore-rescaled and PPL-XGLM paradigm-stability tables in §F.1.
CovariateKGWUnigramSynthIDEXPEditDIPXSIR
Base regime (FLORES+ continuation, n=33)
Fertility (tokens/sent.)0.392∗∗0.249∗∗0.0270.0050.234∗∗0.093
NWM surprisal (XGLM)0.0260.0230.0180.0010.0070.009
Latin script0.136∗0.0360.0620.0520.0830.000
NWM adherence0.0310.0070.1010.0740.219∗∗0.072
Instruct regime (AYA native-speaker, n=33)
Fertility (tokens/sent.)0.346∗∗0.464∗∗0.0270.0160.0190.168∗
NWM surprisal (XGLM)0.0080.152∗0.0220.0520.0050.002
Latin script0.122∗0.0510.0180.0410.0520.082
NWM adherence0.0760.0990.153∗0.230∗∗0.0530.047
Table 16: Single-covariate R2 for BERTScore-rescaled preservation. Stars from the slope p-value: ∗p<0.05, p∗⁣∗<0.01.
CovariateKGWUnigramSynthIDEXPEditDIPXSIR
Base regime (n=33)
Fertility (tokens/sent.)0.0000.0130.0240.0540.0390.095
NWM surprisal (XGLM)0.0360.0160.0000.0320.0390.094
Latin script0.0680.207∗∗0.0040.0220.0030.008
NWM adherence0.239∗∗0.0040.378∗∗0.259∗∗0.298∗∗0.188∗
Instruct regime (n=33)
Fertility (tokens/sent.)0.0720.0690.0430.0260.0470.051
NWM surprisal (XGLM)0.0340.0340.0300.0210.0370.026
Latin script0.0460.0450.0420.0370.0440.056
NWM adherence0.0960.0950.119∗0.125∗0.123∗0.110
Table 17: Single-covariate R2 for PPL-XGLM preservation. Stars from the slope p-value: ∗p<0.05, p∗⁣∗<0.01.
CovariateKGWUnigramSynthIDEXPEditDIPXSIR
Base regime (n=33)
Fertility (tokens/sent.)0.191∗0.327∗∗0.0130.0000.121∗0.148∗
NWM surprisal (XGLM)0.0030.179∗0.0170.0180.0940.183∗
Latin script0.0220.0010.126∗0.1050.0520.056
NWM adherence0.186∗0.0140.164∗0.123∗0.1090.015
Instruct regime (n=33)
Fertility (tokens/sent.)0.446∗∗0.559∗∗0.119∗0.0700.225∗∗0.340∗∗
NWM surprisal (XGLM)0.0440.0290.0100.0000.0470.016
Latin script0.0860.0770.0700.0840.0460.012
NWM adherence0.0200.0240.0960.217∗∗0.0230.006

Why it matters

As AI-generated multilingual content spreads (including flagged multilingual misinformation sites), watermarking is being pitched as a safeguard, but if it silently fails for certain language families, speakers of those languages get weaker protection and pay a bigger quality cost for being watermarked. This work gives developers and policymakers a concrete testing method to catch these gaps before deployment rather than assuming English results generalize.

Terms in this paper

  • 워터마킹(watermarking) · AI가 생성한 텍스트에 눈에 보이지 않는 통계적 흔적을 남겨 나중에 기계 생성물임을 판별하는 기술
  • 탐지 임계값(detection threshold) · 워터마크 점수가 이 값을 넘으면 '워터마크 있음'으로 판정하는 기준선
  • AUC · 임계값에 의존하지 않고 두 그룹(워터마크/비워터마크)이 얼마나 잘 구분되는지 보여주는 지표
  • 언어 유형학적 계열(typological family) · 문법 구조나 어휘 특성이 비슷한 언어들의 묶음, 예: 게르만어파, 셈어파
  • distortion-free 방식 · 이론상 생성되는 텍스트의 확률분포를 바꾸지 않도록 설계된 워터마킹 방식

Original abstract (English)

Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme's detection threshold and a narrow set of quality measurements. Multilingual deployment exposes evaluation-design choices that are inconsequential on English but determine conclusions cross-lingually. We propose an evaluation framework with four components: detection thresholds calibrated empirically per deployment context, a threshold-independent companion measurement that distinguishes calibration failures from detection failures, three disjoint quality measurement paradigms (distributional, paired-semantic, and reference-perplexity), and a generalized-entropy decomposition of cross-language disparity over a typological family partition. Applied to six watermarking schemes, three open-weight generators, eleven languages spanning four scripts and eight typological families, and both base and instruction-tuned regimes, the framework reveals failure modes that single-language single-paradigm evaluation cannot surface. Across detection and quality, observed disparity is predominantly between-family on the typological partition, indicating that cross-lingual fairness gaps in watermarking are structural to language properties rather than idiosyncratic to particular languages.

Authors · Alexander Nemecek, Osama Zafar, Debargha Ganguly, Vikash Singh, Vipin Chaudhary, Erman Ayday

Read on arXiv

Latest papers

All papers →

Latest from METAL LAB

Figures: Alexander Nemecek et al., arXiv:2608.20047, CC BY 4.0