Figure 1: Per-language quality preservation under three paradigms, base-FLORES (top) and instruct-AYA (bottom); n=500 matched pairs per cell. Left: MAUVE on mean-pooled XLM-R-large embeddings (1.0 identical, 0.0 fully separable). Center: BERTScore F1, rescaled against per-language FLORES non-pair baselines (negative: paired generations less similar than random non-pairs). Right: PPL preservation under XGLM-7.5B, exp(−|logPPLR(𝒳+)−logPPLR(𝒳−)|)∈(0,1]. Rows: (scheme, generator) cells; columns: eleven languages grouped by typological family.
Table 1: The eleven evaluation languages with script, typological family, and Joshi resource tier.
ISO 639-3
Language
Script
Typological family
Joshi et al. tier
eng
English
Latin
Germanic
5
nld
Dutch
Latin
Germanic
4
fra
French
Latin
Romance
5
spa
Spanish
Latin
Romance
5
por
Portuguese
Latin
Romance
4
tur
Turkish
Latin
Turkic
4
vie
Vietnamese
Latin
Austroasiatic
4
hin
Hindi
Devanagari
Indic
4
arb
Arabic
Arabic
Semitic
5
zho
Chinese
Han
Sinitic
5
jpn
Japanese
Han + Kana
Japonic
5
Figure 2: Per-scheme disparity decomposition under three quality paradigms, base-FLORES (top) and instruct-AYA (bottom). Bar height is GE2 over the per-language preservation vector, averaged across the three generators per scheme. Stacking shows the typological-partition decomposition into between-family (dark) and within-family (light) components; the percentage above is the between-family share. Absolute disparity is orders of magnitude smaller than under MAUVE or PPL but the between-family pattern persists. Y-axes differ across panels.
Table 2: Detection fairness by scheme, generator, and regime, subset all, α=0.01. Mean TPR shown with 95% bootstrap CIs (B=1000 resamples of the 11-language vector). Min TPR is the realized minimum. DI<.8 counts languages with TPRl<0.8⋅maxlTPRl. GE2 uses the typological partition (Germanic, Romance, Indic, Semitic, Sinitic, Japonic, Turkic, Austroasiatic); Btw. % is its between-family share. Δ=meanTPR(τg)−meanTPR(τl). Values below 10−3 reported as <0.001; entries marked are undefined when the per-language mean is zero.
AUC
TPR at τg (global)
TPR at τl (per-lang)
Scheme
Gen.
Mean
Min
Mean TPR
Min TPR
DI<.8
GE2
Btw. %
Mean TPR
Δ
Base regime (FLORES+ continuation prompts)
KGW
Mistral
0.999
0.997
0.992 [0.987, 0.996]
0.978
0
<0.001
83.6
0.991
+0.001
Gemma
0.997
0.991
0.969 [0.959, 0.979]
0.934
0
<0.001
72.4
0.974
−0.005
Qwen
0.996
0.987
0.965 [0.939, 0.986]
0.880
0
<0.001
85.1
0.966
−0.001
Unigram
Mistral
0.996
0.981
0.868 [0.802, 0.927]
0.638
3
0.008
61.3
0.971
−0.103
Gemma
0.986
0.910
0.000 [0.000, 0.000]
0.000
11
—
—
0.799
−0.799
Qwen
0.994
0.972
0.898 [0.845, 0.942]
0.722
2
0.004
93.8
0.949
−0.051
SynthID
Mistral
1.000
1.000
1.000 [0.999, 1.000]
0.998
0
<0.001
100.0
1.000
−0.000
Gemma
1.000
1.000
1.000 [0.999, 1.000]
0.998
0
<0.001
100.0
1.000
+0.000
Qwen
1.000
0.999
0.997 [0.995, 0.999]
0.990
0
<0.001
97.5
0.998
−0.000
EXPEdit
Mistral
1.000
0.998
0.999 [0.999, 1.000]
0.996
0
<0.001
88.0
1.000
−0.000
Gemma
1.000
0.998
1.000 [0.999, 1.000]
0.998
0
<0.001
100.0
1.000
+0.000
Qwen
0.998
0.993
0.991 [0.982, 0.997]
0.946
0
<0.001
99.2
0.992
−0.001
DIP
Mistral
0.997
0.989
0.975 [0.966, 0.984]
0.938
0
<0.001
92.6
0.975
+0.000
Gemma
0.992
0.965
0.963 [0.947, 0.977]
0.914
0
<0.001
99.5
0.965
−0.002
Qwen
0.992
0.977
0.931 [0.891, 0.968]
0.798
0
0.003
94.7
0.929
+0.002
XSIR
Mistral
0.981
0.962
0.784 [0.743, 0.821]
0.672
2
0.004
91.5
0.833
−0.049
Gemma
0.984
0.977
0.766 [0.712, 0.817]
0.588
4
0.006
96.0
0.829
−0.063
Qwen
0.978
0.930
0.215 [0.118, 0.362]
0.050
10
0.493
98.9
0.794
−0.579
Instruct regime (AYA native-speaker prompts)
KGW
Mistral
0.953
0.909
0.626 [0.540, 0.707]
0.386
7
0.026
95.3
0.632
−0.005
Gemma
0.900
0.848
0.389 [0.325, 0.463]
0.222
10
0.053
82.4
0.411
−0.022
Qwen
0.966
0.937
0.689 [0.628, 0.753]
0.510
8
0.011
78.6
0.723
−0.034
Unigram
Mistral
0.909
0.865
0.311 [0.171, 0.453]
0.022
6
0.297
99.8
0.450
−0.139
Gemma
0.851
0.790
0.037 [0.023, 0.052]
0.010
10
0.226
81.2
0.151
−0.113
Qwen
0.943
0.880
0.546 [0.455, 0.621]
0.210
4
0.032
91.7
0.636
−0.090
SynthID
Mistral
0.991
0.982
0.905 [0.865, 0.940]
0.782
1
0.003
77.7
0.903
+0.002
Gemma
0.946
0.902
0.602 [0.530, 0.675]
0.422
10
0.020
79.7
0.597
+0.005
Qwen
0.993
0.977
0.941 [0.906, 0.971]
0.792
1
0.002
91.9
0.942
−0.002
Figure 3: Joint per-language detection-quality landscape, MAUVE paradigm. Markers: per-language (TPR@τg,MAUVE) pairs at α=0.01. Circles: base-FLORES. Triangles: instruct-AYA. Color: typological family. Annotations sB, sI: mean pairwise Euclidean distance within each regime. BERTScore and PPL-preservation companion panels in Appendix E.
Table 3: Cross-paradigm rank agreement: per-cell Spearman ρ between per-language preservation vectors, mean across (model, regime) cells with per-cell minimum in parentheses (subset=all; n=6 cells per scheme, n=36 pooled). Pairs use BERTScore F1 (rescaled) and PPL preservation under XGLM-7.5B; BS-raw gives nearly identical rankings (ρBS-raw↔BS-rescaled=+0.87). Negative minima identify cells where two paradigms produce opposite per-language orderings.
Scheme
MAUVE ↔ BS
MAUVE ↔ PPL
BS ↔ PPL
KGW
+0.61 (+0.45)
+0.03 (−0.80)
+0.03 (−0.76)
Unigram
+0.56 (+0.23)
+0.14 (−0.39)
+0.20 (−0.67)
SynthID
+0.76 (+0.47)
+0.81 (+0.27)
+0.77 (+0.45)
EXPEdit
+0.68 (+0.26)
+0.75 (+0.39)
+0.70 (+0.46)
DIP
+0.46 (−0.03)
+0.43 (+0.05)
+0.50 (−0.05)
XSIR
+0.42 (−0.09)
+0.10 (−0.71)
+0.00 (−0.76)
All schemes
+0.58 (−0.09)
+0.38 (−0.80)
+0.37 (−0.76)
Figure 4: Per-language detection diagnostics, subset all, α=0.01. Within each panel, left: TPR at the global empirical-FPR threshold τg; right: threshold-free AUC. Rows are (scheme, generator) cells; columns are the eleven evaluation languages, grouped by typological family.
Table 4: Per-cell language adherence on the eleven evaluation languages, base-FLORES (left) and instruct-AYA (right). NWM% is the GlotLID-v3 on-target rate for nowatermark generations; Δ¯ is the mean shift in percentage points across the six watermarked schemes. mean rows are unweighted means across the eleven languages.
Base regime (FLORES+ continuation prompts)
Instruct regime (AYA native-speaker prompts)
Mistral
Gemma
Qwen
Mistral
Gemma
Qwen
Lang
NWM%
Δ¯
NWM%
Δ¯
NWM%
Δ¯
NWM%
Δ¯
NWM%
Δ¯
NWM%
Δ¯
eng
99.6
−0.3
98.8
−2.0
98.4
−0.6
92.6
−0.9
80.6
−1.7
93.2
−2.0
fra
99.2
−1.1
97.6
−4.0
89.2
−0.2
91.4
+0.5
86.2
−1.7
94.8
−4.9
spa
100.0
−1.4
96.8
−2.6
91.2
+1.1
95.2
−1.3
62.6
+2.7
92.0
−5.0
por
99.2
−0.5
96.8
−1.0
98.2
−0.8
96.2
−2.0
70.6
+3.0
96.4
−4.3
nld
96.2
−7.0
95.8
+1.6
80.8
−4.6
73.2
−3.1
89.2
−0.4
90.4
−5.8
zho
97.6
−1.8
92.8
−1.2
96.4
−4.2
89.0
−7.9
95.6
−0.4
98.4
−0.6
jpn
96.8
−3.3
95.6
−1.6
99.8
−3.3
71.0
−10.7
94.0
−1.9
98.4
−7.4
arb
89.2
−3.0
87.6
−3.8
87.8
+2.9
67.6
−15.6
84.2
−1.2
81.0
−15.0
hin
96.2
−9.1
88.2
−2.3
98.8
−3.9
52.2
−6.5
83.8
−1.4
83.2
+0.1
tur
89.0
+1.3
96.8
−4.3
57.2
−1.1
81.8
−8.1
92.8
−0.6
95.6
−13.5
vie
98.6
−4.7
77.0
+4.7
99.0
−1.0
79.2
−7.3
98.8
−0.6
95.2
−6.4
mean
96.5
−2.8
93.1
−1.5
90.6
−1.4
80.9
−5.7
85.3
−0.4
92.6
−5.9
(b) Instruct-AYA regime.
Table 5: Length-distribution similarity between LID subsets and the all-subset baseline. Each row is a (regime, subset) pair; ratio=subset-mean length/all-subset mean length, computed per (regime, model, language, watermark) cell and summarized across cells. Cells require n≥100 in both subsets to enter the summary; n skipped counts cells that fail this floor (typically off_target on high-adherence cells).
Regime
Subset
n kept
n skipped
Mean
Median
Min
Max
Base
on_target_only
231
0
1.003
1.000
0.996
1.081
Base
off_target_only
8
223
0.929
0.908
0.859
1.023
Instruct
on_target_only
231
0
1.000
1.000
0.964
1.057
Instruct
off_target_only
26
205
1.025
1.017
0.947
1.119
Figure 5: PPL preservation under mGPT-1.3B cross-reference. Format and color scale match Figure 1.
Table 6: LID categorization robustness: per-cell adherence under a 3×3 sweep of the GlotLID per-sentence confidence floor (conf_floor) and the per-text on-target sentence fraction (on_target_ratio). ρ is the Spearman rank-correlation between the per-cell on-target-rate vector under each configuration and the default (†). Mean on-target % is the unweighted mean across the 462 cells of the evaluation grid. Mean/Max |Δ| are the mean and maximum absolute per-cell shift in on-target rate (pp) relative to default.
conf_floor
on_target_ratio
ρ vs default
Mean on-target %
Mean |Δ| (pp)
Max |Δ| (pp)
0.3
0.7
0.9347
92.0
2.19
13.60
0.3
0.8
0.9468
90.7
0.91
13.00
0.3
0.9
0.8627
86.3
4.87
31.00
0.5
0.7
0.9910
91.1
1.26
11.80
0.5†
0.8
1.0000
89.8
0.00
0.00
0.5
0.9
0.9317
85.4
4.43
31.20
0.7
0.7
0.9486
87.9
3.76
24.60
0.7
0.8
0.9627
86.8
3.04
24.80
0.7
0.9
0.9385
82.5
7.36
31.80
(b) Instruct-AYA regime.
Table 7: Per-language calibration gap Δl=TPRl(τl)−TPRl(τg), base-FLORES (top) and instruct-AYA (bottom), subset all, α=0.01. Positive: per-language calibration recovers detection. Negative: per-language calibration worsens detection. Zero: detection insensitive to calibration locality. Row means equal −1× the Δ column of Table 2.
Germanic
Romance
Indic
Semitic
Sinitic
Japonic
Turkic
Aus.
Scheme
Gen.
eng
nld
fra
spa
por
hin
arb
zho
jpn
tur
vie
Base regime (FLORES+ continuation prompts)
KGW
Mistral
0.000
+0.004
+0.004
+0.002
0.000
−0.008
0.000
−0.006
0.000
−0.004
+0.002
Gemma
+0.006
+0.030
+0.008
+0.016
+0.038
−0.020
+0.044
−0.010
−0.064
0.000
+0.010
Qwen
+0.028
+0.018
−0.002
−0.002
+0.002
−0.008
0.000
+0.008
0.000
−0.030
0.000
Unigram
Mistral
+0.074
+0.354
+0.218
+0.292
+0.142
−0.062
+0.004
+0.064
−0.018
+0.074
−0.012
Gemma
+0.996
+0.992
+0.968
+0.990
+0.960
0.000
+0.958
+0.976
+0.984
+0.960
0.000
Qwen
+0.024
+0.034
−0.040
+0.002
+0.022
+0.106
+0.230
+0.134
−0.008
+0.092
−0.034
SynthID
Mistral
0.000
0.000
0.000
0.000
0.000
0.000
0.000
+0.002
0.000
0.000
0.000
Gemma
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
Qwen
0.000
0.000
0.000
0.000
0.000
+0.002
0.000
+0.006
0.000
−0.004
0.000
EXPEdit
Mistral
0.000
+0.002
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
Gemma
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
0.000
Qwen
+0.002
−0.006
0.000
0.000
0.000
+0.012
0.000
0.000
0.000
+0.002
0.000
DIP
Mistral
0.000
0.000
+0.002
+0.002
+0.002
+0.002
0.000
−0.014
0.000
+0.004
0.000
Gemma
0.000
−0.004
+0.002
0.000
+0.002
+0.010
+0.004
+0.004
+0.018
−0.010
0.000
Qwen
0.000
−0.008
+0.004
+0.008
−0.004
−0.004
+0.002
+0.010
0.000
−0.028
+0.002
XSIR
Mistral
+0.024
−0.054
+0.166
+0.106
+0.086
−0.022
−0.066
+0.092
+0.014
+0.014
+0.178
Gemma
+0.028
+0.054
+0.054
+0.114
+0.026
+0.146
+0.056
−0.042
+0.168
+0.122
−0.030
Qwen
+0.516
+0.774
+0.770
+0.802
+0.912
−0.286
+0.684
+0.638
+0.464
+0.360
+0.732
Instruct regime (AYA native-speaker prompts)
KGW
Mistral
+0.112
+0.032
−0.250
+0.028
+0.118
−0.016
+0.014
+0.030
−0.018
+0.012
−0.004
Gemma
+0.030
−0.070
+0.108
+0.046
+0.112
−0.038
0.000
−0.036
−0.082
+0.134
+0.034
Qwen
+0.012
+0.014
−0.036
−0.058
+0.110
−0.106
+0.138
+0.016
+0.084
+0.058
+0.140
Unigram
Mistral
+0.340
+0.380
+0.090
+0.282
+0.316
−0.106
0.000
+0.070
−0.178
+0.282
+0.052
Gemma
+0.024
−0.054
+0.174
+0.178
+0.056
−0.070
−0.036
+0.376
+0.276
+0.272
+0.050
Qwen
+0.084
−0.018
−0.108
−0.012
+0.078
+0.128
+0.438
+0.352
+0.094
+0.088
−0.132
SynthID
Mistral
+0.056
+0.020
−0.026
−0.034
−0.048
+0.016
0.000
0.000
−0.022
0.000
+0.016
Gemma
+0.050
−0.030
+0.046
−0.022
−0.020
−0.022
−0.124
+0.038
−0.102
+0.094
+0.036
Qwen
+0.018
−0.002
+0.010
−0.022
−0.002
+0.016
0.000
0.000
0.000
0.000
0.000
Figure 6: On-target minus all-subset preservation, base-FLORES regime (top) and instruct-AYA regime (bottom). Each cell is Δ=preservationon-target−preservationall for the indicated paradigm; rows are (scheme, generator), columns are languages. Positive (red): on-target restriction raises preservation; negative (blue): on-target restriction lowers it. Color scale fixed at [−0.3,+0.3]; values outside this range saturate. The EXPEdit-Mistral, Arabic, BERTScore F1 entry in the instruct panel is omitted (—) because the on-target sample size falls below the n≥100 paired-set floor.
Table 8: Detection fairness summary at α=0.05, base-FLORES (top) and instruct-AYA (bottom), subset all. Format mirrors Table 2; AUC columns are omitted (threshold-independent, identical to Table 2) and bootstrap CIs on Mean TPR are omitted for compactness. L=11 throughout. Values below 10−3 reported as <0.001.
TPR at τg (global)
TPR at τl (per-lang)
Scheme
Gen.
Mean TPR
Min TPR
DI<.8
GE2
Btw. %
Mean TPR
Δ
Base regime (FLORES+ continuation prompts)
KGW
Mistral
0.997
0.992
0
<0.001
61.1
0.997
+0.000
Gemma
0.991
0.980
0
<0.001
93.2
0.993
−0.002
Qwen
0.987
0.952
0
<0.001
57.8
0.986
+0.000
Unigram
Mistral
0.977
0.920
0
<0.001
54.9
0.988
−0.011
Gemma
0.988
0.976
0
<0.001
55.9
0.898
+0.089
Qwen
0.961
0.860
0
0.001
86.3
0.979
−0.018
SynthID
Mistral
1.000
0.998
0
<0.001
100.0
1.000
+0.000
Gemma
1.000
0.998
0
<0.001
100.0
1.000
+0.000
Qwen
0.999
0.996
0
<0.001
74.3
0.999
+0.000
EXPEdit
Mistral
1.000
0.996
0
<0.001
100.0
1.000
+0.000
Gemma
1.000
0.998
0
<0.001
100.0
1.000
+0.000
Qwen
0.995
0.966
0
<0.001
99.2
0.995
−0.000
DIP
Mistral
0.990
0.972
0
<0.001
94.7
0.989
+0.001
Gemma
0.981
0.942
0
<0.001
97.9
0.981
−0.001
Qwen
0.966
0.896
0
0.001
91.6
0.966
+0.000
XSIR
Mistral
0.924
0.868
0
0.001
83.3
0.935
−0.011
Gemma
0.929
0.840
0
0.001
92.8
0.944
−0.015
Qwen
0.503
0.276
9
0.076
94.6
0.919
−0.416
Instruct regime (AYA native-speaker prompts)
KGW
Mistral
0.801
0.622
2
0.006
85.0
0.813
−0.012
Gemma
0.616
0.474
10
0.015
79.7
0.617
−0.001
Qwen
0.860
0.740
1
0.003
78.1
0.869
−0.009
Unigram
Mistral
0.517
0.090
5
0.149
95.8
0.659
−0.143
Gemma
0.380
0.154
5
0.050
80.3
0.408
−0.028
Qwen
0.727
0.366
4
0.018
93.8
0.790
−0.063
SynthID
Mistral
0.967
0.926
0
<0.001
77.8
0.969
−0.002
Gemma
0.791
0.662
4
0.005
75.9
0.785
+0.005
Qwen
0.971
0.898
0
<0.001
95.7
0.973
−0.002
Figure 7: Joint per-language detection–quality landscape, BERTScore-rescaled (top) and PPL-XGLM (bottom) paradigms. Format matches Figure 3: markers are per-language (TPR@τg,quality) pairs at α=0.01, n=33 per regime (11 languages × 3 generators); circles are base-FLORES, triangles are instruct-AYA; color encodes typological family; sB, sI are mean pairwise Euclidean distance within each regime.
Table 9: Detection summary on the on-target subset, α=0.01, global threshold τg, base-FLORES (left) and instruct-AYA (right). Thresholds re-calibrated on each cell’s on-target unwatermarked pool; L=11 throughout. The off-target subset is omitted: in the base regime no cell meets the nneg≥200 inclusion floor, and in the instruct regime only six (Mistral-generator) cells pass the relaxed α=0.05 floor with L≤3, insufficient for the fairness aggregates this appendix reports. The Δ vs. all columns report change in Mean TPR relative to subset all (Table 2).
Base regime (FLORES+ continuation)
Instruct regime (AYA native-speaker)
Scheme
Gen.
Mean TPR
Min TPR
DI<.8
Δ vs. all
Mean TPR
Min TPR
DI<.8
Δ vs. all
KGW
Mistral
0.995
0.983
0
+0.003
0.681
0.425
6
+0.055
Gemma
0.973
0.939
0
+0.004
0.426
0.264
10
+0.037
Qwen
0.987
0.963
0
+0.022
0.697
0.520
8
+0.008
Unigram
Mistral
0.885
0.654
2
+0.017
0.388
0.011
7
+0.077
Gemma
0.987
0.974
0
+0.987
0.020
0.000
9
−0.017
Qwen
0.925
0.762
1
+0.027
0.547
0.176
4
+0.001
SynthID
Mistral
1.000
0.998
0
0.000
0.924
0.774
1
+0.019
Gemma
1.000
0.998
0
0.000
0.610
0.426
10
+0.008
Qwen
0.999
0.996
0
+0.002
0.941
0.764
1
0.000
EXPEdit
Mistral
1.000
1.000
0
+0.001
0.675
0.234
5
+0.098
Gemma
1.000
0.997
0
0.000
0.232
0.091
10
+0.007
Qwen
0.994
0.941
0
+0.003
0.508
0.059
10
+0.012
DIP
Mistral
0.980
0.952
0
+0.005
0.488
0.266
9
+0.033
Gemma
0.977
0.954
0
+0.014
0.198
0.105
10
+0.002
Qwen
0.957
0.848
0
+0.026
0.507
0.200
8
−0.006
XSIR
Mistral
0.799
0.688
2
+0.015
0.296
0.081
10
+0.030
Gemma
0.793
0.600
3
+0.027
0.092
0.019
10
+0.011
Qwen
0.207
0.047
10
−0.008
0.047
0.000
10
+0.004
Table 10: Bottom-quintile-mean Rawlsian floor as a stability companion to Min TPR (Table 2), base-FLORES (left) and instruct-AYA (right), subset all, α=0.01, global threshold τg. With L=11 evaluation languages the bottom quintile is ⌊0.2⋅11⌋=2 languages; the columns report the mean TPR of the two worst-performing languages per cell, with the languages listed minimum first (arbitrary tie-breaking). The Unigram-Gemma base cell is degenerate (all eleven per-language TPRs are zero); its bottom-quintile language identities are reported but uninformative.
Base regime (FLORES+)
Instruct regime (AYA)
Scheme
Gen.
Min TPR
Bot.-q mean
Bot.-q langs
Min TPR
Bot.-q mean
Bot.-q langs
KGW
Mistral
0.978
0.979
zho, hin
0.386
0.447
fra, por
Gemma
0.934
0.943
arb, por
0.222
0.251
por, vie
Qwen
0.880
0.894
tur, nld
0.510
0.564
por, nld
Unigram
Mistral
0.638
0.672
nld, spa
0.022
0.029
spa, fra
Gemma
0.000
0.000
eng, fra
0.010
0.010
por, tur
Qwen
0.722
0.744
tur, arb
0.210
0.306
arb, hin
SynthID
Mistral
0.998
0.998
zho, tur
0.782
0.788
fra, spa
Gemma
0.998
0.999
vie, eng
0.422
0.437
vie, por
Qwen
0.990
0.992
tur, zho
0.792
0.842
hin, por
EXPEdit
Mistral
0.996
0.997
zho, nld
0.238
0.244
eng, spa
Gemma
0.998
0.998
jpn, vie
0.084
0.101
jpn, por
Qwen
0.946
0.968
hin, nld
0.174
0.226
hin, eng
DIP
Mistral
0.938
0.950
zho, jpn
0.264
0.265
fra, spa
Gemma
0.914
0.917
hin, vie
0.110
0.115
por, jpn
Qwen
0.798
0.819
tur, hin
0.296
0.345
hin, fra
XSIR
Mistral
0.672
0.673
zho, vie
0.088
0.116
fra, spa
Gemma
0.588
0.635
jpn, hin
0.016
0.019
tur, jpn
Qwen
0.050
0.064
por, fra
0.000
0.000
nld, vie
Table 11: Robustness of the generalized-entropy decomposition to choice of inequality measure: GE2 (half the squared coefficient of variation, more sensitive to inequality at the top of the per-language TPR distribution) versus GE0 (mean log deviation, more sensitive to inequality at the bottom). Base-FLORES (left) and instruct-AYA (right), subset all, α=0.01, global threshold τg; typological partition as in Table 2. Total values below 10−3 reported as <0.001. The Unigram-Gemma base GE0 entry is obtained by replacing zero TPRs with a small ϵ (per the strict-positivity requirement of GE0); the resulting Total=0 and Btw.%=100.0 are regularization artifacts, marked †.
Base regime (FLORES+)
Instruct regime (AYA)
GE2
GE0
GE2
GE0
Scheme
Gen.
Total
Btw. %
Total
Btw. %
Total
Btw. %
Total
Btw. %
KGW
Mistral
<0.001
83.6
<0.001
83.7
0.026
95.3
0.027
91.4
Gemma
<0.001
72.4
<0.001
72.7
0.053
82.4
0.046
73.3
Qwen
0.001
85.1
0.001
85.1
0.011
78.6
0.011
72.5
Unigram
Mistral
0.008
61.3
0.009
55.7
0.297
99.8
0.556
94.8
Gemma
—
—
<0.001†
100.0†
0.226
81.2
0.243
73.9
Qwen
0.004
93.8
0.005
94.6
0.032
91.7
0.046
94.6
SynthID
Mistral
<0.001
100.0
<0.001
100.0
0.003
77.7
0.003
75.6
Gemma
<0.001
100.0
<0.001
100.0
0.020
79.7
0.019
75.6
Qwen
<0.001
97.5
<0.001
97.5
0.002
91.9
0.002
92.2
EXPEdit
Mistral
<0.001
88.0
<0.001
88.1
0.067
80.0
0.087
68.2
Gemma
<0.001
100.0
<0.001
100.0
0.071
84.2
0.077
79.0
Qwen
<0.001
99.2
<0.001
99.2
0.080
84.4
0.093
81.4
DIP
Mistral
<0.001
92.6
<0.001
92.8
0.058
99.5
0.059
99.2
Gemma
<0.001
99.5
<0.001
99.5
0.084
91.6
0.068
84.4
Qwen
0.003
94.7
0.003
94.6
0.035
99.6
0.035
99.6
XSIR
Mistral
0.004
91.5
0.004
92.1
0.185
91.8
0.160
83.4
Gemma
0.006
96.0
0.007
96.3
0.175
76.5
0.220
81.1
Qwen
0.493
98.9
0.326
92.5
3.787
99.9
3.244
79.0
Table 12: Robustness of the GE2 between-share to choice of partition: typological family (8 groups: Germanic, Romance, Indic, Semitic, Sinitic, Japonic, Turkic, Austroasiatic, with 5 of 11 languages in non-singleton groups) vs. script (4 groups: Latin pooling eng, nld, fra, spa, por, tur, vie; Devanagari=hin; Arabic=arb; Han pooling zho, jpn; with 9 of 11 languages in non-singleton groups). Base-FLORES (left) and instruct-AYA (right), subset all, α=0.01, global threshold τg. GE2 Total depends only on the per-language vector and is partition-independent (shown once per regime). Totals below 10−3 reported as <0.001; for these cells the between-share values involve ratios of near-zero variance components and should not be over-interpreted.
Base regime (FLORES+)
Instruct regime (AYA)
Btw. %
Btw. %
Scheme
Gen.
GE2
Family
Script
GE2
Family
Script
KGW
Mistral
<0.001
83.6
46.2
0.026
95.3
32.8
Gemma
<0.001
72.4
62.2
0.053
82.4
29.8
Qwen
0.001
85.1
26.2
0.011
78.6
26.7
Unigram
Mistral
0.008
61.3
28.8
0.297
99.8
46.3
Gemma
—
—
—
0.226
81.2
61.7
Qwen
0.004
93.8
26.1
0.032
91.7
75.0
SynthID
Mistral
<0.001
100.0
17.1
0.003
77.7
33.0
Gemma
<0.001
100.0
5.7
0.020
79.7
47.1
Qwen
<0.001
97.5
7.5
0.002
91.9
79.3
EXPEdit
Mistral
<0.001
88.0
31.7
0.067
80.0
30.2
Gemma
<0.001
100.0
17.1
0.071
84.2
8.5
Qwen
<0.001
99.2
95.8
0.080
84.4
41.6
DIP
Mistral
<0.001
92.6
69.3
0.058
99.5
38.9
Gemma
<0.001
99.5
48.2
0.084
91.6
25.2
Qwen
0.003
94.7
32.5
0.035
99.6
53.4
XSIR
Mistral
0.004
91.5
54.8
0.185
91.8
64.7
Gemma
0.006
96.0
12.7
0.175
76.5
7.6
Qwen
0.493
98.9
93.5
3.787
99.9
99.8
Table 13: Quality fairness summary by scheme, generator, regime, and measurement, subset all. Each row reports mean per-language preservation with the per-cell minimum and floor language (ISO 639-3 superscript) in parentheses. DI<.8 counts languages with per-language preservation below 0.8×maxl. GE2 uses the typological partition (Germanic, Romance, Indic, Semitic, Sinitic, Japonic, Turkic, Austroasiatic); Btw. % is its between-family share. BERTScore mean (min) is reported on rescaled F1 [28]; DI<.8 and GE2 are omitted for BERTScore because neither registers signal on the bounded raw scale and both behave pathologically on the rescaled scale (§5.2).
MAUVE
BERTScore F1 (rescaled)
PPL preservation (XGLM-7.5B)
Scheme
Gen
mean (min)
DI<.8
GE2
Btw. %
mean (min)
mean (min)
DI<.8
GE2
Btw. %
Base regime (FLORES+ continuation)
KGW
Mistral
0.947 (0.879)jpn
0
0.001
86.0
-0.041 (-0.079)vie
0.866 (0.793)spa
0
0.002
66.4
KGW
Gemma
0.855 (0.435)arb
1
0.013
98.9
-0.081 (-0.194)vie
0.899 (0.812)fra
0
0.002
91.0
KGW
Qwen
0.889 (0.471)hin
1
0.012
95.6
-0.034 (-0.091)tur
0.834 (0.686)hin
3
0.005
94.0
Unigram
Mistral
0.747 (0.420)arb
5
0.026
49.5
-0.070 (-0.134)arb
0.872 (0.811)spa
0
0.002
96.4
Unigram
Gemma
0.638 (0.444)spa
6
0.017
67.8
-0.098 (-0.191)jpn
0.911 (0.815)nld
0
0.002
68.3
Unigram
Qwen
0.734 (0.123)hin
3
0.043
97.2
-0.055 (-0.128)zho
0.896 (0.647)hin
1
0.005
98.0
SynthID
Mistral
0.187 (0.020)arb
10
0.435
59.0
-0.115 (-0.163)vie
0.299 (0.145)arb
7
0.048
88.4
SynthID
Gemma
0.095 (0.029)tur
10
0.641
43.0
-0.146 (-0.254)vie
0.269 (0.195)tur
7
0.021
73.2
SynthID
Qwen
0.183 (0.035)vie
10
0.455
53.0
-0.112 (-0.188)tur
0.308 (0.203)arb
9
0.028
78.2
EXPEdit
Mistral
0.079 (0.006)arb
10
0.536
57.1
-0.149 (-0.228)vie
0.218 (0.079)arb
6
0.085
99.1
EXPEdit
Gemma
0.013 (0.006)jpn
10
0.574
43.4
-0.220 (-0.333)vie
0.135 (0.076)zho
9
0.033
82.7
EXPEdit
Qwen
0.067 (0.013)arb
10
0.874
47.9
-0.145 (-0.245)tur
0.224 (0.135)tur
9
0.042
74.6
DIP
Mistral
0.973 (0.933)hin
0
0.000
95.6
-0.015 (-0.040)vie
0.809 (0.750)hin
0
0.001
77.6
DIP
Gemma
0.932 (0.878)jpn
0
0.001
94.7
-0.053 (-0.183)vie
0.774 (0.702)nld
0
0.001
68.0
DIP
Qwen
0.947 (0.864)hin
0
0.001
81.6
-0.008 (-0.071)tur
0.811 (0.714)tur
0
0.002
92.7
XSIR
Mistral
0.859 (0.642)hin
2
0.006
86.6
-0.062 (-0.110)hin
0.922 (0.816)nld
0
0.002
73.2
XSIR
Gemma
0.683 (0.461)fra
6
0.018
80.9
-0.123 (-0.226)vie
0.925 (0.785)nld
1
0.002
75.8
XSIR
Qwen
0.850 (0.603)hin
2
0.008
89.9
-0.042 (-0.129)tur
0.883 (0.745)hin
1
0.003
86.4
Instruct regime (AYA native-speaker)
KGW
Mistral
0.989 (0.970)arb
0
0.000
58.9
0.255 (0.088)arb
0.892 (0.853)jpn
0
0.001
91.0
KGW
Gemma
0.990 (0.959)arb
0
0.000
65.3
0.397 (0.338)tur
0.967 (0.941)hin
0
0.000
80.0
KGW
Qwen
0.987 (0.954)hin
0
0.000
90.8
0.300 (0.192)tur
0.885 (0.749)hin
1
0.002
92.5
Unigram
Mistral
0.981 (0.938)nld
0
0.000
40.3
0.240 (0.087)arb
0.897 (0.835)hin
0
0.001
97.4
Unigram
Gemma
0.996 (0.991)jpn
0
0.000
54.1
0.400 (0.329)tur
0.965 (0.941)nld
0
0.000
78.0
Unigram
Qwen
0.980 (0.927)hin
0
0.000
97.7
0.297 (0.184)tur
0.885 (0.700)hin
1
0.003
95.9
SynthID
Mistral
0.721 (0.066)arb
4
0.107
99.6
0.149 (-0.117)arb
0.595 (0.185)arb
7
0.070
95.4
SynthID
Gemma
0.994 (0.979)fra
0
0.000
50.0
0.365 (0.296)tur
0.931 (0.885)zho
0
0.000
90.0
SynthID
Qwen
0.833 (0.237)tur
2
0.033
97.1
0.219 (0.047)tur
0.645 (0.442)tur
5
0.014
90.7
Table 14: Single-covariate R2 for per-language detection TPR at τg, α=0.01, subset all. Each fit regresses the per-language outcome vector on a single covariate over 33 (language, generator) points. Stars from the slope p-value: ∗p<0.05, p∗∗<0.01.
Covariate
KGW
Unigram
SynthID
EXPEdit
DIP
XSIR
Base regime (FLORES+ continuation, n=33)
Fertility (tokens/sent.)
0.016
0.052
0.008
0.834∗∗
0.238∗∗
0.001
NWM surprisal (XGLM)
0.066
0.074
0.058
0.261∗∗
0.156∗
0.003
Latin script
0.029
0.001
0.005
0.051
0.009
0.019
NWM adherence
0.600∗∗
0.006
0.415∗∗
0.001
0.461∗∗
0.091
Instruct regime (AYA native-speaker, n=33)
Fertility (tokens/sent.)
0.017
0.046
0.009
0.001
0.000
0.107
NWM surprisal (XGLM)
0.024
0.003
0.009
0.040
0.017
0.011
Latin script
0.046
0.014
0.035
0.049
0.040
0.058
NWM adherence
0.000
0.000
0.003
0.047
0.000
0.291∗∗
Table 15: Single-covariate R2 for per-language MAUVE preservation. Each fit regresses the per-language outcome vector on a single covariate over 33 (language, generator) points. Stars from the slope p-value: ∗p<0.05, p∗∗<0.01. BERTScore-rescaled and PPL-XGLM paradigm-stability tables in §F.1.
Covariate
KGW
Unigram
SynthID
EXPEdit
DIP
XSIR
Base regime (FLORES+ continuation, n=33)
Fertility (tokens/sent.)
0.392∗∗
0.249∗∗
0.027
0.005
0.234∗∗
0.093
NWM surprisal (XGLM)
0.026
0.023
0.018
0.001
0.007
0.009
Latin script
0.136∗
0.036
0.062
0.052
0.083
0.000
NWM adherence
0.031
0.007
0.101
0.074
0.219∗∗
0.072
Instruct regime (AYA native-speaker, n=33)
Fertility (tokens/sent.)
0.346∗∗
0.464∗∗
0.027
0.016
0.019
0.168∗
NWM surprisal (XGLM)
0.008
0.152∗
0.022
0.052
0.005
0.002
Latin script
0.122∗
0.051
0.018
0.041
0.052
0.082
NWM adherence
0.076
0.099
0.153∗
0.230∗∗
0.053
0.047
Table 16: Single-covariate R2 for BERTScore-rescaled preservation. Stars from the slope p-value: ∗p<0.05, p∗∗<0.01.
Covariate
KGW
Unigram
SynthID
EXPEdit
DIP
XSIR
Base regime (n=33)
Fertility (tokens/sent.)
0.000
0.013
0.024
0.054
0.039
0.095
NWM surprisal (XGLM)
0.036
0.016
0.000
0.032
0.039
0.094
Latin script
0.068
0.207∗∗
0.004
0.022
0.003
0.008
NWM adherence
0.239∗∗
0.004
0.378∗∗
0.259∗∗
0.298∗∗
0.188∗
Instruct regime (n=33)
Fertility (tokens/sent.)
0.072
0.069
0.043
0.026
0.047
0.051
NWM surprisal (XGLM)
0.034
0.034
0.030
0.021
0.037
0.026
Latin script
0.046
0.045
0.042
0.037
0.044
0.056
NWM adherence
0.096
0.095
0.119∗
0.125∗
0.123∗
0.110
Table 17: Single-covariate R2 for PPL-XGLM preservation. Stars from the slope p-value: ∗p<0.05, p∗∗<0.01.
Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme's detection threshold and a narrow set of quality measurements. Multilingual deployment exposes evaluation-design choices that are inconsequential on English but determine conclusions cross-lingually. We propose an evaluation framework with four components: detection thresholds calibrated empirically per deployment context, a threshold-independent companion measurement that distinguishes calibration failures from detection failures, three disjoint quality measurement paradigms (distributional, paired-semantic, and reference-perplexity), and a generalized-entropy decomposition of cross-language disparity over a typological family partition. Applied to six watermarking schemes, three open-weight generators, eleven languages spanning four scripts and eight typological families, and both base and instruction-tuned regimes, the framework reveals failure modes that single-language single-paradigm evaluation cannot surface. Across detection and quality, observed disparity is predominantly between-family on the typological partition, indicating that cross-lingual fairness gaps in watermarking are structural to language properties rather than idiosyncratic to particular languages.