每天早上一封邮件,把昨天的 AI 梳理好订阅邮件

METAL LAB

Auditing Cross-Lingual Fairness in Language Model Watermarking

arXiv:2608.200472026-08-21

本该识别AI生成文本的水印技术在非英语语言中表现明显更差,而且这种差距按语系而非单个语言呈现

为了事后能识别出AI生成的文本,业界越来越多地在生成内容中嵌入肉眼不可见的统计标记,即水印,但此前几乎所有测试都只用英语进行。这项研究搭建了一套新的评测框架,在11种语言、6种水印方案和3个开源大模型上进行测试,发现无论是检测效果还是文本质量保留,都会因语言不同而出现明显且不均衡的下降。这种不均衡是结构性的,源于相近语言共有的特征,而不是某个语言的个别问题。

他们做了什么

  1. 研究覆盖11种语言(涉及4种文字体系、8个语系)、6种水印方案、3个开源大模型,并分别在基础续写模式和指令微调(类似聊天机器人)模式下进行测试。
  2. 团队没有直接使用各水印方案自带的、针对英语调优的判定阈值,而是为每种语言和模型重新校准检测阈值,并专门区分了水印信号是真的不存在,还是仅仅阈值设置不合适。
  3. 文本质量的评估同时使用了三种互相独立的方法,即整体分布相似度、逐句语义相似度和基于参照模型的流畅度,以避免单一指标掩盖问题。
  4. 两种号称理论上不改变文本分布的水印方案SynthID-Text和EXPEdit,在实际测试中反而对文本质量破坏最大;而在指令微调模式下,大多数方案的检测性能都明显下降。
  5. 统计分析显示,语言间的不公平主要来自语系层面的差异,比如日耳曼语系、闪米特语系、汉藏语系之间的差异,而不是某一个语言的特殊情况,说明问题内嵌于语言结构本身。
Figure 1: Per-language quality preservation under three paradigms, base-FLORES (top) and instruct-AYA (bottom); n=500 matched pairs per cell. Left: MAUVE on mean-pooled XLM-R-large embeddings (1.0 identical, 0.0 fully separable). Center: BERTScore F1, rescaled against per-language FLORES non-pair baselines (negative: paired generations less similar than random non-pairs). Right: PPL preservation under XGLM-7.5B, exp⁡(−|log⁡PPLR​(𝒳+)−log⁡PPLR​(𝒳−)|)∈(0,1]. Rows: (scheme, generator) cells; columns: eleven languages grouped by typological family.
Figure 1: Per-language quality preservation under three paradigms, base-FLORES (top) and instruct-AYA (bottom); n=500 matched pairs per cell. Left: MAUVE on mean-pooled XLM-R-large embeddings (1.0 identical, 0.0 fully separable). Center: BERTScore F1, rescaled against per-language FLORES non-pair baselines (negative: paired generations less similar than random non-pairs). Right: PPL preservation under XGLM-7.5B, exp⁡(−|log⁡PPLR​(𝒳+)−log⁡PPLR​(𝒳−)|)∈(0,1]. Rows: (scheme, generator) cells; columns: eleven languages grouped by typological family.
Table 1: The eleven evaluation languages with script, typological family, and Joshi resource tier.
ISO 639-3LanguageScriptTypological familyJoshi et al. tier
engEnglishLatinGermanic5
nldDutchLatinGermanic4
fraFrenchLatinRomance5
spaSpanishLatinRomance5
porPortugueseLatinRomance4
turTurkishLatinTurkic4
vieVietnameseLatinAustroasiatic4
hinHindiDevanagariIndic4
arbArabicArabicSemitic5
zhoChineseHanSinitic5
jpnJapaneseHan + KanaJaponic5
Figure 2: Per-scheme disparity decomposition under three quality paradigms, base-FLORES (top) and instruct-AYA (bottom). Bar height is GE2 over the per-language preservation vector, averaged across the three generators per scheme. Stacking shows the typological-partition decomposition into between-family (dark) and within-family (light) components; the percentage above is the between-family share. Absolute disparity is orders of magnitude smaller than under MAUVE or PPL but the between-family pattern persists. Y-axes differ across panels.
Figure 2: Per-scheme disparity decomposition under three quality paradigms, base-FLORES (top) and instruct-AYA (bottom). Bar height is GE2 over the per-language preservation vector, averaged across the three generators per scheme. Stacking shows the typological-partition decomposition into between-family (dark) and within-family (light) components; the percentage above is the between-family share. Absolute disparity is orders of magnitude smaller than under MAUVE or PPL but the between-family pattern persists. Y-axes differ across panels.
Table 2: Detection fairness by scheme, generator, and regime, subset all, α=0.01. Mean TPR shown with 95% bootstrap CIs (B=1000 resamples of the 11-language vector). Min TPR is the realized minimum. DI<.8 counts languages with TPRl<0.8⋅maxl⁡TPRl. GE2 uses the typological partition (Germanic, Romance, Indic, Semitic, Sinitic, Japonic, Turkic, Austroasiatic); Btw. % is its between-family share. Δ=mean​TPR​(τg)−mean​TPR​(τl). Values below 10−3 reported as <0.001; entries marked are undefined when the per-language mean is zero.
AUCTPR at τg (global)TPR at τl (per-lang)
SchemeGen.MeanMinMean TPRMin TPRDI<.8GE2Btw. %Mean TPRΔ
Base regime (FLORES+ continuation prompts)
KGWMistral0.9990.9970.992 [0.987, 0.996]0.9780<0.00183.60.991+0.001
Gemma0.9970.9910.969 [0.959, 0.979]0.9340<0.00172.40.974−0.005
Qwen0.9960.9870.965 [0.939, 0.986]0.8800<0.00185.10.966−0.001
UnigramMistral0.9960.9810.868 [0.802, 0.927]0.63830.00861.30.971−0.103
Gemma0.9860.9100.000 [0.000, 0.000]0.000110.799−0.799
Qwen0.9940.9720.898 [0.845, 0.942]0.72220.00493.80.949−0.051
SynthIDMistral1.0001.0001.000 [0.999, 1.000]0.9980<0.001100.01.000−0.000
Gemma1.0001.0001.000 [0.999, 1.000]0.9980<0.001100.01.000+0.000
Qwen1.0000.9990.997 [0.995, 0.999]0.9900<0.00197.50.998−0.000
EXPEditMistral1.0000.9980.999 [0.999, 1.000]0.9960<0.00188.01.000−0.000
Gemma1.0000.9981.000 [0.999, 1.000]0.9980<0.001100.01.000+0.000
Qwen0.9980.9930.991 [0.982, 0.997]0.9460<0.00199.20.992−0.001
DIPMistral0.9970.9890.975 [0.966, 0.984]0.9380<0.00192.60.975+0.000
Gemma0.9920.9650.963 [0.947, 0.977]0.9140<0.00199.50.965−0.002
Qwen0.9920.9770.931 [0.891, 0.968]0.79800.00394.70.929+0.002
XSIRMistral0.9810.9620.784 [0.743, 0.821]0.67220.00491.50.833−0.049
Gemma0.9840.9770.766 [0.712, 0.817]0.58840.00696.00.829−0.063
Qwen0.9780.9300.215 [0.118, 0.362]0.050100.49398.90.794−0.579
Instruct regime (AYA native-speaker prompts)
KGWMistral0.9530.9090.626 [0.540, 0.707]0.38670.02695.30.632−0.005
Gemma0.9000.8480.389 [0.325, 0.463]0.222100.05382.40.411−0.022
Qwen0.9660.9370.689 [0.628, 0.753]0.51080.01178.60.723−0.034
UnigramMistral0.9090.8650.311 [0.171, 0.453]0.02260.29799.80.450−0.139
Gemma0.8510.7900.037 [0.023, 0.052]0.010100.22681.20.151−0.113
Qwen0.9430.8800.546 [0.455, 0.621]0.21040.03291.70.636−0.090
SynthIDMistral0.9910.9820.905 [0.865, 0.940]0.78210.00377.70.903+0.002
Gemma0.9460.9020.602 [0.530, 0.675]0.422100.02079.70.597+0.005
Qwen0.9930.9770.941 [0.906, 0.971]0.79210.00291.90.942−0.002
Figure 3: Joint per-language detection-quality landscape, MAUVE paradigm. Markers: per-language (TPR​@​τg,MAUVE) pairs at α=0.01. Circles: base-FLORES. Triangles: instruct-AYA. Color: typological family. Annotations sB, sI: mean pairwise Euclidean distance within each regime. BERTScore and PPL-preservation companion panels in Appendix E.
Figure 3: Joint per-language detection-quality landscape, MAUVE paradigm. Markers: per-language (TPR​@​τg,MAUVE) pairs at α=0.01. Circles: base-FLORES. Triangles: instruct-AYA. Color: typological family. Annotations sB, sI: mean pairwise Euclidean distance within each regime. BERTScore and PPL-preservation companion panels in Appendix E.
Table 3: Cross-paradigm rank agreement: per-cell Spearman ρ between per-language preservation vectors, mean across (model, regime) cells with per-cell minimum in parentheses (subset=all; n=6 cells per scheme, n=36 pooled). Pairs use BERTScore F1 (rescaled) and PPL preservation under XGLM-7.5B; BS-raw gives nearly identical rankings (ρBS-raw↔BS-rescaled=+0.87). Negative minima identify cells where two paradigms produce opposite per-language orderings.
SchemeMAUVE ↔ BSMAUVE ↔ PPLBS ↔ PPL
KGW+0.61 (+0.45)+0.03 (−0.80)+0.03 (−0.76)
Unigram+0.56 (+0.23)+0.14 (−0.39)+0.20 (−0.67)
SynthID+0.76 (+0.47)+0.81 (+0.27)+0.77 (+0.45)
EXPEdit+0.68 (+0.26)+0.75 (+0.39)+0.70 (+0.46)
DIP+0.46 (−0.03)+0.43 (+0.05)+0.50 (−0.05)
XSIR+0.42 (−0.09)+0.10 (−0.71)+0.00 (−0.76)
All schemes+0.58 (−0.09)+0.38 (−0.80)+0.37 (−0.76)
Figure 4: Per-language detection diagnostics, subset all, α=0.01. Within each panel, left: TPR at the global empirical-FPR threshold τg; right: threshold-free AUC. Rows are (scheme, generator) cells; columns are the eleven evaluation languages, grouped by typological family.
Figure 4: Per-language detection diagnostics, subset all, α=0.01. Within each panel, left: TPR at the global empirical-FPR threshold τg; right: threshold-free AUC. Rows are (scheme, generator) cells; columns are the eleven evaluation languages, grouped by typological family.
Table 4: Per-cell language adherence on the eleven evaluation languages, base-FLORES (left) and instruct-AYA (right). NWM% is the GlotLID-v3 on-target rate for nowatermark generations; Δ¯ is the mean shift in percentage points across the six watermarked schemes. mean rows are unweighted means across the eleven languages.
Base regime (FLORES+ continuation prompts)Instruct regime (AYA native-speaker prompts)
MistralGemmaQwenMistralGemmaQwen
LangNWM%Δ¯NWM%Δ¯NWM%Δ¯NWM%Δ¯NWM%Δ¯NWM%Δ¯
eng99.6−0.398.8−2.098.4−0.692.6−0.980.6−1.793.2−2.0
fra99.2−1.197.6−4.089.2−0.291.4+0.586.2−1.794.8−4.9
spa100.0−1.496.8−2.691.2+1.195.2−1.362.6+2.792.0−5.0
por99.2−0.596.8−1.098.2−0.896.2−2.070.6+3.096.4−4.3
nld96.2−7.095.8+1.680.8−4.673.2−3.189.2−0.490.4−5.8
zho97.6−1.892.8−1.296.4−4.289.0−7.995.6−0.498.4−0.6
jpn96.8−3.395.6−1.699.8−3.371.0−10.794.0−1.998.4−7.4
arb89.2−3.087.6−3.887.8+2.967.6−15.684.2−1.281.0−15.0
hin96.2−9.188.2−2.398.8−3.952.2−6.583.8−1.483.2+0.1
tur89.0+1.396.8−4.357.2−1.181.8−8.192.8−0.695.6−13.5
vie98.6−4.777.0+4.799.0−1.079.2−7.398.8−0.695.2−6.4
mean96.5−2.893.1−1.590.6−1.480.9−5.785.3−0.492.6−5.9
(b) Instruct-AYA regime.
(b) Instruct-AYA regime.
Table 5: Length-distribution similarity between LID subsets and the all-subset baseline. Each row is a (regime, subset) pair; ratio=subset-mean length/all-subset mean length, computed per (regime, model, language, watermark) cell and summarized across cells. Cells require n≥100 in both subsets to enter the summary; n skipped counts cells that fail this floor (typically off_target on high-adherence cells).
RegimeSubsetn keptn skippedMeanMedianMinMax
Baseon_target_only23101.0031.0000.9961.081
Baseoff_target_only82230.9290.9080.8591.023
Instructon_target_only23101.0001.0000.9641.057
Instructoff_target_only262051.0251.0170.9471.119
Figure 5: PPL preservation under mGPT-1.3B cross-reference. Format and color scale match Figure 1.
Figure 5: PPL preservation under mGPT-1.3B cross-reference. Format and color scale match Figure 1.
Table 6: LID categorization robustness: per-cell adherence under a 3×3 sweep of the GlotLID per-sentence confidence floor (conf_floor) and the per-text on-target sentence fraction (on_target_ratio). ρ is the Spearman rank-correlation between the per-cell on-target-rate vector under each configuration and the default (†). Mean on-target % is the unweighted mean across the 462 cells of the evaluation grid. Mean/Max |Δ| are the mean and maximum absolute per-cell shift in on-target rate (pp) relative to default.
conf_flooron_target_ratioρ vs defaultMean on-target %Mean |Δ| (pp)Max |Δ| (pp)
0.30.70.934792.02.1913.60
0.30.80.946890.70.9113.00
0.30.90.862786.34.8731.00
0.50.70.991091.11.2611.80
0.5†0.81.000089.80.000.00
0.50.90.931785.44.4331.20
0.70.70.948687.93.7624.60
0.70.80.962786.83.0424.80
0.70.90.938582.57.3631.80
(b) Instruct-AYA regime.
(b) Instruct-AYA regime.
Table 7: Per-language calibration gap Δl=TPRl​(τl)−TPRl​(τg), base-FLORES (top) and instruct-AYA (bottom), subset all, α=0.01. Positive: per-language calibration recovers detection. Negative: per-language calibration worsens detection. Zero: detection insensitive to calibration locality. Row means equal −1× the Δ column of Table 2.
GermanicRomanceIndicSemiticSiniticJaponicTurkicAus.
SchemeGen.engnldfraspaporhinarbzhojpnturvie
Base regime (FLORES+ continuation prompts)
KGWMistral0.000+0.004+0.004+0.0020.000−0.0080.000−0.0060.000−0.004+0.002
Gemma+0.006+0.030+0.008+0.016+0.038−0.020+0.044−0.010−0.0640.000+0.010
Qwen+0.028+0.018−0.002−0.002+0.002−0.0080.000+0.0080.000−0.0300.000
UnigramMistral+0.074+0.354+0.218+0.292+0.142−0.062+0.004+0.064−0.018+0.074−0.012
Gemma+0.996+0.992+0.968+0.990+0.9600.000+0.958+0.976+0.984+0.9600.000
Qwen+0.024+0.034−0.040+0.002+0.022+0.106+0.230+0.134−0.008+0.092−0.034
SynthIDMistral0.0000.0000.0000.0000.0000.0000.000+0.0020.0000.0000.000
Gemma0.0000.0000.0000.0000.0000.0000.0000.0000.0000.0000.000
Qwen0.0000.0000.0000.0000.000+0.0020.000+0.0060.000−0.0040.000
EXPEditMistral0.000+0.0020.0000.0000.0000.0000.0000.0000.0000.0000.000
Gemma0.0000.0000.0000.0000.0000.0000.0000.0000.0000.0000.000
Qwen+0.002−0.0060.0000.0000.000+0.0120.0000.0000.000+0.0020.000
DIPMistral0.0000.000+0.002+0.002+0.002+0.0020.000−0.0140.000+0.0040.000
Gemma0.000−0.004+0.0020.000+0.002+0.010+0.004+0.004+0.018−0.0100.000
Qwen0.000−0.008+0.004+0.008−0.004−0.004+0.002+0.0100.000−0.028+0.002
XSIRMistral+0.024−0.054+0.166+0.106+0.086−0.022−0.066+0.092+0.014+0.014+0.178
Gemma+0.028+0.054+0.054+0.114+0.026+0.146+0.056−0.042+0.168+0.122−0.030
Qwen+0.516+0.774+0.770+0.802+0.912−0.286+0.684+0.638+0.464+0.360+0.732
Instruct regime (AYA native-speaker prompts)
KGWMistral+0.112+0.032−0.250+0.028+0.118−0.016+0.014+0.030−0.018+0.012−0.004
Gemma+0.030−0.070+0.108+0.046+0.112−0.0380.000−0.036−0.082+0.134+0.034
Qwen+0.012+0.014−0.036−0.058+0.110−0.106+0.138+0.016+0.084+0.058+0.140
UnigramMistral+0.340+0.380+0.090+0.282+0.316−0.1060.000+0.070−0.178+0.282+0.052
Gemma+0.024−0.054+0.174+0.178+0.056−0.070−0.036+0.376+0.276+0.272+0.050
Qwen+0.084−0.018−0.108−0.012+0.078+0.128+0.438+0.352+0.094+0.088−0.132
SynthIDMistral+0.056+0.020−0.026−0.034−0.048+0.0160.0000.000−0.0220.000+0.016
Gemma+0.050−0.030+0.046−0.022−0.020−0.022−0.124+0.038−0.102+0.094+0.036
Qwen+0.018−0.002+0.010−0.022−0.002+0.0160.0000.0000.0000.0000.000
Figure 6: On-target minus all-subset preservation, base-FLORES regime (top) and instruct-AYA regime (bottom). Each cell is Δ=preservationon-target−preservationall for the indicated paradigm; rows are (scheme, generator), columns are languages. Positive (red): on-target restriction raises preservation; negative (blue): on-target restriction lowers it. Color scale fixed at [−0.3,+0.3]; values outside this range saturate. The EXPEdit-Mistral, Arabic, BERTScore F1 entry in the instruct panel is omitted (—) because the on-target sample size falls below the n≥100 paired-set floor.
Figure 6: On-target minus all-subset preservation, base-FLORES regime (top) and instruct-AYA regime (bottom). Each cell is Δ=preservationon-target−preservationall for the indicated paradigm; rows are (scheme, generator), columns are languages. Positive (red): on-target restriction raises preservation; negative (blue): on-target restriction lowers it. Color scale fixed at [−0.3,+0.3]; values outside this range saturate. The EXPEdit-Mistral, Arabic, BERTScore F1 entry in the instruct panel is omitted (—) because the on-target sample size falls below the n≥100 paired-set floor.
Table 8: Detection fairness summary at α=0.05, base-FLORES (top) and instruct-AYA (bottom), subset all. Format mirrors Table 2; AUC columns are omitted (threshold-independent, identical to Table 2) and bootstrap CIs on Mean TPR are omitted for compactness. L=11 throughout. Values below 10−3 reported as <0.001.
TPR at τg (global)TPR at τl (per-lang)
SchemeGen.Mean TPRMin TPRDI<.8GE2Btw. %Mean TPRΔ
Base regime (FLORES+ continuation prompts)
KGWMistral0.9970.9920<0.00161.10.997+0.000
Gemma0.9910.9800<0.00193.20.993−0.002
Qwen0.9870.9520<0.00157.80.986+0.000
UnigramMistral0.9770.9200<0.00154.90.988−0.011
Gemma0.9880.9760<0.00155.90.898+0.089
Qwen0.9610.86000.00186.30.979−0.018
SynthIDMistral1.0000.9980<0.001100.01.000+0.000
Gemma1.0000.9980<0.001100.01.000+0.000
Qwen0.9990.9960<0.00174.30.999+0.000
EXPEditMistral1.0000.9960<0.001100.01.000+0.000
Gemma1.0000.9980<0.001100.01.000+0.000
Qwen0.9950.9660<0.00199.20.995−0.000
DIPMistral0.9900.9720<0.00194.70.989+0.001
Gemma0.9810.9420<0.00197.90.981−0.001
Qwen0.9660.89600.00191.60.966+0.000
XSIRMistral0.9240.86800.00183.30.935−0.011
Gemma0.9290.84000.00192.80.944−0.015
Qwen0.5030.27690.07694.60.919−0.416
Instruct regime (AYA native-speaker prompts)
KGWMistral0.8010.62220.00685.00.813−0.012
Gemma0.6160.474100.01579.70.617−0.001
Qwen0.8600.74010.00378.10.869−0.009
UnigramMistral0.5170.09050.14995.80.659−0.143
Gemma0.3800.15450.05080.30.408−0.028
Qwen0.7270.36640.01893.80.790−0.063
SynthIDMistral0.9670.9260<0.00177.80.969−0.002
Gemma0.7910.66240.00575.90.785+0.005
Qwen0.9710.8980<0.00195.70.973−0.002
Figure 7: Joint per-language detection–quality landscape, BERTScore-rescaled (top) and PPL-XGLM (bottom) paradigms. Format matches Figure 3: markers are per-language (TPR​@​τg,quality) pairs at α=0.01, n=33 per regime (11 languages × 3 generators); circles are base-FLORES, triangles are instruct-AYA; color encodes typological family; sB, sI are mean pairwise Euclidean distance within each regime.
Figure 7: Joint per-language detection–quality landscape, BERTScore-rescaled (top) and PPL-XGLM (bottom) paradigms. Format matches Figure 3: markers are per-language (TPR​@​τg,quality) pairs at α=0.01, n=33 per regime (11 languages × 3 generators); circles are base-FLORES, triangles are instruct-AYA; color encodes typological family; sB, sI are mean pairwise Euclidean distance within each regime.
Table 9: Detection summary on the on-target subset, α=0.01, global threshold τg, base-FLORES (left) and instruct-AYA (right). Thresholds re-calibrated on each cell’s on-target unwatermarked pool; L=11 throughout. The off-target subset is omitted: in the base regime no cell meets the nneg≥200 inclusion floor, and in the instruct regime only six (Mistral-generator) cells pass the relaxed α=0.05 floor with L≤3, insufficient for the fairness aggregates this appendix reports. The Δ vs. all columns report change in Mean TPR relative to subset all (Table 2).
Base regime (FLORES+ continuation)Instruct regime (AYA native-speaker)
SchemeGen.Mean TPRMin TPRDI<.8Δ vs. allMean TPRMin TPRDI<.8Δ vs. all
KGWMistral0.9950.9830+0.0030.6810.4256+0.055
Gemma0.9730.9390+0.0040.4260.26410+0.037
Qwen0.9870.9630+0.0220.6970.5208+0.008
UnigramMistral0.8850.6542+0.0170.3880.0117+0.077
Gemma0.9870.9740+0.9870.0200.0009−0.017
Qwen0.9250.7621+0.0270.5470.1764+0.001
SynthIDMistral1.0000.99800.0000.9240.7741+0.019
Gemma1.0000.99800.0000.6100.42610+0.008
Qwen0.9990.9960+0.0020.9410.76410.000
EXPEditMistral1.0001.0000+0.0010.6750.2345+0.098
Gemma1.0000.99700.0000.2320.09110+0.007
Qwen0.9940.9410+0.0030.5080.05910+0.012
DIPMistral0.9800.9520+0.0050.4880.2669+0.033
Gemma0.9770.9540+0.0140.1980.10510+0.002
Qwen0.9570.8480+0.0260.5070.2008−0.006
XSIRMistral0.7990.6882+0.0150.2960.08110+0.030
Gemma0.7930.6003+0.0270.0920.01910+0.011
Qwen0.2070.04710−0.0080.0470.00010+0.004
Table 10: Bottom-quintile-mean Rawlsian floor as a stability companion to Min TPR (Table 2), base-FLORES (left) and instruct-AYA (right), subset all, α=0.01, global threshold τg. With L=11 evaluation languages the bottom quintile is ⌊0.2⋅11⌋=2 languages; the columns report the mean TPR of the two worst-performing languages per cell, with the languages listed minimum first (arbitrary tie-breaking). The Unigram-Gemma base cell is degenerate (all eleven per-language TPRs are zero); its bottom-quintile language identities are reported but uninformative.
Base regime (FLORES+)Instruct regime (AYA)
SchemeGen.Min TPRBot.-q meanBot.-q langsMin TPRBot.-q meanBot.-q langs
KGWMistral0.9780.979zho, hin0.3860.447fra, por
Gemma0.9340.943arb, por0.2220.251por, vie
Qwen0.8800.894tur, nld0.5100.564por, nld
UnigramMistral0.6380.672nld, spa0.0220.029spa, fra
Gemma0.0000.000eng, fra0.0100.010por, tur
Qwen0.7220.744tur, arb0.2100.306arb, hin
SynthIDMistral0.9980.998zho, tur0.7820.788fra, spa
Gemma0.9980.999vie, eng0.4220.437vie, por
Qwen0.9900.992tur, zho0.7920.842hin, por
EXPEditMistral0.9960.997zho, nld0.2380.244eng, spa
Gemma0.9980.998jpn, vie0.0840.101jpn, por
Qwen0.9460.968hin, nld0.1740.226hin, eng
DIPMistral0.9380.950zho, jpn0.2640.265fra, spa
Gemma0.9140.917hin, vie0.1100.115por, jpn
Qwen0.7980.819tur, hin0.2960.345hin, fra
XSIRMistral0.6720.673zho, vie0.0880.116fra, spa
Gemma0.5880.635jpn, hin0.0160.019tur, jpn
Qwen0.0500.064por, fra0.0000.000nld, vie
Table 11: Robustness of the generalized-entropy decomposition to choice of inequality measure: GE2 (half the squared coefficient of variation, more sensitive to inequality at the top of the per-language TPR distribution) versus GE0 (mean log deviation, more sensitive to inequality at the bottom). Base-FLORES (left) and instruct-AYA (right), subset all, α=0.01, global threshold τg; typological partition as in Table 2. Total values below 10−3 reported as <0.001. The Unigram-Gemma base GE0 entry is obtained by replacing zero TPRs with a small ϵ (per the strict-positivity requirement of GE0); the resulting Total=0 and Btw.%=100.0 are regularization artifacts, marked †.
Base regime (FLORES+)Instruct regime (AYA)
GE2GE0GE2GE0
SchemeGen.TotalBtw. %TotalBtw. %TotalBtw. %TotalBtw. %
KGWMistral<0.00183.6<0.00183.70.02695.30.02791.4
Gemma<0.00172.4<0.00172.70.05382.40.04673.3
Qwen0.00185.10.00185.10.01178.60.01172.5
UnigramMistral0.00861.30.00955.70.29799.80.55694.8
Gemma<0.001†100.0†0.22681.20.24373.9
Qwen0.00493.80.00594.60.03291.70.04694.6
SynthIDMistral<0.001100.0<0.001100.00.00377.70.00375.6
Gemma<0.001100.0<0.001100.00.02079.70.01975.6
Qwen<0.00197.5<0.00197.50.00291.90.00292.2
EXPEditMistral<0.00188.0<0.00188.10.06780.00.08768.2
Gemma<0.001100.0<0.001100.00.07184.20.07779.0
Qwen<0.00199.2<0.00199.20.08084.40.09381.4
DIPMistral<0.00192.6<0.00192.80.05899.50.05999.2
Gemma<0.00199.5<0.00199.50.08491.60.06884.4
Qwen0.00394.70.00394.60.03599.60.03599.6
XSIRMistral0.00491.50.00492.10.18591.80.16083.4
Gemma0.00696.00.00796.30.17576.50.22081.1
Qwen0.49398.90.32692.53.78799.93.24479.0
Table 12: Robustness of the GE2 between-share to choice of partition: typological family (8 groups: Germanic, Romance, Indic, Semitic, Sinitic, Japonic, Turkic, Austroasiatic, with 5 of 11 languages in non-singleton groups) vs. script (4 groups: Latin pooling eng, nld, fra, spa, por, tur, vie; Devanagari=hin; Arabic=arb; Han pooling zho, jpn; with 9 of 11 languages in non-singleton groups). Base-FLORES (left) and instruct-AYA (right), subset all, α=0.01, global threshold τg. GE2 Total depends only on the per-language vector and is partition-independent (shown once per regime). Totals below 10−3 reported as <0.001; for these cells the between-share values involve ratios of near-zero variance components and should not be over-interpreted.
Base regime (FLORES+)Instruct regime (AYA)
Btw. %Btw. %
SchemeGen.GE2FamilyScriptGE2FamilyScript
KGWMistral<0.00183.646.20.02695.332.8
Gemma<0.00172.462.20.05382.429.8
Qwen0.00185.126.20.01178.626.7
UnigramMistral0.00861.328.80.29799.846.3
Gemma0.22681.261.7
Qwen0.00493.826.10.03291.775.0
SynthIDMistral<0.001100.017.10.00377.733.0
Gemma<0.001100.05.70.02079.747.1
Qwen<0.00197.57.50.00291.979.3
EXPEditMistral<0.00188.031.70.06780.030.2
Gemma<0.001100.017.10.07184.28.5
Qwen<0.00199.295.80.08084.441.6
DIPMistral<0.00192.669.30.05899.538.9
Gemma<0.00199.548.20.08491.625.2
Qwen0.00394.732.50.03599.653.4
XSIRMistral0.00491.554.80.18591.864.7
Gemma0.00696.012.70.17576.57.6
Qwen0.49398.993.53.78799.999.8
Table 13: Quality fairness summary by scheme, generator, regime, and measurement, subset all. Each row reports mean per-language preservation with the per-cell minimum and floor language (ISO 639-3 superscript) in parentheses. DI<.8 counts languages with per-language preservation below 0.8×maxl. GE2 uses the typological partition (Germanic, Romance, Indic, Semitic, Sinitic, Japonic, Turkic, Austroasiatic); Btw. % is its between-family share. BERTScore mean (min) is reported on rescaled F1 [28]; DI<.8 and GE2 are omitted for BERTScore because neither registers signal on the bounded raw scale and both behave pathologically on the rescaled scale (§5.2).
MAUVEBERTScore F1 (rescaled)PPL preservation (XGLM-7.5B)
SchemeGenmean (min)DI<.8GE2Btw. %mean (min)mean (min)DI<.8GE2Btw. %
Base regime (FLORES+ continuation)
KGWMistral0.947 (0.879)jpn00.00186.0-0.041 (-0.079)vie0.866 (0.793)spa00.00266.4
KGWGemma0.855 (0.435)arb10.01398.9-0.081 (-0.194)vie0.899 (0.812)fra00.00291.0
KGWQwen0.889 (0.471)hin10.01295.6-0.034 (-0.091)tur0.834 (0.686)hin30.00594.0
UnigramMistral0.747 (0.420)arb50.02649.5-0.070 (-0.134)arb0.872 (0.811)spa00.00296.4
UnigramGemma0.638 (0.444)spa60.01767.8-0.098 (-0.191)jpn0.911 (0.815)nld00.00268.3
UnigramQwen0.734 (0.123)hin30.04397.2-0.055 (-0.128)zho0.896 (0.647)hin10.00598.0
SynthIDMistral0.187 (0.020)arb100.43559.0-0.115 (-0.163)vie0.299 (0.145)arb70.04888.4
SynthIDGemma0.095 (0.029)tur100.64143.0-0.146 (-0.254)vie0.269 (0.195)tur70.02173.2
SynthIDQwen0.183 (0.035)vie100.45553.0-0.112 (-0.188)tur0.308 (0.203)arb90.02878.2
EXPEditMistral0.079 (0.006)arb100.53657.1-0.149 (-0.228)vie0.218 (0.079)arb60.08599.1
EXPEditGemma0.013 (0.006)jpn100.57443.4-0.220 (-0.333)vie0.135 (0.076)zho90.03382.7
EXPEditQwen0.067 (0.013)arb100.87447.9-0.145 (-0.245)tur0.224 (0.135)tur90.04274.6
DIPMistral0.973 (0.933)hin00.00095.6-0.015 (-0.040)vie0.809 (0.750)hin00.00177.6
DIPGemma0.932 (0.878)jpn00.00194.7-0.053 (-0.183)vie0.774 (0.702)nld00.00168.0
DIPQwen0.947 (0.864)hin00.00181.6-0.008 (-0.071)tur0.811 (0.714)tur00.00292.7
XSIRMistral0.859 (0.642)hin20.00686.6-0.062 (-0.110)hin0.922 (0.816)nld00.00273.2
XSIRGemma0.683 (0.461)fra60.01880.9-0.123 (-0.226)vie0.925 (0.785)nld10.00275.8
XSIRQwen0.850 (0.603)hin20.00889.9-0.042 (-0.129)tur0.883 (0.745)hin10.00386.4
Instruct regime (AYA native-speaker)
KGWMistral0.989 (0.970)arb00.00058.90.255 (0.088)arb0.892 (0.853)jpn00.00191.0
KGWGemma0.990 (0.959)arb00.00065.30.397 (0.338)tur0.967 (0.941)hin00.00080.0
KGWQwen0.987 (0.954)hin00.00090.80.300 (0.192)tur0.885 (0.749)hin10.00292.5
UnigramMistral0.981 (0.938)nld00.00040.30.240 (0.087)arb0.897 (0.835)hin00.00197.4
UnigramGemma0.996 (0.991)jpn00.00054.10.400 (0.329)tur0.965 (0.941)nld00.00078.0
UnigramQwen0.980 (0.927)hin00.00097.70.297 (0.184)tur0.885 (0.700)hin10.00395.9
SynthIDMistral0.721 (0.066)arb40.10799.60.149 (-0.117)arb0.595 (0.185)arb70.07095.4
SynthIDGemma0.994 (0.979)fra00.00050.00.365 (0.296)tur0.931 (0.885)zho00.00090.0
SynthIDQwen0.833 (0.237)tur20.03397.10.219 (0.047)tur0.645 (0.442)tur50.01490.7
Table 14: Single-covariate R2 for per-language detection TPR at τg, α=0.01, subset all. Each fit regresses the per-language outcome vector on a single covariate over 33 (language, generator) points. Stars from the slope p-value: ∗p<0.05, p∗⁣∗<0.01.
CovariateKGWUnigramSynthIDEXPEditDIPXSIR
Base regime (FLORES+ continuation, n=33)
Fertility (tokens/sent.)0.0160.0520.0080.834∗∗0.238∗∗0.001
NWM surprisal (XGLM)0.0660.0740.0580.261∗∗0.156∗0.003
Latin script0.0290.0010.0050.0510.0090.019
NWM adherence0.600∗∗0.0060.415∗∗0.0010.461∗∗0.091
Instruct regime (AYA native-speaker, n=33)
Fertility (tokens/sent.)0.0170.0460.0090.0010.0000.107
NWM surprisal (XGLM)0.0240.0030.0090.0400.0170.011
Latin script0.0460.0140.0350.0490.0400.058
NWM adherence0.0000.0000.0030.0470.0000.291∗∗
Table 15: Single-covariate R2 for per-language MAUVE preservation. Each fit regresses the per-language outcome vector on a single covariate over 33 (language, generator) points. Stars from the slope p-value: ∗p<0.05, p∗⁣∗<0.01. BERTScore-rescaled and PPL-XGLM paradigm-stability tables in §F.1.
CovariateKGWUnigramSynthIDEXPEditDIPXSIR
Base regime (FLORES+ continuation, n=33)
Fertility (tokens/sent.)0.392∗∗0.249∗∗0.0270.0050.234∗∗0.093
NWM surprisal (XGLM)0.0260.0230.0180.0010.0070.009
Latin script0.136∗0.0360.0620.0520.0830.000
NWM adherence0.0310.0070.1010.0740.219∗∗0.072
Instruct regime (AYA native-speaker, n=33)
Fertility (tokens/sent.)0.346∗∗0.464∗∗0.0270.0160.0190.168∗
NWM surprisal (XGLM)0.0080.152∗0.0220.0520.0050.002
Latin script0.122∗0.0510.0180.0410.0520.082
NWM adherence0.0760.0990.153∗0.230∗∗0.0530.047
Table 16: Single-covariate R2 for BERTScore-rescaled preservation. Stars from the slope p-value: ∗p<0.05, p∗⁣∗<0.01.
CovariateKGWUnigramSynthIDEXPEditDIPXSIR
Base regime (n=33)
Fertility (tokens/sent.)0.0000.0130.0240.0540.0390.095
NWM surprisal (XGLM)0.0360.0160.0000.0320.0390.094
Latin script0.0680.207∗∗0.0040.0220.0030.008
NWM adherence0.239∗∗0.0040.378∗∗0.259∗∗0.298∗∗0.188∗
Instruct regime (n=33)
Fertility (tokens/sent.)0.0720.0690.0430.0260.0470.051
NWM surprisal (XGLM)0.0340.0340.0300.0210.0370.026
Latin script0.0460.0450.0420.0370.0440.056
NWM adherence0.0960.0950.119∗0.125∗0.123∗0.110
Table 17: Single-covariate R2 for PPL-XGLM preservation. Stars from the slope p-value: ∗p<0.05, p∗⁣∗<0.01.
CovariateKGWUnigramSynthIDEXPEditDIPXSIR
Base regime (n=33)
Fertility (tokens/sent.)0.191∗0.327∗∗0.0130.0000.121∗0.148∗
NWM surprisal (XGLM)0.0030.179∗0.0170.0180.0940.183∗
Latin script0.0220.0010.126∗0.1050.0520.056
NWM adherence0.186∗0.0140.164∗0.123∗0.1090.015
Instruct regime (n=33)
Fertility (tokens/sent.)0.446∗∗0.559∗∗0.119∗0.0700.225∗∗0.340∗∗
NWM surprisal (XGLM)0.0440.0290.0100.0000.0470.016
Latin script0.0860.0770.0700.0840.0460.012
NWM adherence0.0200.0240.0960.217∗∗0.0230.006

为什么重要

随着多语言AI生成的虚假信息不断扩散,水印技术被寄予厚望作为应对手段,但如果它在某些语系上悄悄失效,使用这些语言的用户获得的保护更弱,同时承担的文本质量损失反而更大。这项研究打破了英语测试结果可以直接套用到其他语言的假设,并提供了一套具体方法,能在实际部署前发现这类语言差距。

本文术语

  • 水印(watermarking) · 在AI生成文本中嵌入肉眼不可见的统计痕迹,以便事后识别其为机器生成
  • 检测阈值 · 水印得分超过该数值即被判定为含有水印的临界线
  • AUC · 不依赖具体阈值、衡量两组数据可分离程度的指标
  • 语系(typological family) · 语法结构或词汇特征相近的一组语言,例如日耳曼语系、闪米特语系
  • distortion-free方案 · 理论上不改变生成文本概率分布的水印设计方法

论文原文摘要(英文)

Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme's detection threshold and a narrow set of quality measurements. Multilingual deployment exposes evaluation-design choices that are inconsequential on English but determine conclusions cross-lingually. We propose an evaluation framework with four components: detection thresholds calibrated empirically per deployment context, a threshold-independent companion measurement that distinguishes calibration failures from detection failures, three disjoint quality measurement paradigms (distributional, paired-semantic, and reference-perplexity), and a generalized-entropy decomposition of cross-language disparity over a typological family partition. Applied to six watermarking schemes, three open-weight generators, eleven languages spanning four scripts and eight typological families, and both base and instruction-tuned regimes, the framework reveals failure modes that single-language single-paradigm evaluation cannot surface. Across detection and quality, observed disparity is predominantly between-family on the typological partition, indicating that cross-lingual fairness gaps in watermarking are structural to language properties rather than idiosyncratic to particular languages.

作者 · Alexander Nemecek, Osama Zafar, Debargha Ganguly, Vikash Singh, Vipin Chaudhary, Erman Ayday

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道

图片来源: Alexander Nemecek et al., arXiv:2608.20047, CC BY 4.0