매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

Auditing Cross-Lingual Fairness in Language Model Watermarking

arXiv:2608.200472026-08-21

AI 생성 텍스트를 잡아내는 워터마크 기술이 영어 아닌 언어에서는 훨씬 부실하게 작동하고, 그 격차는 개별 언어가 아니라 언어 계열 단위로 나타난다

AI가 쓴 글임을 나중에 알아볼 수 있도록 눈에 안 보이는 통계적 표시를 심는 '워터마킹' 기술이 확산되고 있지만, 지금까지 성능 검증은 거의 영어로만 이뤄졌다. 이 연구는 11개 언어, 6가지 워터마킹 기법, 3개 공개 AI 모델을 대상으로 새로운 평가 방법을 만들어 적용했고, 탐지 성능과 글의 품질 보존 모두 언어에 따라 큰 편차를 보인다는 것을 확인했다. 이 편차는 특정 언어 하나의 문제가 아니라, 비슷한 언어들이 묶인 '언어 계열' 수준에서 구조적으로 나타났다.

무엇을 했나

  1. 11개 언어(4개 문자, 8개 언어 계열), 6가지 워터마킹 기법, 3개 공개 AI 모델을 기본 문장완성 모드와 지시문 응답(챗봇) 모드 두 가지로 나눠 총 그리드 형태로 검증했다.
  2. 각 기법이 원래 정해둔 영어 기준 판정선을 그대로 쓰지 않고, 언어와 모델별로 탐지 기준선을 다시 계산했으며, 워터마크 신호가 아예 없는 것인지 아니면 기준선만 잘못 맞춰진 것인지를 구분해서 진단했다.
  3. 텍스트 품질 훼손 여부를 한 가지 방법이 아니라 전체 분포 비교, 문장 단위 의미 유사도, 참조 모델 기반 자연스러움까지 서로 다른 세 가지 방법으로 교차 검증했다.
  4. 이론적으로 텍스트를 왜곡하지 않는다고 알려진 SynthID-Text, EXPEdit 두 기법이 실제로는 품질을 가장 많이 훼손했고, 지시문 응답 모드에서는 대부분의 기법에서 탐지 성능이 크게 떨어졌다.
  5. 통계 분석 결과 언어 간 불균형의 대부분은 개별 언어 하나의 특이성이 아니라 게르만어파, 셈어파, 중국어파 같은 언어 계열 간 차이에서 비롯됐으며, 이는 문제가 언어 구조 자체에 내재해 있음을 시사한다.
Figure 1: Per-language quality preservation under three paradigms, base-FLORES (top) and instruct-AYA (bottom); n=500 matched pairs per cell. Left: MAUVE on mean-pooled XLM-R-large embeddings (1.0 identical, 0.0 fully separable). Center: BERTScore F1, rescaled against per-language FLORES non-pair baselines (negative: paired generations less similar than random non-pairs). Right: PPL preservation under XGLM-7.5B, exp⁡(−|log⁡PPLR​(𝒳+)−log⁡PPLR​(𝒳−)|)∈(0,1]. Rows: (scheme, generator) cells; columns: eleven languages grouped by typological family.
Figure 1: Per-language quality preservation under three paradigms, base-FLORES (top) and instruct-AYA (bottom); n=500 matched pairs per cell. Left: MAUVE on mean-pooled XLM-R-large embeddings (1.0 identical, 0.0 fully separable). Center: BERTScore F1, rescaled against per-language FLORES non-pair baselines (negative: paired generations less similar than random non-pairs). Right: PPL preservation under XGLM-7.5B, exp⁡(−|log⁡PPLR​(𝒳+)−log⁡PPLR​(𝒳−)|)∈(0,1]. Rows: (scheme, generator) cells; columns: eleven languages grouped by typological family.
Table 1: The eleven evaluation languages with script, typological family, and Joshi resource tier.
ISO 639-3LanguageScriptTypological familyJoshi et al. tier
engEnglishLatinGermanic5
nldDutchLatinGermanic4
fraFrenchLatinRomance5
spaSpanishLatinRomance5
porPortugueseLatinRomance4
turTurkishLatinTurkic4
vieVietnameseLatinAustroasiatic4
hinHindiDevanagariIndic4
arbArabicArabicSemitic5
zhoChineseHanSinitic5
jpnJapaneseHan + KanaJaponic5
Figure 2: Per-scheme disparity decomposition under three quality paradigms, base-FLORES (top) and instruct-AYA (bottom). Bar height is GE2 over the per-language preservation vector, averaged across the three generators per scheme. Stacking shows the typological-partition decomposition into between-family (dark) and within-family (light) components; the percentage above is the between-family share. Absolute disparity is orders of magnitude smaller than under MAUVE or PPL but the between-family pattern persists. Y-axes differ across panels.
Figure 2: Per-scheme disparity decomposition under three quality paradigms, base-FLORES (top) and instruct-AYA (bottom). Bar height is GE2 over the per-language preservation vector, averaged across the three generators per scheme. Stacking shows the typological-partition decomposition into between-family (dark) and within-family (light) components; the percentage above is the between-family share. Absolute disparity is orders of magnitude smaller than under MAUVE or PPL but the between-family pattern persists. Y-axes differ across panels.
Table 2: Detection fairness by scheme, generator, and regime, subset all, α=0.01. Mean TPR shown with 95% bootstrap CIs (B=1000 resamples of the 11-language vector). Min TPR is the realized minimum. DI<.8 counts languages with TPRl<0.8⋅maxl⁡TPRl. GE2 uses the typological partition (Germanic, Romance, Indic, Semitic, Sinitic, Japonic, Turkic, Austroasiatic); Btw. % is its between-family share. Δ=mean​TPR​(τg)−mean​TPR​(τl). Values below 10−3 reported as <0.001; entries marked are undefined when the per-language mean is zero.
AUCTPR at τg (global)TPR at τl (per-lang)
SchemeGen.MeanMinMean TPRMin TPRDI<.8GE2Btw. %Mean TPRΔ
Base regime (FLORES+ continuation prompts)
KGWMistral0.9990.9970.992 [0.987, 0.996]0.9780<0.00183.60.991+0.001
Gemma0.9970.9910.969 [0.959, 0.979]0.9340<0.00172.40.974−0.005
Qwen0.9960.9870.965 [0.939, 0.986]0.8800<0.00185.10.966−0.001
UnigramMistral0.9960.9810.868 [0.802, 0.927]0.63830.00861.30.971−0.103
Gemma0.9860.9100.000 [0.000, 0.000]0.000110.799−0.799
Qwen0.9940.9720.898 [0.845, 0.942]0.72220.00493.80.949−0.051
SynthIDMistral1.0001.0001.000 [0.999, 1.000]0.9980<0.001100.01.000−0.000
Gemma1.0001.0001.000 [0.999, 1.000]0.9980<0.001100.01.000+0.000
Qwen1.0000.9990.997 [0.995, 0.999]0.9900<0.00197.50.998−0.000
EXPEditMistral1.0000.9980.999 [0.999, 1.000]0.9960<0.00188.01.000−0.000
Gemma1.0000.9981.000 [0.999, 1.000]0.9980<0.001100.01.000+0.000
Qwen0.9980.9930.991 [0.982, 0.997]0.9460<0.00199.20.992−0.001
DIPMistral0.9970.9890.975 [0.966, 0.984]0.9380<0.00192.60.975+0.000
Gemma0.9920.9650.963 [0.947, 0.977]0.9140<0.00199.50.965−0.002
Qwen0.9920.9770.931 [0.891, 0.968]0.79800.00394.70.929+0.002
XSIRMistral0.9810.9620.784 [0.743, 0.821]0.67220.00491.50.833−0.049
Gemma0.9840.9770.766 [0.712, 0.817]0.58840.00696.00.829−0.063
Qwen0.9780.9300.215 [0.118, 0.362]0.050100.49398.90.794−0.579
Instruct regime (AYA native-speaker prompts)
KGWMistral0.9530.9090.626 [0.540, 0.707]0.38670.02695.30.632−0.005
Gemma0.9000.8480.389 [0.325, 0.463]0.222100.05382.40.411−0.022
Qwen0.9660.9370.689 [0.628, 0.753]0.51080.01178.60.723−0.034
UnigramMistral0.9090.8650.311 [0.171, 0.453]0.02260.29799.80.450−0.139
Gemma0.8510.7900.037 [0.023, 0.052]0.010100.22681.20.151−0.113
Qwen0.9430.8800.546 [0.455, 0.621]0.21040.03291.70.636−0.090
SynthIDMistral0.9910.9820.905 [0.865, 0.940]0.78210.00377.70.903+0.002
Gemma0.9460.9020.602 [0.530, 0.675]0.422100.02079.70.597+0.005
Qwen0.9930.9770.941 [0.906, 0.971]0.79210.00291.90.942−0.002
Figure 3: Joint per-language detection-quality landscape, MAUVE paradigm. Markers: per-language (TPR​@​τg,MAUVE) pairs at α=0.01. Circles: base-FLORES. Triangles: instruct-AYA. Color: typological family. Annotations sB, sI: mean pairwise Euclidean distance within each regime. BERTScore and PPL-preservation companion panels in Appendix E.
Figure 3: Joint per-language detection-quality landscape, MAUVE paradigm. Markers: per-language (TPR​@​τg,MAUVE) pairs at α=0.01. Circles: base-FLORES. Triangles: instruct-AYA. Color: typological family. Annotations sB, sI: mean pairwise Euclidean distance within each regime. BERTScore and PPL-preservation companion panels in Appendix E.
Table 3: Cross-paradigm rank agreement: per-cell Spearman ρ between per-language preservation vectors, mean across (model, regime) cells with per-cell minimum in parentheses (subset=all; n=6 cells per scheme, n=36 pooled). Pairs use BERTScore F1 (rescaled) and PPL preservation under XGLM-7.5B; BS-raw gives nearly identical rankings (ρBS-raw↔BS-rescaled=+0.87). Negative minima identify cells where two paradigms produce opposite per-language orderings.
SchemeMAUVE ↔ BSMAUVE ↔ PPLBS ↔ PPL
KGW+0.61 (+0.45)+0.03 (−0.80)+0.03 (−0.76)
Unigram+0.56 (+0.23)+0.14 (−0.39)+0.20 (−0.67)
SynthID+0.76 (+0.47)+0.81 (+0.27)+0.77 (+0.45)
EXPEdit+0.68 (+0.26)+0.75 (+0.39)+0.70 (+0.46)
DIP+0.46 (−0.03)+0.43 (+0.05)+0.50 (−0.05)
XSIR+0.42 (−0.09)+0.10 (−0.71)+0.00 (−0.76)
All schemes+0.58 (−0.09)+0.38 (−0.80)+0.37 (−0.76)
Figure 4: Per-language detection diagnostics, subset all, α=0.01. Within each panel, left: TPR at the global empirical-FPR threshold τg; right: threshold-free AUC. Rows are (scheme, generator) cells; columns are the eleven evaluation languages, grouped by typological family.
Figure 4: Per-language detection diagnostics, subset all, α=0.01. Within each panel, left: TPR at the global empirical-FPR threshold τg; right: threshold-free AUC. Rows are (scheme, generator) cells; columns are the eleven evaluation languages, grouped by typological family.
Table 4: Per-cell language adherence on the eleven evaluation languages, base-FLORES (left) and instruct-AYA (right). NWM% is the GlotLID-v3 on-target rate for nowatermark generations; Δ¯ is the mean shift in percentage points across the six watermarked schemes. mean rows are unweighted means across the eleven languages.
Base regime (FLORES+ continuation prompts)Instruct regime (AYA native-speaker prompts)
MistralGemmaQwenMistralGemmaQwen
LangNWM%Δ¯NWM%Δ¯NWM%Δ¯NWM%Δ¯NWM%Δ¯NWM%Δ¯
eng99.6−0.398.8−2.098.4−0.692.6−0.980.6−1.793.2−2.0
fra99.2−1.197.6−4.089.2−0.291.4+0.586.2−1.794.8−4.9
spa100.0−1.496.8−2.691.2+1.195.2−1.362.6+2.792.0−5.0
por99.2−0.596.8−1.098.2−0.896.2−2.070.6+3.096.4−4.3
nld96.2−7.095.8+1.680.8−4.673.2−3.189.2−0.490.4−5.8
zho97.6−1.892.8−1.296.4−4.289.0−7.995.6−0.498.4−0.6
jpn96.8−3.395.6−1.699.8−3.371.0−10.794.0−1.998.4−7.4
arb89.2−3.087.6−3.887.8+2.967.6−15.684.2−1.281.0−15.0
hin96.2−9.188.2−2.398.8−3.952.2−6.583.8−1.483.2+0.1
tur89.0+1.396.8−4.357.2−1.181.8−8.192.8−0.695.6−13.5
vie98.6−4.777.0+4.799.0−1.079.2−7.398.8−0.695.2−6.4
mean96.5−2.893.1−1.590.6−1.480.9−5.785.3−0.492.6−5.9
(b) Instruct-AYA regime.
(b) Instruct-AYA regime.
Table 5: Length-distribution similarity between LID subsets and the all-subset baseline. Each row is a (regime, subset) pair; ratio=subset-mean length/all-subset mean length, computed per (regime, model, language, watermark) cell and summarized across cells. Cells require n≥100 in both subsets to enter the summary; n skipped counts cells that fail this floor (typically off_target on high-adherence cells).
RegimeSubsetn keptn skippedMeanMedianMinMax
Baseon_target_only23101.0031.0000.9961.081
Baseoff_target_only82230.9290.9080.8591.023
Instructon_target_only23101.0001.0000.9641.057
Instructoff_target_only262051.0251.0170.9471.119
Figure 5: PPL preservation under mGPT-1.3B cross-reference. Format and color scale match Figure 1.
Figure 5: PPL preservation under mGPT-1.3B cross-reference. Format and color scale match Figure 1.
Table 6: LID categorization robustness: per-cell adherence under a 3×3 sweep of the GlotLID per-sentence confidence floor (conf_floor) and the per-text on-target sentence fraction (on_target_ratio). ρ is the Spearman rank-correlation between the per-cell on-target-rate vector under each configuration and the default (†). Mean on-target % is the unweighted mean across the 462 cells of the evaluation grid. Mean/Max |Δ| are the mean and maximum absolute per-cell shift in on-target rate (pp) relative to default.
conf_flooron_target_ratioρ vs defaultMean on-target %Mean |Δ| (pp)Max |Δ| (pp)
0.30.70.934792.02.1913.60
0.30.80.946890.70.9113.00
0.30.90.862786.34.8731.00
0.50.70.991091.11.2611.80
0.5†0.81.000089.80.000.00
0.50.90.931785.44.4331.20
0.70.70.948687.93.7624.60
0.70.80.962786.83.0424.80
0.70.90.938582.57.3631.80
(b) Instruct-AYA regime.
(b) Instruct-AYA regime.
Table 7: Per-language calibration gap Δl=TPRl​(τl)−TPRl​(τg), base-FLORES (top) and instruct-AYA (bottom), subset all, α=0.01. Positive: per-language calibration recovers detection. Negative: per-language calibration worsens detection. Zero: detection insensitive to calibration locality. Row means equal −1× the Δ column of Table 2.
GermanicRomanceIndicSemiticSiniticJaponicTurkicAus.
SchemeGen.engnldfraspaporhinarbzhojpnturvie
Base regime (FLORES+ continuation prompts)
KGWMistral0.000+0.004+0.004+0.0020.000−0.0080.000−0.0060.000−0.004+0.002
Gemma+0.006+0.030+0.008+0.016+0.038−0.020+0.044−0.010−0.0640.000+0.010
Qwen+0.028+0.018−0.002−0.002+0.002−0.0080.000+0.0080.000−0.0300.000
UnigramMistral+0.074+0.354+0.218+0.292+0.142−0.062+0.004+0.064−0.018+0.074−0.012
Gemma+0.996+0.992+0.968+0.990+0.9600.000+0.958+0.976+0.984+0.9600.000
Qwen+0.024+0.034−0.040+0.002+0.022+0.106+0.230+0.134−0.008+0.092−0.034
SynthIDMistral0.0000.0000.0000.0000.0000.0000.000+0.0020.0000.0000.000
Gemma0.0000.0000.0000.0000.0000.0000.0000.0000.0000.0000.000
Qwen0.0000.0000.0000.0000.000+0.0020.000+0.0060.000−0.0040.000
EXPEditMistral0.000+0.0020.0000.0000.0000.0000.0000.0000.0000.0000.000
Gemma0.0000.0000.0000.0000.0000.0000.0000.0000.0000.0000.000
Qwen+0.002−0.0060.0000.0000.000+0.0120.0000.0000.000+0.0020.000
DIPMistral0.0000.000+0.002+0.002+0.002+0.0020.000−0.0140.000+0.0040.000
Gemma0.000−0.004+0.0020.000+0.002+0.010+0.004+0.004+0.018−0.0100.000
Qwen0.000−0.008+0.004+0.008−0.004−0.004+0.002+0.0100.000−0.028+0.002
XSIRMistral+0.024−0.054+0.166+0.106+0.086−0.022−0.066+0.092+0.014+0.014+0.178
Gemma+0.028+0.054+0.054+0.114+0.026+0.146+0.056−0.042+0.168+0.122−0.030
Qwen+0.516+0.774+0.770+0.802+0.912−0.286+0.684+0.638+0.464+0.360+0.732
Instruct regime (AYA native-speaker prompts)
KGWMistral+0.112+0.032−0.250+0.028+0.118−0.016+0.014+0.030−0.018+0.012−0.004
Gemma+0.030−0.070+0.108+0.046+0.112−0.0380.000−0.036−0.082+0.134+0.034
Qwen+0.012+0.014−0.036−0.058+0.110−0.106+0.138+0.016+0.084+0.058+0.140
UnigramMistral+0.340+0.380+0.090+0.282+0.316−0.1060.000+0.070−0.178+0.282+0.052
Gemma+0.024−0.054+0.174+0.178+0.056−0.070−0.036+0.376+0.276+0.272+0.050
Qwen+0.084−0.018−0.108−0.012+0.078+0.128+0.438+0.352+0.094+0.088−0.132
SynthIDMistral+0.056+0.020−0.026−0.034−0.048+0.0160.0000.000−0.0220.000+0.016
Gemma+0.050−0.030+0.046−0.022−0.020−0.022−0.124+0.038−0.102+0.094+0.036
Qwen+0.018−0.002+0.010−0.022−0.002+0.0160.0000.0000.0000.0000.000
Figure 6: On-target minus all-subset preservation, base-FLORES regime (top) and instruct-AYA regime (bottom). Each cell is Δ=preservationon-target−preservationall for the indicated paradigm; rows are (scheme, generator), columns are languages. Positive (red): on-target restriction raises preservation; negative (blue): on-target restriction lowers it. Color scale fixed at [−0.3,+0.3]; values outside this range saturate. The EXPEdit-Mistral, Arabic, BERTScore F1 entry in the instruct panel is omitted (—) because the on-target sample size falls below the n≥100 paired-set floor.
Figure 6: On-target minus all-subset preservation, base-FLORES regime (top) and instruct-AYA regime (bottom). Each cell is Δ=preservationon-target−preservationall for the indicated paradigm; rows are (scheme, generator), columns are languages. Positive (red): on-target restriction raises preservation; negative (blue): on-target restriction lowers it. Color scale fixed at [−0.3,+0.3]; values outside this range saturate. The EXPEdit-Mistral, Arabic, BERTScore F1 entry in the instruct panel is omitted (—) because the on-target sample size falls below the n≥100 paired-set floor.
Table 8: Detection fairness summary at α=0.05, base-FLORES (top) and instruct-AYA (bottom), subset all. Format mirrors Table 2; AUC columns are omitted (threshold-independent, identical to Table 2) and bootstrap CIs on Mean TPR are omitted for compactness. L=11 throughout. Values below 10−3 reported as <0.001.
TPR at τg (global)TPR at τl (per-lang)
SchemeGen.Mean TPRMin TPRDI<.8GE2Btw. %Mean TPRΔ
Base regime (FLORES+ continuation prompts)
KGWMistral0.9970.9920<0.00161.10.997+0.000
Gemma0.9910.9800<0.00193.20.993−0.002
Qwen0.9870.9520<0.00157.80.986+0.000
UnigramMistral0.9770.9200<0.00154.90.988−0.011
Gemma0.9880.9760<0.00155.90.898+0.089
Qwen0.9610.86000.00186.30.979−0.018
SynthIDMistral1.0000.9980<0.001100.01.000+0.000
Gemma1.0000.9980<0.001100.01.000+0.000
Qwen0.9990.9960<0.00174.30.999+0.000
EXPEditMistral1.0000.9960<0.001100.01.000+0.000
Gemma1.0000.9980<0.001100.01.000+0.000
Qwen0.9950.9660<0.00199.20.995−0.000
DIPMistral0.9900.9720<0.00194.70.989+0.001
Gemma0.9810.9420<0.00197.90.981−0.001
Qwen0.9660.89600.00191.60.966+0.000
XSIRMistral0.9240.86800.00183.30.935−0.011
Gemma0.9290.84000.00192.80.944−0.015
Qwen0.5030.27690.07694.60.919−0.416
Instruct regime (AYA native-speaker prompts)
KGWMistral0.8010.62220.00685.00.813−0.012
Gemma0.6160.474100.01579.70.617−0.001
Qwen0.8600.74010.00378.10.869−0.009
UnigramMistral0.5170.09050.14995.80.659−0.143
Gemma0.3800.15450.05080.30.408−0.028
Qwen0.7270.36640.01893.80.790−0.063
SynthIDMistral0.9670.9260<0.00177.80.969−0.002
Gemma0.7910.66240.00575.90.785+0.005
Qwen0.9710.8980<0.00195.70.973−0.002
Figure 7: Joint per-language detection–quality landscape, BERTScore-rescaled (top) and PPL-XGLM (bottom) paradigms. Format matches Figure 3: markers are per-language (TPR​@​τg,quality) pairs at α=0.01, n=33 per regime (11 languages × 3 generators); circles are base-FLORES, triangles are instruct-AYA; color encodes typological family; sB, sI are mean pairwise Euclidean distance within each regime.
Figure 7: Joint per-language detection–quality landscape, BERTScore-rescaled (top) and PPL-XGLM (bottom) paradigms. Format matches Figure 3: markers are per-language (TPR​@​τg,quality) pairs at α=0.01, n=33 per regime (11 languages × 3 generators); circles are base-FLORES, triangles are instruct-AYA; color encodes typological family; sB, sI are mean pairwise Euclidean distance within each regime.
Table 9: Detection summary on the on-target subset, α=0.01, global threshold τg, base-FLORES (left) and instruct-AYA (right). Thresholds re-calibrated on each cell’s on-target unwatermarked pool; L=11 throughout. The off-target subset is omitted: in the base regime no cell meets the nneg≥200 inclusion floor, and in the instruct regime only six (Mistral-generator) cells pass the relaxed α=0.05 floor with L≤3, insufficient for the fairness aggregates this appendix reports. The Δ vs. all columns report change in Mean TPR relative to subset all (Table 2).
Base regime (FLORES+ continuation)Instruct regime (AYA native-speaker)
SchemeGen.Mean TPRMin TPRDI<.8Δ vs. allMean TPRMin TPRDI<.8Δ vs. all
KGWMistral0.9950.9830+0.0030.6810.4256+0.055
Gemma0.9730.9390+0.0040.4260.26410+0.037
Qwen0.9870.9630+0.0220.6970.5208+0.008
UnigramMistral0.8850.6542+0.0170.3880.0117+0.077
Gemma0.9870.9740+0.9870.0200.0009−0.017
Qwen0.9250.7621+0.0270.5470.1764+0.001
SynthIDMistral1.0000.99800.0000.9240.7741+0.019
Gemma1.0000.99800.0000.6100.42610+0.008
Qwen0.9990.9960+0.0020.9410.76410.000
EXPEditMistral1.0001.0000+0.0010.6750.2345+0.098
Gemma1.0000.99700.0000.2320.09110+0.007
Qwen0.9940.9410+0.0030.5080.05910+0.012
DIPMistral0.9800.9520+0.0050.4880.2669+0.033
Gemma0.9770.9540+0.0140.1980.10510+0.002
Qwen0.9570.8480+0.0260.5070.2008−0.006
XSIRMistral0.7990.6882+0.0150.2960.08110+0.030
Gemma0.7930.6003+0.0270.0920.01910+0.011
Qwen0.2070.04710−0.0080.0470.00010+0.004
Table 10: Bottom-quintile-mean Rawlsian floor as a stability companion to Min TPR (Table 2), base-FLORES (left) and instruct-AYA (right), subset all, α=0.01, global threshold τg. With L=11 evaluation languages the bottom quintile is ⌊0.2⋅11⌋=2 languages; the columns report the mean TPR of the two worst-performing languages per cell, with the languages listed minimum first (arbitrary tie-breaking). The Unigram-Gemma base cell is degenerate (all eleven per-language TPRs are zero); its bottom-quintile language identities are reported but uninformative.
Base regime (FLORES+)Instruct regime (AYA)
SchemeGen.Min TPRBot.-q meanBot.-q langsMin TPRBot.-q meanBot.-q langs
KGWMistral0.9780.979zho, hin0.3860.447fra, por
Gemma0.9340.943arb, por0.2220.251por, vie
Qwen0.8800.894tur, nld0.5100.564por, nld
UnigramMistral0.6380.672nld, spa0.0220.029spa, fra
Gemma0.0000.000eng, fra0.0100.010por, tur
Qwen0.7220.744tur, arb0.2100.306arb, hin
SynthIDMistral0.9980.998zho, tur0.7820.788fra, spa
Gemma0.9980.999vie, eng0.4220.437vie, por
Qwen0.9900.992tur, zho0.7920.842hin, por
EXPEditMistral0.9960.997zho, nld0.2380.244eng, spa
Gemma0.9980.998jpn, vie0.0840.101jpn, por
Qwen0.9460.968hin, nld0.1740.226hin, eng
DIPMistral0.9380.950zho, jpn0.2640.265fra, spa
Gemma0.9140.917hin, vie0.1100.115por, jpn
Qwen0.7980.819tur, hin0.2960.345hin, fra
XSIRMistral0.6720.673zho, vie0.0880.116fra, spa
Gemma0.5880.635jpn, hin0.0160.019tur, jpn
Qwen0.0500.064por, fra0.0000.000nld, vie
Table 11: Robustness of the generalized-entropy decomposition to choice of inequality measure: GE2 (half the squared coefficient of variation, more sensitive to inequality at the top of the per-language TPR distribution) versus GE0 (mean log deviation, more sensitive to inequality at the bottom). Base-FLORES (left) and instruct-AYA (right), subset all, α=0.01, global threshold τg; typological partition as in Table 2. Total values below 10−3 reported as <0.001. The Unigram-Gemma base GE0 entry is obtained by replacing zero TPRs with a small ϵ (per the strict-positivity requirement of GE0); the resulting Total=0 and Btw.%=100.0 are regularization artifacts, marked †.
Base regime (FLORES+)Instruct regime (AYA)
GE2GE0GE2GE0
SchemeGen.TotalBtw. %TotalBtw. %TotalBtw. %TotalBtw. %
KGWMistral<0.00183.6<0.00183.70.02695.30.02791.4
Gemma<0.00172.4<0.00172.70.05382.40.04673.3
Qwen0.00185.10.00185.10.01178.60.01172.5
UnigramMistral0.00861.30.00955.70.29799.80.55694.8
Gemma<0.001†100.0†0.22681.20.24373.9
Qwen0.00493.80.00594.60.03291.70.04694.6
SynthIDMistral<0.001100.0<0.001100.00.00377.70.00375.6
Gemma<0.001100.0<0.001100.00.02079.70.01975.6
Qwen<0.00197.5<0.00197.50.00291.90.00292.2
EXPEditMistral<0.00188.0<0.00188.10.06780.00.08768.2
Gemma<0.001100.0<0.001100.00.07184.20.07779.0
Qwen<0.00199.2<0.00199.20.08084.40.09381.4
DIPMistral<0.00192.6<0.00192.80.05899.50.05999.2
Gemma<0.00199.5<0.00199.50.08491.60.06884.4
Qwen0.00394.70.00394.60.03599.60.03599.6
XSIRMistral0.00491.50.00492.10.18591.80.16083.4
Gemma0.00696.00.00796.30.17576.50.22081.1
Qwen0.49398.90.32692.53.78799.93.24479.0
Table 12: Robustness of the GE2 between-share to choice of partition: typological family (8 groups: Germanic, Romance, Indic, Semitic, Sinitic, Japonic, Turkic, Austroasiatic, with 5 of 11 languages in non-singleton groups) vs. script (4 groups: Latin pooling eng, nld, fra, spa, por, tur, vie; Devanagari=hin; Arabic=arb; Han pooling zho, jpn; with 9 of 11 languages in non-singleton groups). Base-FLORES (left) and instruct-AYA (right), subset all, α=0.01, global threshold τg. GE2 Total depends only on the per-language vector and is partition-independent (shown once per regime). Totals below 10−3 reported as <0.001; for these cells the between-share values involve ratios of near-zero variance components and should not be over-interpreted.
Base regime (FLORES+)Instruct regime (AYA)
Btw. %Btw. %
SchemeGen.GE2FamilyScriptGE2FamilyScript
KGWMistral<0.00183.646.20.02695.332.8
Gemma<0.00172.462.20.05382.429.8
Qwen0.00185.126.20.01178.626.7
UnigramMistral0.00861.328.80.29799.846.3
Gemma0.22681.261.7
Qwen0.00493.826.10.03291.775.0
SynthIDMistral<0.001100.017.10.00377.733.0
Gemma<0.001100.05.70.02079.747.1
Qwen<0.00197.57.50.00291.979.3
EXPEditMistral<0.00188.031.70.06780.030.2
Gemma<0.001100.017.10.07184.28.5
Qwen<0.00199.295.80.08084.441.6
DIPMistral<0.00192.669.30.05899.538.9
Gemma<0.00199.548.20.08491.625.2
Qwen0.00394.732.50.03599.653.4
XSIRMistral0.00491.554.80.18591.864.7
Gemma0.00696.012.70.17576.57.6
Qwen0.49398.993.53.78799.999.8
Table 13: Quality fairness summary by scheme, generator, regime, and measurement, subset all. Each row reports mean per-language preservation with the per-cell minimum and floor language (ISO 639-3 superscript) in parentheses. DI<.8 counts languages with per-language preservation below 0.8×maxl. GE2 uses the typological partition (Germanic, Romance, Indic, Semitic, Sinitic, Japonic, Turkic, Austroasiatic); Btw. % is its between-family share. BERTScore mean (min) is reported on rescaled F1 [28]; DI<.8 and GE2 are omitted for BERTScore because neither registers signal on the bounded raw scale and both behave pathologically on the rescaled scale (§5.2).
MAUVEBERTScore F1 (rescaled)PPL preservation (XGLM-7.5B)
SchemeGenmean (min)DI<.8GE2Btw. %mean (min)mean (min)DI<.8GE2Btw. %
Base regime (FLORES+ continuation)
KGWMistral0.947 (0.879)jpn00.00186.0-0.041 (-0.079)vie0.866 (0.793)spa00.00266.4
KGWGemma0.855 (0.435)arb10.01398.9-0.081 (-0.194)vie0.899 (0.812)fra00.00291.0
KGWQwen0.889 (0.471)hin10.01295.6-0.034 (-0.091)tur0.834 (0.686)hin30.00594.0
UnigramMistral0.747 (0.420)arb50.02649.5-0.070 (-0.134)arb0.872 (0.811)spa00.00296.4
UnigramGemma0.638 (0.444)spa60.01767.8-0.098 (-0.191)jpn0.911 (0.815)nld00.00268.3
UnigramQwen0.734 (0.123)hin30.04397.2-0.055 (-0.128)zho0.896 (0.647)hin10.00598.0
SynthIDMistral0.187 (0.020)arb100.43559.0-0.115 (-0.163)vie0.299 (0.145)arb70.04888.4
SynthIDGemma0.095 (0.029)tur100.64143.0-0.146 (-0.254)vie0.269 (0.195)tur70.02173.2
SynthIDQwen0.183 (0.035)vie100.45553.0-0.112 (-0.188)tur0.308 (0.203)arb90.02878.2
EXPEditMistral0.079 (0.006)arb100.53657.1-0.149 (-0.228)vie0.218 (0.079)arb60.08599.1
EXPEditGemma0.013 (0.006)jpn100.57443.4-0.220 (-0.333)vie0.135 (0.076)zho90.03382.7
EXPEditQwen0.067 (0.013)arb100.87447.9-0.145 (-0.245)tur0.224 (0.135)tur90.04274.6
DIPMistral0.973 (0.933)hin00.00095.6-0.015 (-0.040)vie0.809 (0.750)hin00.00177.6
DIPGemma0.932 (0.878)jpn00.00194.7-0.053 (-0.183)vie0.774 (0.702)nld00.00168.0
DIPQwen0.947 (0.864)hin00.00181.6-0.008 (-0.071)tur0.811 (0.714)tur00.00292.7
XSIRMistral0.859 (0.642)hin20.00686.6-0.062 (-0.110)hin0.922 (0.816)nld00.00273.2
XSIRGemma0.683 (0.461)fra60.01880.9-0.123 (-0.226)vie0.925 (0.785)nld10.00275.8
XSIRQwen0.850 (0.603)hin20.00889.9-0.042 (-0.129)tur0.883 (0.745)hin10.00386.4
Instruct regime (AYA native-speaker)
KGWMistral0.989 (0.970)arb00.00058.90.255 (0.088)arb0.892 (0.853)jpn00.00191.0
KGWGemma0.990 (0.959)arb00.00065.30.397 (0.338)tur0.967 (0.941)hin00.00080.0
KGWQwen0.987 (0.954)hin00.00090.80.300 (0.192)tur0.885 (0.749)hin10.00292.5
UnigramMistral0.981 (0.938)nld00.00040.30.240 (0.087)arb0.897 (0.835)hin00.00197.4
UnigramGemma0.996 (0.991)jpn00.00054.10.400 (0.329)tur0.965 (0.941)nld00.00078.0
UnigramQwen0.980 (0.927)hin00.00097.70.297 (0.184)tur0.885 (0.700)hin10.00395.9
SynthIDMistral0.721 (0.066)arb40.10799.60.149 (-0.117)arb0.595 (0.185)arb70.07095.4
SynthIDGemma0.994 (0.979)fra00.00050.00.365 (0.296)tur0.931 (0.885)zho00.00090.0
SynthIDQwen0.833 (0.237)tur20.03397.10.219 (0.047)tur0.645 (0.442)tur50.01490.7
Table 14: Single-covariate R2 for per-language detection TPR at τg, α=0.01, subset all. Each fit regresses the per-language outcome vector on a single covariate over 33 (language, generator) points. Stars from the slope p-value: ∗p<0.05, p∗⁣∗<0.01.
CovariateKGWUnigramSynthIDEXPEditDIPXSIR
Base regime (FLORES+ continuation, n=33)
Fertility (tokens/sent.)0.0160.0520.0080.834∗∗0.238∗∗0.001
NWM surprisal (XGLM)0.0660.0740.0580.261∗∗0.156∗0.003
Latin script0.0290.0010.0050.0510.0090.019
NWM adherence0.600∗∗0.0060.415∗∗0.0010.461∗∗0.091
Instruct regime (AYA native-speaker, n=33)
Fertility (tokens/sent.)0.0170.0460.0090.0010.0000.107
NWM surprisal (XGLM)0.0240.0030.0090.0400.0170.011
Latin script0.0460.0140.0350.0490.0400.058
NWM adherence0.0000.0000.0030.0470.0000.291∗∗
Table 15: Single-covariate R2 for per-language MAUVE preservation. Each fit regresses the per-language outcome vector on a single covariate over 33 (language, generator) points. Stars from the slope p-value: ∗p<0.05, p∗⁣∗<0.01. BERTScore-rescaled and PPL-XGLM paradigm-stability tables in §F.1.
CovariateKGWUnigramSynthIDEXPEditDIPXSIR
Base regime (FLORES+ continuation, n=33)
Fertility (tokens/sent.)0.392∗∗0.249∗∗0.0270.0050.234∗∗0.093
NWM surprisal (XGLM)0.0260.0230.0180.0010.0070.009
Latin script0.136∗0.0360.0620.0520.0830.000
NWM adherence0.0310.0070.1010.0740.219∗∗0.072
Instruct regime (AYA native-speaker, n=33)
Fertility (tokens/sent.)0.346∗∗0.464∗∗0.0270.0160.0190.168∗
NWM surprisal (XGLM)0.0080.152∗0.0220.0520.0050.002
Latin script0.122∗0.0510.0180.0410.0520.082
NWM adherence0.0760.0990.153∗0.230∗∗0.0530.047
Table 16: Single-covariate R2 for BERTScore-rescaled preservation. Stars from the slope p-value: ∗p<0.05, p∗⁣∗<0.01.
CovariateKGWUnigramSynthIDEXPEditDIPXSIR
Base regime (n=33)
Fertility (tokens/sent.)0.0000.0130.0240.0540.0390.095
NWM surprisal (XGLM)0.0360.0160.0000.0320.0390.094
Latin script0.0680.207∗∗0.0040.0220.0030.008
NWM adherence0.239∗∗0.0040.378∗∗0.259∗∗0.298∗∗0.188∗
Instruct regime (n=33)
Fertility (tokens/sent.)0.0720.0690.0430.0260.0470.051
NWM surprisal (XGLM)0.0340.0340.0300.0210.0370.026
Latin script0.0460.0450.0420.0370.0440.056
NWM adherence0.0960.0950.119∗0.125∗0.123∗0.110
Table 17: Single-covariate R2 for PPL-XGLM preservation. Stars from the slope p-value: ∗p<0.05, p∗⁣∗<0.01.
CovariateKGWUnigramSynthIDEXPEditDIPXSIR
Base regime (n=33)
Fertility (tokens/sent.)0.191∗0.327∗∗0.0130.0000.121∗0.148∗
NWM surprisal (XGLM)0.0030.179∗0.0170.0180.0940.183∗
Latin script0.0220.0010.126∗0.1050.0520.056
NWM adherence0.186∗0.0140.164∗0.123∗0.1090.015
Instruct regime (n=33)
Fertility (tokens/sent.)0.446∗∗0.559∗∗0.119∗0.0700.225∗∗0.340∗∗
NWM surprisal (XGLM)0.0440.0290.0100.0000.0470.016
Latin script0.0860.0770.0700.0840.0460.012
NWM adherence0.0200.0240.0960.217∗∗0.0230.006

왜 중요한가

다국어로 퍼지는 AI 생성 허위정보에 대응하기 위해 워터마킹이 대책으로 제시되고 있는데, 특정 언어 계열에서 이 기술이 조용히 실패한다면 그 언어를 쓰는 사람들은 보호를 덜 받으면서도 품질 저하라는 대가는 더 크게 치르게 된다. 이 연구는 영어 결과를 그대로 다른 언어에 적용해도 된다는 가정을 깨고, 배포 전에 이런 격차를 잡아낼 수 있는 구체적인 검증 방법을 제시한다.

이 논문의 용어

  • 워터마킹 · AI가 생성한 텍스트에 보이지 않는 통계적 흔적을 남겨 나중에 기계 생성물임을 판별하게 하는 기술
  • 탐지 임계값 · 워터마크 점수가 이 값을 넘으면 워터마크가 있다고 판정하는 기준선
  • AUC · 특정 기준선에 의존하지 않고 두 집단이 얼마나 잘 구분되는지 나타내는 지표
  • 언어 유형학적 계열 · 문법이나 어휘 구조가 비슷한 언어들의 묶음, 예를 들어 게르만어파나 셈어파
  • distortion-free 방식 · 생성되는 텍스트의 확률분포를 이론상 바꾸지 않도록 설계된 워터마킹 기법

논문 원문 초록 (영문)

Watermarking schemes for large language model output are evaluated almost exclusively on English text using each scheme's detection threshold and a narrow set of quality measurements. Multilingual deployment exposes evaluation-design choices that are inconsequential on English but determine conclusions cross-lingually. We propose an evaluation framework with four components: detection thresholds calibrated empirically per deployment context, a threshold-independent companion measurement that distinguishes calibration failures from detection failures, three disjoint quality measurement paradigms (distributional, paired-semantic, and reference-perplexity), and a generalized-entropy decomposition of cross-language disparity over a typological family partition. Applied to six watermarking schemes, three open-weight generators, eleven languages spanning four scripts and eight typological families, and both base and instruction-tuned regimes, the framework reveals failure modes that single-language single-paradigm evaluation cannot surface. Across detection and quality, observed disparity is predominantly between-family on the typological partition, indicating that cross-lingual fairness gaps in watermarking are structural to language properties rather than idiosyncratic to particular languages.

저자 · Alexander Nemecek, Osama Zafar, Debargha Ganguly, Vikash Singh, Vipin Chaudhary, Erman Ayday

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사

그림 출처: Alexander Nemecek et al., arXiv:2608.20047, CC BY 4.0