매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs

arXiv:2608.181312026-08-20

Current safety alignment training for Large Language Models (LLMs) are heavily English-centric. When such safety filters fail for non-English languages, the consequences are immediate and user-facing: voice assistants and spoken dialogue systems may produce stereotype-reinforcing outputs, bypassing the standard English-focused safety alignments and propagating harmful bias to non-English speaking communities. For spoken language technologies deploy

저자 · Namya Bhatnagar

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사