매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

Abliteration Mitigation via Refusal Aliases

arXiv:2608.180932026-08-20

Abliteration, the removal of refusal capabilities from large language models by projecting weight matrices orthogonal to an extracted refusal direction, has emerged as a prominent safety concern through its ability to bypass post-training alignment using only a small set of contrastive prompts. We find that existing defenses commonly overlook the cause of abliteration; that is, how easily the refusal direction can be extracted. To hinder this proce

저자 · Nathan Truong

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사