每天早上一封邮件,把昨天的 AI 梳理好订阅邮件

METAL LAB

Abliteration Mitigation via Refusal Aliases

arXiv:2608.180932026-08-20

Abliteration, the removal of refusal capabilities from large language models by projecting weight matrices orthogonal to an extracted refusal direction, has emerged as a prominent safety concern through its ability to bypass post-training alignment using only a small set of contrastive prompts. We find that existing defenses commonly overlook the cause of abliteration; that is, how easily the refusal direction can be extracted. To hinder this proce

作者 · Nathan Truong

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道