One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

Abliteration Mitigation via Refusal Aliases

arXiv:2608.180932026-08-20

Abliteration, the removal of refusal capabilities from large language models by projecting weight matrices orthogonal to an extracted refusal direction, has emerged as a prominent safety concern through its ability to bypass post-training alignment using only a small set of contrastive prompts. We find that existing defenses commonly overlook the cause of abliteration; that is, how easily the refusal direction can be extracted. To hinder this proce

Authors · Nathan Truong

Read on arXiv

Latest papers

All papers →

Latest from METAL LAB