매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

Calling the Bluff: Detecting Ever-Shifting Harmful Chat Dialogue via Ordered Reasoning Chain Regularization

arXiv:2608.084512026-08-11

arXiv:2608.08451v1 Announce Type: new Abstract: Harmful chat dialogues are ever-shifting through type-shifting and lexical evasion, yet we find they share invariant principles, i.e., an Ordered Reasoning Chain (ORC) of recurring topics, harm language indicators, severity hierarchies, and type characteristics, which can help us capture the key information in the frequently changing lexical expressions. We propose BRACE, which encodes the ORC as four differentiable stages (Topic -> Indicator -> Se

저자 · Haojie Yu, Ziyou Jiang, Junjie Wang, Mingyang Li, Yuekai Huang, Jie Huang, Qing Wang

arXiv에서 원문 보기