每天早上一封邮件,把昨天的 AI 梳理好订阅邮件

METAL LAB

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

arXiv:2608.175972026-08-17

Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent state, permissions, and external actions. Existing safety benchmarks mainly target individual attack mechanisms or a limited subset of operational settings, making it difficult to compare how safety failures emerge across different harness responsibilities. We present HarnessRisk, a lifecycle oriented benchmark that organizes agent harness safety into six operational phases including

作者 · Yajing Bai

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道