One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

arXiv:2608.175972026-08-17

Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent state, permissions, and external actions. Existing safety benchmarks mainly target individual attack mechanisms or a limited subset of operational settings, making it difficult to compare how safety failures emerge across different harness responsibilities. We present HarnessRisk, a lifecycle oriented benchmark that organizes agent harness safety into six operational phases including

Authors · Yajing Bai

Read on arXiv

Latest papers

All papers →

Latest from METAL LAB