HazardAuditor: From Executable Threats to Safer Computer-Use Agents
HazardAuditor trains guards on what agents actually do while they run.
The paper says current guard models miss risks that appear during browser, terminal, file-system, and service interactions. HazardAuditor runs agents including Claude Code, Codex, Hermes, and OpenClaw in controlled environments, then converts their actions into a shared event format. Its Guard Policy Optimization method treats the safety verdict as the main training target instead of letting long rationales dominate. The authors report accuracy gains of up to 16.5 percentage points over the strongest prior guard. HF Daily Papers' note
The paper says current guard models miss risks that appear during browser, terminal, file-system, and service interactions. HazardAuditor runs agents including Claude Code, Codex, Hermes, and OpenClaw in controlled environments, then converts their actions into a shared event format. Its Guard Policy Optimization method treats the safety verdict as the main training target instead of letting long rationales dominate. The authors report accuracy gains of up to 16.5 percentage points over the strongest prior guard. HF Daily Papers' note
score 5