StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
The paper puts the guardrail before the tool call, checking each agent action before it runs.
StepGuard is trained to judge safe and unsafe step-level actions in otherwise matched agent trajectories. Its data engine, StepGen, creates paired trajectories that differ at the risky step, and Balance-GRPO adjusts learning to reduce both over-blocking and missed attacks. In the reported tests, it led open-weight guard models on average accuracy and was comparable to GPT-5.4. Used on AgentDojo and AgentDyn, it cut mean attack success by 77.3% versus no guard, with mean utility down 2.8 percentage points. ArXiv · AI/CL/LG's note
StepGuard is trained to judge safe and unsafe step-level actions in otherwise matched agent trajectories. Its data engine, StepGen, creates paired trajectories that differ at the risky step, and Balance-GRPO adjusts learning to reduce both over-blocking and missed attacks. In the reported tests, it led open-weight guard models on average accuracy and was comparable to GPT-5.4. Used on AgentDojo and AgentDyn, it cut mean attack success by 77.3% versus no guard, with mean utility down 2.8 percentage points. ArXiv · AI/CL/LG's note
score 5