AdaGuard: An Adaptive Guard Model with User-defined Policies
AdaGuard is built to judge agent behavior against policies supplied at inference time, not a fixed safety taxonomy.
The paper introduces AdaptiveSafety, a dataset with 10,939 training examples and 1,000 test examples spanning policies with 1 to 100 rules. Each example includes explanations and the full set of violated rules, with counterfactuals meant to show when a policy or behavior change flips compliance. The authors also propose SafePO, a reinforcement learning method for improving violation identification while balancing explanations and final verdicts. Their 4B AdaGuard model reports 89.30% binary accuracy on AdaptiveSafety and 71.82% on DynaBench. HF Daily Papers' note
The paper introduces AdaptiveSafety, a dataset with 10,939 training examples and 1,000 test examples spanning policies with 1 to 100 rules. Each example includes explanations and the full set of violated rules, with counterfactuals meant to show when a policy or behavior change flips compliance. The authors also propose SafePO, a reinforcement learning method for improving violation identification while balancing explanations and final verdicts. Their 4B AdaGuard model reports 89.30% binary accuracy on AdaptiveSafety and 71.82% on DynaBench. HF Daily Papers' note
score 4