POLAR: Ontology-Guided Risk Prevention for Tool-Calling LLM Agents
POLAR blocks risky tool calls before execution by scoring whether an action can be reversed.
The paper frames that as a structured, auditable alternative to safety checks that only react after a tool-use agent has already made a mistake. Its ontology produces a graded reversibility score from a candidate inverse sequence, then prunes calls below a threshold. On tau^2-bench, the gains are uneven: airline tasks improve for four of six agents, while only eight of eighteen model-domain cells improve overall. The authors also note that reward is not the same thing as prevented harm. ArXiv · AI/CL/LG's note
The paper frames that as a structured, auditable alternative to safety checks that only react after a tool-use agent has already made a mistake. Its ontology produces a graded reversibility score from a candidate inverse sequence, then prunes calls below a threshold. On tau^2-bench, the gains are uneven: airline tasks improve for four of six agents, while only eight of eighteen model-domain cells improve overall. The authors also note that reward is not the same thing as prevented harm. ArXiv · AI/CL/LG's note
score 5