PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety
The benchmark finds current LLMs still struggle to intervene at the right moment in risky agent workflows.
PASTABench covers 1,139 multi-turn trajectories across five risk categories and 13 subcategories. The paper frames safety monitoring around whether to intervene, when to intervene, and what risk is present. Its best-tested model hit only 40.74% optimal-timing interventions. The authors also report that some smaller models’ safety scores collapse when hazard keywords are neutralized, suggesting keyword sensitivity rather than deeper risk understanding. ArXiv · AI/CL/LG's note
PASTABench covers 1,139 multi-turn trajectories across five risk categories and 13 subcategories. The paper frames safety monitoring around whether to intervene, when to intervene, and what risk is present. Its best-tested model hit only 40.74% optimal-timing interventions. The authors also report that some smaller models’ safety scores collapse when hazard keywords are neutralized, suggesting keyword sensitivity rather than deeper risk understanding. ArXiv · AI/CL/LG's note
score 5