Megadose AI progress, ranked and analyzed.

PASTABench: Proactive Assessment of Sequential Trajectories for Agent Safety

· ArXiv · AI/CL/LG ·
The benchmark finds current LLMs still struggle to intervene at the right moment in risky agent workflows.

PASTABench covers 1,139 multi-turn trajectories across five risk categories and 13 subcategories. The paper frames safety monitoring around whether to intervene, when to intervene, and what risk is present. Its best-tested model hit only 40.74% optimal-timing interventions. The authors also report that some smaller models’ safety scores collapse when hazard keywords are neutralized, suggesting keyword sensitivity rather than deeper risk understanding. ArXiv · AI/CL/LG's note

score 5

Categories: Research