When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
The paper says physical danger and text-level unsafe content appear as separable signals inside model hidden states.
The authors test this across several Qwen2.5, Phi-3.5, and SmolLM2 models, then build PRISM as a single-layer probe over hidden states. PRISM reports 86.2-87.7% accuracy on SafeAgentBench with lower false positives than same-scale LLM judges. On the authors’ PSB-1K benchmark, built from 1,000 physical-risk pairs without direct harm keywords, PRISM reaches 99.6% accuracy and 0.7% false positives. ArXiv · AI/CL/LG's note
The authors test this across several Qwen2.5, Phi-3.5, and SmolLM2 models, then build PRISM as a single-layer probe over hidden states. PRISM reports 86.2-87.7% accuracy on SafeAgentBench with lower false positives than same-scale LLM judges. On the authors’ PSB-1K benchmark, built from 1,000 physical-risk pairs without direct harm keywords, PRISM reaches 99.6% accuracy and 0.7% false positives. ArXiv · AI/CL/LG's note
score 4