Megadose AI progress, ranked and analyzed.

ShieldVLA: Feasibility-Aware Safety Alignment for Vision-Language-Action Models

· HF Daily Papers ·
ShieldVLA uses a learned reachability-based safety critic to decide when a robot policy should pursue rewards and when it should recover from unsafe states.

The paper targets VLA models for robotic manipulation and navigation, where standard safety fine-tuning can still leave constraint violations or become too conservative. Its framework approximates a Hamilton-Jacobi reachability value function from visual observations, then uses that estimate to gate policy optimization. The authors also use rubric-based VLM safety scores to create structured critic targets without dense manual cost labels. Across five benchmarks and multiple VLA backbones, they report a 57% average reduction in cumulative safety cost and a +0.13 task-success gain over SafeVLA. HF Daily Papers' note

score 5

Categories: Research