Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models
FailBank turns runtime safety corrections into training targets, improving VLA policies instead of only blocking bad actions.
The paper proposes a four-stage self-evolving framework for robotic manipulation models, using an observe-only CBF safety module to generate counterfactual corrections while the policy still acts. It stores useful failures as corrective targets and keeps successful uncorrected actions as anchors for guarded LoRA updates. On VLA-Arena, FailBank improved success rates by 8.5 and 6.9 points over base policies while cutting policy-induced cumulative cost by 35.6% and 23.8%. Against runtime shielding, it raised success by 25.4 and 9.5 points while keeping costs comparable. HF Daily Papers' note
The paper proposes a four-stage self-evolving framework for robotic manipulation models, using an observe-only CBF safety module to generate counterfactual corrections while the policy still acts. It stores useful failures as corrective targets and keeps successful uncorrected actions as anchors for guarded LoRA updates. On VLA-Arena, FailBank improved success rates by 8.5 and 6.9 points over base policies while cutting policy-induced cumulative cost by 35.6% and 23.8%. Against runtime shielding, it raised success by 25.4 and 9.5 points while keeping costs comparable. HF Daily Papers' note
score 5