Redistribution-based Cost Inference Improves Sparse Safe Offline RL
The paper turns one-time “stop” feedback into per-step safety costs for offline RL training.
RCI treats sparse unsafe-stop labels as a temporal credit assignment problem, redistributing trajectory-level feedback into dense costs. The authors argue the redistribution is theoretically lossless for the constrained decision problem while making cost-critic learning easier. In highway driving and robotic manipulation tests, it lowers violation rates versus sparse-label and classifier-based baselines, including under mixed datasets and label noise. ArXiv · AI/CL/LG's note
RCI treats sparse unsafe-stop labels as a temporal credit assignment problem, redistributing trajectory-level feedback into dense costs. The authors argue the redistribution is theoretically lossless for the constrained decision problem while making cost-critic learning easier. In highway driving and robotic manipulation tests, it lowers violation rates versus sparse-label and classifier-based baselines, including under mixed datasets and label noise. ArXiv · AI/CL/LG's note
score 4