Megadose AI progress, ranked and analyzed.

Redistribution-based Cost Inference Improves Sparse Safe Offline RL

· ArXiv · AI/CL/LG ·
The paper turns one-time “stop” feedback into per-step safety costs for offline RL training.

RCI treats sparse unsafe-stop labels as a temporal credit assignment problem, redistributing trajectory-level feedback into dense costs. The authors argue the redistribution is theoretically lossless for the constrained decision problem while making cost-critic learning easier. In highway driving and robotic manipulation tests, it lowers violation rates versus sparse-label and classifier-based baselines, including under mixed datasets and label noise. ArXiv · AI/CL/LG's note

score 4

Categories: Research