A Unified Bellman Operator for Safety-Critical Reinforcement Learning
The paper proposes a Bellman operator meant to learn task reward and hard safety constraints in one value function.
It targets safe reinforcement learning settings where methods often split between strict safety with prior knowledge and joint learning with only average safety. The authors describe a two-timescale temporal-difference setup: safety is estimated faster, while the joint value is learned more slowly. Their convergence argument uses occupation-averaged differential inclusions and ends at optimal safety-constrained task value functions. In continuous-control tests with neural approximations, they report stable convergence and near-zero test-time safety violations. ArXiv · AI/CL/LG's note
It targets safe reinforcement learning settings where methods often split between strict safety with prior knowledge and joint learning with only average safety. The authors describe a two-timescale temporal-difference setup: safety is estimated faster, while the joint value is learned more slowly. Their convergence argument uses occupation-averaged differential inclusions and ends at optimal safety-constrained task value functions. In continuous-control tests with neural approximations, they report stable convergence and near-zero test-time safety violations. ArXiv · AI/CL/LG's note
score 4