Safe Meta-Reinforcement Learning via Information Space Reachability
The paper puts the safety check inside the agent’s belief state, not just its physical state.
Li, Tang, and Azizan propose a meta-RL framework that evaluates whether an adapting agent can avoid unsafe regions indefinitely. Their “information space” combines the environment state with the agent’s belief about the task it is facing. They define a safety value function, show it has Bellman-style structure, and use it for safety filtering and constrained policy optimization. The authors report effectiveness on meta-RL benchmarks. ArXiv · AI/CL/LG's note
Li, Tang, and Azizan propose a meta-RL framework that evaluates whether an adapting agent can avoid unsafe regions indefinitely. Their “information space” combines the environment state with the agent’s belief about the task it is facing. They define a safety value function, show it has Bellman-style structure, and use it for safety filtering and constrained policy optimization. The authors report effectiveness on meta-RL benchmarks. ArXiv · AI/CL/LG's note
score 4