Megadose AI progress, ranked and analyzed.

NashDreamer: Model-Based Reinforcement Learning for Zero-Sum Imperfect-Information Games

· ArXiv · AI/CL/LG ·
NashDreamer uses a centralized world model to make model-based RL workable in two-player zero-sum imperfect-information games.

The paper argues decentralized model learning runs into identifiability barriers in these settings. Its MARSSM architecture separates environment dynamics from how player strategies shape individual observations. Across four benchmark games, the authors report better early sample efficiency than model-free baselines. They also flag posterior collapse in stochastic environments as an unresolved weakness for Dreamer-style methods. ArXiv · AI/CL/LG's note

score 4

Categories: Research