Megadose Built for builders and researchers.

Learning in Dreams, Winning in Reality: A Continuous Dyna Loop for a Ten-Hero MOBA

· ArXiv · AI/CL/LG ·
A policy trained only in a learned MOBA simulator reached a 70.2% real-game win rate as radiant, but failed entirely as dire.

The paper tests a continuous asynchronous Dyna loop on a full ten-hero MOBA, using real games only to feed the world model and evaluate policies. Dream training alone won 0%, and the pre-loop policy was at 33.7%, before the loop reached 421 wins in 600 held-out radiant games. The author says internal dream metrics did not reveal model exploitation, while real-game anchoring exposed collapses. The resulting agent won through macro pressure, not coordinated fighting, reflecting what the world model captured and missed. ArXiv · AI/CL/LG's note

score 5

Categories: Research