Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding
DMM replaces one-shot decentralized action sampling with repeated intent refinement before agents commit to moves.
The paper argues that valid per-agent choices can still combine into bad joint actions when sampled independently. Its method starts agents with random action intents, then refines them through local communication rounds, borrowing the denoising idea from diffusion models. DMM is pretrained on expert MAPF solutions and further tuned with MICPO, a critic-free group-relative reinforcement method. In the reported tests, it solves 1,598 of 1,600 MovingAI tasks and scales to more than one million simultaneously acting agents in obstacle-rich settings. HF Daily Papers' note
The paper argues that valid per-agent choices can still combine into bad joint actions when sampled independently. Its method starts agents with random action intents, then refines them through local communication rounds, borrowing the denoising idea from diffusion models. DMM is pretrained on expert MAPF solutions and further tuned with MICPO, a critic-free group-relative reinforcement method. In the reported tests, it solves 1,598 of 1,600 MovingAI tasks and scales to more than one million simultaneously acting agents in obstacle-rich settings. HF Daily Papers' note
score 4