Controlling Collectives of AI Agents in Reasoning Space with Spatial Transformers
COMPASS reports natural-language control of robot flocks up to 1,024 agents, far beyond its training scale.
The paper proposes a decentralized architecture where each robot uses a spatial transformer to turn multi-hop fleet messages into a learned feedback token. In the authors’ experiments, that setup outperforms a centralized frontier LLM policy and a language-only communication ablation on cohesive flocking formations. They also report that structured variation in the input command can offset model biases across larger teams. Hand-engineered feedback with raw state in the language channel, by contrast, is said to destroy cohesion. Source: ArXiv · AI/CL/LG's note.
The paper proposes a decentralized architecture where each robot uses a spatial transformer to turn multi-hop fleet messages into a learned feedback token. In the authors’ experiments, that setup outperforms a centralized frontier LLM policy and a language-only communication ablation on cohesive flocking formations. They also report that structured variation in the input command can offset model biases across larger teams. Hand-engineered feedback with raw state in the language channel, by contrast, is said to destroy cohesion. Source: ArXiv · AI/CL/LG's note.
score 5