CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents
CERA-MoA trains the router and the LLM agents together instead of treating routing as fixed around changing agents.
The paper proposes an iterative reinforcement-learning setup where agent policies and a dynamic router co-evolve. Its router estimates each agent’s semantic familiarity from mid-layer hidden states, avoiding full rollouts. It then activates a minimal agent subset through cumulative-threshold routing to balance performance and efficiency. The authors report gains over static-agent routing and fixed-workflow fine-tuning baselines across multiple domains. ArXiv · AI/CL/LG's note
The paper proposes an iterative reinforcement-learning setup where agent policies and a dynamic router co-evolve. Its router estimates each agent’s semantic familiarity from mid-layer hidden states, avoiding full rollouts. It then activates a minimal agent subset through cumulative-threshold routing to balance performance and efficiency. The authors report gains over static-agent routing and fixed-workflow fine-tuning baselines across multiple domains. ArXiv · AI/CL/LG's note
score 4