Structuring MoE Expert Selection for Agentic Reinforcement Learning
The paper says MoE routing already tracks agent-like operations, and RL training can make use of that structure.
The authors find that off-the-shelf MoE models tend to route semantically similar agent turns, such as READ or UPDATE, through overlapping experts. They argue standard RL post-training leaves that routing uncontrolled, hurting both task success and inference efficiency. Their proposed framework aligns turn-level expert choices with agent operations, keeps token-level routing locally consistent, and uses entropy gating for stability. They report more than 10-point success-rate gains across all evaluated benchmarks. HF Daily Papers' note
The authors find that off-the-shelf MoE models tend to route semantically similar agent turns, such as READ or UPDATE, through overlapping experts. They argue standard RL post-training leaves that routing uncontrolled, hurting both task success and inference efficiency. Their proposed framework aligns turn-level expert choices with agent operations, keeps token-level routing locally consistent, and uses entropy gating for stability. They report more than 10-point success-rate gains across all evaluated benchmarks. HF Daily Papers' note
score 5