BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models
BRANCH-MoE replaces flat expert routing with a balanced binary-tree router for large embedding models.
The paper puts experts at tree leaves and uses an exponential moving average at each internal node to keep traffic from collapsing into one subtree. It argues this can balance expert use without an auxiliary load-balancing loss, while also giving the experts a topology that flat routers lack. The authors test it against several routing baselines on four tabular prediction datasets with 16 experts and top-4 routing. They report comparable task quality, balanced utilization, localized co-activation, and reduced communication when experts are placed by tree prefix. ArXiv · AI/CL/LG's note
The paper puts experts at tree leaves and uses an exponential moving average at each internal node to keep traffic from collapsing into one subtree. It argues this can balance expert use without an auxiliary load-balancing loss, while also giving the experts a topology that flat routers lack. The authors test it against several routing baselines on four tabular prediction datasets with 16 experts and top-4 routing. They report comparable task quality, balanced utilization, localized co-activation, and reduced communication when experts are placed by tree prefix. ArXiv · AI/CL/LG's note
score 5