DeaMoE: Efficient MoE Structure for Fast Small-Batch Decoding
DeaMoE targets the weight-loading bottleneck that slows MoE models during small-batch decoding.
The paper groups experts into “departments,” letting experts in the same group share most parameters while keeping a small private portion. Its routing is designed in two stages to avoid redundant expert loading during decoding. The authors report up to 50.9% fewer loaded weights per step versus vanilla MoE, with up to 1.33x end-to-end TPOT speedup on a 7B model using A40 hardware. Microbenchmarks on DeepSeek-V3 show peak speedups near 2x on A40 and H100. ArXiv · AI/CL/LG's note
The paper groups experts into “departments,” letting experts in the same group share most parameters while keeping a small private portion. Its routing is designed in two stages to avoid redundant expert loading during decoding. The authors report up to 50.9% fewer loaded weights per step versus vanilla MoE, with up to 1.33x end-to-end TPOT speedup on a 7B model using A40 hardware. Microbenchmarks on DeepSeek-V3 show peak speedups near 2x on A40 and H100. ArXiv · AI/CL/LG's note
score 5