Higher-order pruning of experts in mixture-of-experts language models
HOPE prunes MoE models by accounting for how experts work together, not just what each expert contributes alone.
The paper derives a second-order pruning objective meant to minimize an upper bound on pruning error. It frames REAP as the same approach with the expert-interaction terms removed. Across three MoE models up to 122B parameters, HOPE beats tested baselines most clearly at high pruning rates and on agentic workloads. At 50% pruning, it reports the best average rank and gains up to 6.1% on agentic coding.
ArXiv · AI/CL/LG's note
The paper derives a second-order pruning objective meant to minimize an upper bound on pruning error. It frames REAP as the same approach with the expert-interaction terms removed. Across three MoE models up to 122B parameters, HOPE beats tested baselines most clearly at high pruning rates and on agentic workloads. At 50% pruning, it reports the best average rank and gains up to 6.1% on agentic coding.
ArXiv · AI/CL/LG's note
score 6