DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning

  • 2025-09-19 15:47:42
  • Sikai Bai, Haoxi Li, Jie Zhang, Zicong Hong, Song Guo
  • 0

Abstract

Despite the significant breakthrough of Mixture-of-Experts (MoE), theincreasing scale of these MoE models presents huge memory and storagechallenges. Existing MoE pruning methods, which involve reducing parameter sizewith a uniform sparsity across all layers, often lead to suboptimal outcomesand performance degradation due to varying expert redundancy in different MoElayers. To address this, we propose a non-uniform pruning strategy, dubbed\textbf{Di}fferentiable \textbf{E}xpert \textbf{P}runing (\textbf{DiEP}), whichadaptively adjusts pruning rates at the layer level while jointly learninginter-layer importance, effectively capturing the varying redundancy acrossdifferent MoE layers. By transforming the global discrete search space into acontinuous one, our method handles exponentially growing non-uniform expertcombinations, enabling adaptive gradient-based pruning. Extensive experimentson five advanced MoE models demonstrate the efficacy of our method acrossvarious NLP tasks. Notably, \textbf{DiEP} retains around 92\% of originalperformance on Mixtral 8$\times$7B with only half the experts, outperformingother pruning methods by up to 7.1\% on the challenging MMLU dataset.

 

Quick Read (beta)

loading the full paper ...