arXiv · 2610.09342
Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression
Abstract
Mixture-of-Experts (MoE) large language models decouple capacity from compute through sparse routing, but their large parameter count creates storage and serving challenges. We analyze three MoE compression families: expert pruning, expert merging, and weight reconstruction, and derive structural error bounds showing that pruning and merging can incur non-vanishing errors tied to routing and expert heterogeneity. In contrast, weight reconstruction avoids these structural costs by preserving expert structure and routing. Motivated by the analysis, we propose Shared Low-rank Basis Factorization (SLBF), a data-free weight reconstruction method that uses rank-$k$ bases shared among experts, enabling richer cross-expert sharing, faster convergence, and lower reconstruction error. A post-hoc gauge fixing removes redundant parameters at no representational cost. Across five MoE architectures spanning 16B to 122B parameters, SLBF consistently outperforms methods from all three compression families.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Tianxiao Cao, Jiahe Shao, Yuning Qiu, Kyohei Atarashi, Hisashi Kashima, Qibin Zhao. 2026-10-07. Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression. https://arxiv.org/abs/2610.09342
Cite the original work for its findings. Save a collection to share your selection of sources.