arXiv · 2609.32581
HERO-MoE: Historical Expert Routing with Scale-Preserving Fusion
Abstract
Mixture-of-Experts (MoE) architectures have become a standard way to scale model capacity while keeping computation sparse, yet routing remains a key determinant of MoE quality and training behavior. Prior empirical studies suggest that MoE routing reflects input semantics and upstream computation across depth, but standard routers do not explicitly use the routing distributions produced by preceding layers. We propose HERO-MoE, Historical Expert ROuting with Scale-Preserving Fusion, a routing framework that injects historical routing priors into MoE routers by reusing detached, dense routing distributions collected from preceding MoE layers. The key idea is simple: HERO-MoE preserves the original token-conditioned routing branch and adds a residual historical routing contribution before the standard softmax and top-$k$ dispatch. To stabilize this historical signal, HERO-MoE introduces a scale-preserving fusion mechanism that matches the magnitude of historical routing memory to the current hidden representation and accounts for the number of visible historical layers, without introducing an auxiliary routing loss or a fusion-specific tuning parameter. By reusing routing distributions already computed by preceding MoE layers, HERO-MoE improves training-loss reduction with modest end-to-end overhead. The resulting router remains compatible with standard sparse dispatch, including top-$k$ and group-limited routing, and can be inserted into existing MoE backbones with minimal architectural changes. Experiments on an approximately 8B-parameter MoE model with 0.5B active parameters, trained from scratch on 100B tokens, show that HERO-MoE reduces the final loss from 1.6393 to 1.6184, while peak memory and FLOPs increase by only 0.44\% and 0.64\%, respectively.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Junxiang Qiu, Zhengsu Chen, Xinting Hu, Shuo Wang, Hengheng Zhang, Shaofeng Zhang, Changcheng Li, Boyu Shi, Qi Tian. 2026-09-26. HERO-MoE: Historical Expert Routing with Scale-Preserving Fusion. https://arxiv.org/abs/2609.32581
Cite the original work for its findings. Save a collection to share your selection of sources.