arXiv · 2609.37751
Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models
Abstract
Sparse mixture-of-experts (MoE) large language models scale model capacity by routing each token to a small subset of experts. Their routers are regularized with load balancing terms and learn affinity scores through the language-model objective. However, these objectives do not provide direct alignment between routing affinities and token-level error. We introduce token-error supervision for sparse routing in two forms. The first form predicts an error score per expert. The affinity-weighted aggregate of these scores is aligned to the next-token cross-entropy loss, while the individual scores attenuate affinity before top-$K$ selection. The second directly aligns the router's affinities to the model's objective without requiring an additional head or inference-time modification. Both formulations use the Itakura--Saito divergence or an exponential negative log-likelihood for aligning affinities and token errors. Across two sparse MoE backbones and four multiple-choice question-answering benchmarks, we evaluate both supervision mechanisms. On Granite, our method improves accuracy by approximately 2.3 percentage points on average over a parameter-matched routing baseline. With stronger supervision, the gain on ARC-Challenge reaches 2.94 points. Both mechanisms preserve the native sparse execution budget and aggregation policy. Our code is available in the supplementary materials.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yury Nahshan, Nati Daniel, Jacob Goldberger, Yoli Shavit. 2026-09-29. Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models. https://arxiv.org/abs/2609.37751
Cite the original work for its findings. Save a collection to share your selection of sources.