UniPool: Learning Expert-to-Layer Ownership from Brief Global Access
Most Mixture-of-Experts (MoE) transformers preassign each expert to one layer for the whole of training. We study whether a brief global-access phase can learn a compatible expert-to-layer allocation that is then executed privately. UniPool-Lock gives every layer its own router over one global pool, balances aggregate pool usage, and scores experts with a scale-stable NormRouter. After the short ownership-learning phase, it assigns each expert to one layer, locks the disjoint allocation, and continues training as a layer-private MoE. With the same expert-FFN budget and routed expert FLOPs, this recipe lowers held-out loss relative to vanilla MoE by 0.024-0.037 across five dense-equivalent scales from 182M to 1.5B, including -0.0247 at 1.5B after 60B tokens. The ownership-learning phase lasts 2K steps, about 3.3% of training, and post-lock step time is within -0.7% to +2.8% of vanilla MoE in our throughput measurements. Controls locate the gain in full-pool training and in the allocation it produces: a random disjoint allocation fixed at initialization, trained with the same router and losses, matches vanilla, whereas locking the allocation learned during the full-pool phase recovers nearly all of the persistent full-pool gain, and substituting a random allocation at the lock forfeits about half of it. Keeping full-pool access throughout training (UniPool-Full) also improves over vanilla MoE at four scales and outperforms it with only 66.7% (182M) to 50% (469M and 650M) of its expert parameters.