Search arXiv⌕ Search

arXiv · 2610.03741

Logit-Aware MIMO AirComp for Distributed Mixture-of-Experts LLM Inference over Wireless Edge Networks

Abstract

Distributed mixture-of-experts (MoE) inference is a promising architecture for deploying large language models (LLMs) at wireless edge networks because sparse experts can be placed across coordinated base stations (BSs), while the anchor node and user equipment (UE) can offload LLM inference tasks to BSs. The communication bottleneck is the MoE aggregation, where the anchor BS must recover a weighted sum of selected expert outputs before each decoding step. Over-the-air computation (AirComp) is well matched to this operation because the wireless multiple-access channel naturally superposes simultaneous transmissions. However, conventional AirComp minimizes communication distortion, whereas MoE aggregation errors have unequal impact on LLM outputs. We propose a logit-aware MIMO AirComp framework that estimates local logit sensitivity as block-level weights and jointly optimizes receive combiners and BS precoders under per-BS power constraints. We also develop an alternating algorithm that protects decoding decisions from aggregation perturbations. Using OpenCompass, we evaluate Qwen3-30B-A3B-Instruct-2507-FP8 on GSM8K and ARC-Challenge. At 30 dB aggregation SNR, perturbed Qwen3 retains 99.3% of clean GSM8K accuracy and 94.8% of clean ARC-Challenge accuracy. In a direct ARC-Challenge closed-loop audit over 295 examples, SW-AirComp achieves 43.73% accuracy at 30 dB and 26.78% at 20 dB, outperforming unweighted, matched-filter, and zero-forcing AirComp. In the wireless simulator, SW-AirComp reduces decision-relevant aggregation distortion; at 20 dB, its final weighted-sum mean-squared error is 53.5% lower than unweighted AirComp under the same channels, samples, and sensitivity weights. Logit-RMSE and logit-gap diagnostics further show that the gain comes from reducing output-logit perturbation and lowering the risk of top-token changes.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Lyutianyang Zhang, Yunjian Jia, Liu Cao, Dengke Wang, Jinke Ren, Shuguang Cui. 2026-09-22. Logit-Aware MIMO AirComp for Distributed Mixture-of-Experts LLM Inference over Wireless Edge Networks. https://arxiv.org/abs/2610.03741

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Parametric and structure-aware information theory for multi-scale analysis of composition

Compositional data is common across the natural and social sciences, requiring methods to measure the diversity of compositions and the heterogeneity of collections of them. Via Bregman geometry, a strictly concave diversity index induces a measure of heterogeneity that decomposes across scales. A canonical example is Shannon entropy inducing mutual information. We develop this framework by incorporating category similarity into $α$-logarithmic entropy, a parametric diversity index. We disprove an existing general concavity result and prove that, for $α=3$, the region of strict concavity is always convex, complementing previous results for $α=1$ and $2$. The resulting geometry provides methods to compare compositions substantially faster than optimal transport. Applied to occupation compositions across England and Wales, they reveal distinct regionalisations based on different notions of occupation relatedness. Applied to ecological data, they recover established patterns of functional and taxonomic $β$-diversity, and reveal sensitivity to the diversity parameter.

cs.IT↗

Improved Characterization of the Memory-Rate Tradeoff for Demand-Private Coded Caching With Multiple Demands

We consider a coded caching problem with multiple demands under a privacy constraint. In this problem, a server with access to $N$ files serves $K$ users over a shared link, and each user requests $L$ distinct files. The privacy constraint requires that each user obtain no information about the demands of the other users. We first propose a new achievable scheme for arbitrary $N$, $K$, and $L$ by applying a novel transformation to a suitable non-private coded caching scheme. We then derive a new converse bound, and show that the proposed scheme is order optimal within a multiplicative factor of $6$ of this bound. Finally, for the case of two users, we completely characterize the optimal memory-rate tradeoff for arbitrary $N$ and $L$ through three new achievable schemes and matching converse bounds.

cs.IT↗

Learning LDPC codes with density evolution over relaxed protographs

We consider the design of low-density parity-check (LDPC) codes for a given iterative decoder. While LDPC performance can be evaluated using simulation, density evolution (DE), or EXIT-chart analysis, selecting a parity-check matrix (PCM) remains a difficult combinatorial optimization problem. Existing approaches often rely on population-based search, random mutations, or genetic algorithms, which require careful tuning and incur high computational cost. Recent gradient descent (GD)-based methods optimize relaxed PCMs by differentiating through decoder simulations, but rely on noisy Monte Carlo estimates, line searches over soft matrix representations, and remain costly for long codes. Moreover, the loss is typically evaluated only at integer-valued PCMs. We focus on long protograph-based LDPC codes and propose a deterministic GD-based framework operating directly on a relaxed protograph representation. The loss is based on DE bit error rate (BER) and can be evaluated directly for relaxed protographs. To justify this relaxation, we associate the relaxed representation with an ensemble of binary PCMs and show that the proposed relaxed DE yields the ensemble-averaged DE performance. The resulting procedure supports standard GD optimization and achieves fast, reliable convergence through deterministic DE evaluation and informative gradients. Numerical results show that the optimized protographs outperform 5G LDPC codes with matching dimensions.

cs.IT↗