Search arXiv⌕ Search

arXiv · 2609.31177

Floating-Point Microformat Quantization and Pruning for Efficient MU-MIMO Neural Receivers

Abstract

Neural receivers outperform conventional 5G NR processing chains, but their compute and memory demands hinder real-time deployment. For a standard-compliant multi-user MIMO neural receiver, the 4-bit number format, not merely the bit width, determines whether compression preserves the gain over classical receivers. We apply weight and activation quantization-aware training (QAT) and, separately, 50% magnitude pruning, comparing INT8/INT4 and FP8 (E4M3)/FP4 (E2M1) weights with INT8 post-ReLU activations. Trained on 3GPP UMi channels and evaluated on TDL-B and TDL-C, 8-bit weight-activation models remain within 0.05 dB of FP32 at 10% and 1% block error rate (BLER). At 4 bits, uniform INT4 loses 3.3-3.7 dB and falls below LS-LMMSE, whereas FP4 more than halves this loss (1.3-1.4 dB) and still outperforms it by about 0.5 dB, even after pruning. FP4's denser near-zero grid matches the trained weight distribution, and FP4 avoids the residual-path over-pruning seen with INT4. An analytic cost model projects 66x fewer bit-operations and 8.8x less weight storage for pruned 4-bit-weight inference.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

SaiKrishna Saketh Yellapragada, Esa Ollila, Mário Costa, Yawei Li. 2026-09-25. Floating-Point Microformat Quantization and Pruning for Efficient MU-MIMO Neural Receivers. https://arxiv.org/abs/2609.31177

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Secure Over-the-Air Computation Against Multiple Eavesdroppers using Correlated Artificial Noise

Over-the-air (OtA) computation enables scalable analog aggregation by exploiting the superposition property of wireless channels, making it an attractive joint communication and computation paradigm for distributed sensing and learning. However, the uncoded nature of analog transmission exposes the computation result to eavesdropping, and the fundamental security limits of OtA computation against multiple cooperating adversaries remain poorly understood. In this paper, we develop an estimation-theoretic framework for analyzing the security of analog OtA computation in the presence of multiple distributed eavesdroppers that may jointly process their observations. We first derive the optimal estimator for cooperating eavesdroppers and bounds on the achievable estimation accuracy of both the legitimate receiver and the adversaries. Our analysis reveals a key insight: while random channel phase misalignment provides significant inherent MSE-security against individual eavesdroppers, this protection largely disappears once multiple eavesdroppers cooperate. Motivated by this observation, we propose a correlated artificial noise design based on zero-forcing that preserves the aggregation accuracy at the legitimate receiver while maximizing the estimation error at the cooperative eavesdroppers. Numerical results demonstrate that the proposed design substantially reduces the security advantage gained through eavesdropper cooperation and achieves security close to uncorrelated artificial-noise schemes without sacrificing computation accuracy. These results provide both a theoretical characterization of the security limits of analog OtA computation and a practical design guideline for its secure deployment in real-world wireless systems.

eess.SP↗

Distributed Riemannian Optimization in Geodesically Non-convex Environments

This paper studies the problem of distributed Riemannian optimization over a network of agents whose cost functions are geodesically smooth but possibly geodesically non-convex. Extending a well-known distributed optimization strategy called diffusion adaptation to Riemannian manifolds, we show that the resulting algorithm, the Riemannian diffusion adaptation, provably exhibits several desirable behaviors when minimizing a sum of geodesically smooth non-convex functions over manifolds of bounded curvature. More specifically, we establish that the algorithm can approximately achieve network agreement in the sense that Fréchet variance of the iterates among the agents is small. Moreover, the algorithm is guaranteed to converge to a neighborhood of a first-order stationary point for general geodesically non-convex cost functions. When the global cost function additionally satisfies the local Riemannian Polyak-Lojasiewicz (PL) condition, we also show that it converges linearly under a constant step size up to a steady-state error. Finally, we apply this algorithm to decentralized robust principal component analysis (PCA) and low-rank matrix completion problems and illustrate its convergence and performance through numerical simulations.

eess.SP↗

Integrated sensing and communications in the 3GPP New Radio: sensing limits

Integrated Sensing and Communications (ISAC) is regarded as a key element of the beyond-fifth-generation (5G) and sixth-generation (6G) systems, raising the question of whether current 5G New Radio (NR) signal structures can meet the sensing accuracy requirements specified by the Third Generation Partnership Project (3GPP). This paper addresses this issue by analyzing the fundamental limits of range and velocity estimation through the Cramér-Rao lower bound (CRLB) for a monostatic unmanned aerial vehicle (UAV) sensing use case currently under consideration in the 3GPP standardization process. The study focuses on standardized signals and also evaluates the potential performance gains achievable with reference signals specifically designed for sensing purposes. The compact CRLB expressions derived in this work highlight the fundamental trade-offs between estimation accuracy and system parameters. The results further indicate that information from multiple slots must be exploited in the estimation process to attain the performance targets defined by the 3GPP. As a result, the 5G NR positioning reference signal (PRS), whose patterns may be suboptimal for velocity estimation when using single-slot resources, becomes suitable when multislot estimation is employed. Finally, we propose a two-step iterative range and radial-velocity estimator that attains the CRLB over a significantly wider range of distances than conventional maximum-likelihood (ML) estimators, for which the well-known threshold effect severely limits the distance range over which the accuracy requirements imposed by the 3GPP are satisfied.

eess.SP↗