Search arXiv⌕ Search

arXiv · 2610.04289

Scaling of Wireless Foundation Models via Representation Diversity and Multi-Branch Architectures

Abstract

Wireless foundation models learn representations from unlabeled radio signals for reuse across downstream tasks. Scaling model capacity is a common strategy for learning richer representations and improving downstream performance. However, its gains are less consistent in wireless self-supervised learning when pretraining data are limited. We investigate objective diversity as an alternative scaling axis: different self-supervised objectives emphasize different signal properties, and combining their representations can preserve more information to enable diverse tasks. We develop a fusion framework that combines frozen representations from independent encoders trained through reconstructive, predictive, and contrastive learning. This provides a reference for the benefits of diversity, but requires the maintenance of multiple encoders. To retain these benefits within the parameter budget of a standard single-objective encoder, we introduce a jointly trained multi-branch architecture with a shared trunk and objective-specific branches. We pretrain on a heterogeneous corpus of spectrogram and channel state information data and evaluate six downstream tasks that span communication, sensing, and positioning. At an equal output embedding dimension, fusion outperforms the evaluated single-objective encoders on all six tasks while using roughly one-third of their parameters. The multi-branch architecture retains much of the fusion benefit within a single-encoder parameter budget. Representation analyses indicate complementary contributions across objectives, with much of the added benefit retained in components orthogonal to the reconstruction representation subspace. These findings support objective diversity as an effective strategy for scaling wireless foundation models.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ahmed Mohamed, Ahmed Aboulfotouh, Hatem Abou-Zeid. 2026-10-03. Scaling of Wireless Foundation Models via Representation Diversity and Multi-Branch Architectures. https://arxiv.org/abs/2610.04289

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

LatentWave: JEPA Pretraining for Wireless Foundation Models

Wireless foundation models have emerged as a promising alternative to building separate models for each wireless task. However, existing approaches rely on masked input reconstruction, which can bias representations toward low-level signal details. In this paper, we propose LatentWave, a wireless foundation model pretrained using a Joint-Embedding Predictive Architecture (JEPA) on diverse wireless spectrograms and channel state information (CSI). By predicting masked regions in latent space, LatentWave learns representations that are more transferable out of the box across diverse downstream tasks. The proposed architecture employs per-channel patch embeddings with stochastic channel sampling during pretraining, allowing it to process variable antenna counts and improving usability across heterogeneous wireless configurations. We evaluate LatentWave on four downstream tasks: RF signal classification, 5G NR positioning, beam prediction, and LoS/NLoS classification, comparing against a masked-modeling baseline (WavesFM) pretrained on the same data. Additionally, we show that the masking geometry introduces a task-dependent inductive bias: frequency masking strongly favors channel-related tasks such as positioning and beam prediction, while region masking better preserves discriminability for signal classification.

eess.SP↗

Pinching-Antenna Differential Spatial Modulation with RSSI-Assisted Selection and Low-Complexity Detection: Methods and Performance Analysis

Pinching antenna (PA) systems provide macroscopic spatial reconfigurability by enabling signal radiation from favorable locations along a dielectric waveguide. Exploiting this flexibility, however, typically requires channel knowledge for PA configuration, which can introduce substantial overhead in dynamic propagation environments. This paper proposes PA-assisted differential spatial modulation (PA-DSM), enabling data transfer with spatial index selection and non-coherent data detection without instantaneous channel state information (CSI) at the receiver. To exploit PA reconfigurability while preserving this CSI-efficient operation, we introduce a slow-timescale received signal strength indicator (RSSI)-assisted selection strategy that identifies a favorable subset of PAs from scalar power measurements. We further specialize a two-stage low-complexity maximum-likelihood (LC-ML) detector to phase shift keying (PSK) signaling, obtaining closed-form symbol decisions that remove the joint modulation-combination search while preserving the minimizer of exhaustive differential detection. Furthermore, a moment generating function (MGF)-based pairwise error analysis and the corresponding bit error probability (BEP) bound are developed under normalized common-K Rician fading, revealing how the difference rank and spectrum, number of receive antennas, and previous-state-dependent line-of-sight (LoS) alignment affect the error performance. Numerical results across representative indoor-office, indoor-factory, and street-canyon sixth-generation (6G) Frequency Range 3 (FR3) scenarios validate the analytical trends and quantify when PA-DSM outperforms coherent benchmark schemes whose channel estimates age between pilot transmissions.

eess.SP↗

Quantized Transformers for Massive MIMO Precoding with Automatic Resolution Tuning

Deep learning precoders, particularly those based on Transformer architectures, offer superior spectral efficiency for Massive MIMO systems but may incur high computational cost. While existing mixed-precision quantization studies on convolutional neural networks establish important energy baselines, traditional discrete search methods fail to scale to the large parameter spaces of modern Transformers. To bridge this gap, we apply a differentiable precision learning technique to massive MIMO precoding. This approach autonomously and jointly optimizes weights, activations, quantization step sizes, and layer-wise bit-widths in a single training loop, directly optimizing for energy efficiency. We further expand this approach by incorporating Neural Architecture Search across various Transformer model sizes and evaluating the impact of random versus floating-point trained initializations on resolution tuning. Ultimately, we demonstrate that the proposed training method enables the deployment of highly compact Transformer precoders, improving energy efficiency by up to 288\(\times\) compared to the classical Weighted Minimum Mean Square Error algorithm at equal sum-rate performance in a dense downtown environment.

eess.SP↗