Search arXiv⌕ Search

arXiv · 2507.01728

Token Communication in the Era of Large Models: An Information Bottleneck-Based Approach

Abstract

This letter proposes UniToCom, a unified token communication paradigm that treats tokens as the fundamental units for both processing and wireless transmission. Specifically, to enable efficient token representations, we propose a generative information bottleneck (GenIB) principle, which facilitates the learning of tokens that preserve essential information while supporting reliable generation across multiple modalities. By doing this, GenIB-based tokenization is conducive to improving the communication efficiency and reducing computational complexity. Additionally, we develop $σ$-GenIB to address the challenges of variance collapse in autoregressive modeling, maintaining representational diversity and stability. Moreover, we employ a causal Transformer-based multimodal large language model (MLLM) at the receiver to unify the processing of both discrete and continuous tokens under the next-token prediction paradigm. Simulation results validate the effectiveness and superiority of the proposed UniToCom compared to baselines under dynamic channel conditions. By integrating token processing with MLLMs, UniToCom enables scalable and generalizable communication in favor of multimodal understanding and generation, providing a potential solution for next-generation intelligent communications.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Hao Wei, Wanli Ni, Wen Wang, Wenjun Xu, Dusit Niyato, Ping Zhang. 2025-07-02. Token Communication in the Era of Large Models: An Information Bottleneck-Based Approach. https://arxiv.org/abs/2507.01728

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Low-Interference N-Continuous OFDM via Optimized Time-Domain Smoothing

A novel basis signal optimization method is proposed for reducing the interference in the N-continuous orthogonal frequency division multiplexing (NC-OFDM) system. Compared to conventional NC-OFDM, the proposed scheme is capable of improving the transmission performance while maintaining an identical sidelobe suppression performance imposed by the linear combination of two groups of basis signals. Our performance results demonstrate that with a low-complexity overhead, the proposed scheme is capable of striking a better trade-off among the bit error rate (BER), complexity, and the sidelobe suppression performance compared to its conventional counterparts.

eess.SP↗

Few-Shot Specific Emitter Identification via Integrated Complex Variational Mode Decomposition and Spatial Attention Transfer

Specific emitter identification (SEI) utilizes passive hardware characteristics to authenticate transmitters, providing a robust physical-layer security solution. However, most deep-learning-based methods rely on extensive data or require prior information, which poses challenges in real-world scenarios with limited labeled data. We propose an integrated complex variational mode decomposition algorithm that decomposes and reconstructs complex-valued signals to approximate the original transmitted signals, thereby enabling more accurate feature extraction. We further utilize a temporal convolutional network to effectively model the sequential signal characteristics, and introduce a spatial attention mechanism to adaptively weight informative signal segments, significantly enhancing identification performance. Additionally, the branch network allows leveraging pre-trained weights from other data while reducing the need for auxiliary datasets. Ablation experiments on the simulated data demonstrate the effectiveness of each component of the model. An accuracy comparison on a public dataset reveals that our method achieves 96% accuracy using only 10 symbols without requiring any prior knowledge.

eess.SP↗

A Foundation Model for Large-Scale Wireless Network Planning , Operation and Optimization

Wireless cellular networks form the connective tissue of human society, sustained by a continuous physical dialogue between engineered infrastructure and its surroundings. Radio signals emitted from base stations traverse terrain, diffract around buildings and scatter through streets before reaching billions of users. Together, these interactions produce the city-wide radio environment on which every network decision rests. Shaping this environment through deployment and optimization determines the connectivity societies rely on, yet learning it effectively at city scale and generalizing across diverse cities and deployments remain open challenges. Here we answer positively by introducing ChaRT, a foundation model that learns transferable radio representations from measurement reports generated by deployed cellular networks. These reports provide abundant multi-cell, multi-beam observations without dedicated campaigns, forming a scalable data foundation for city-scale learning. ChaRT embeds beam-level angular structure, network hierarchy and propagation-regime diversity in its architecture, and is pretrained through context-aware masked beam modelling and self-distillation with channel-model-constrained augmentation. We pretrain ChaRT on over one billion reports comprising 18.2 billion beam-level observations from 3,503 cells in one city. With a single set of weights, ChaRT reconstructs radio environments in unseen cities and transfers to radio map construction, new-site prediction and network parameter tuning. With only 1% of labelled data, it supports user localization, beam prediction, propagation scenario classification and estimation of the signal-to-interference-plus-noise ratio. The learned representation further enables beamspace clustering for reusable radio-grid construction. These results establish ChaRT as a transferable foundation for network-wide intelligence.

eess.SP↗