Search arXivSearch

arXiv subjects

Yulin Shao

Publications and source records attributed to Yulin Shao.

At least 19 recordsLinked to original sources

Networked Embodied Communication: From Collective Distinguishability to Communication Reliability

Embodied agents need to convey information to surrounding infrastructure, but their active communication interfaces may be unavailable, constrained, or intentionally inactive. Their ability to manipulate physical states offers a complementary path: messages can be encoded in deliberately selected configurations and recovered through infrastructure sensing. This principle underlies embodied communication. Yet physical differences do not guarantee distinguishable messages: a single sensing viewpoint may leave ambiguities that repeated sensing cannot resolve. This paper develops networked embodied communication, where distributed access points (APs) jointly observe message-bearing scatterer positions under fixed illumination. Under a correlated Gaussian sensing model, we characterize the additional distinguishability supplied by receive APs, establish exact redundancy conditions, and reveal how distinctions absent from individual observations can emerge through cross-AP statistical relationships. We then establish the exact asymptotic optimal maximum-error behavior of a finite alphabet under repeated independent sensing. The largest group of indistinguishable messages determines the error floor; once all messages are distinguishable, the minimum pairwise Chernoff information determines the error exponent. For a given alphabet, receiver cooperation can therefore eliminate an error floor that repetition at any individual AP cannot overcome. Building on these results, we derive finite-budget reliability conditions and jointly design the receive AP set and message-bearing positions. Numerical results show that the proposed search closely approaches exact benchmarks on reduced instances with substantially fewer candidate evaluations than exhaustive enumeration, while receiver cooperation reduces the sensing intervals needed to guarantee reliable decoding.

eess.SP

Resolving the Discontinuity of Continuous-Time AFDM Waveforms

Continuous-time affine frequency division multiplexing (AFDM) waveforms, constructed via frequency wrapping and phase correction, are known to be sample-wise equivalent to the widely adopted discrete AFDM framework. In this paper, we uncover a fundamental and previously overlooked flaw in this construction: its complex envelope is inherently discontinuous for generic chirp parameters. We show that these discontinuities are the direct cause of the high out-of-band emission (OOBE). To resolve this issue, we propose a fundamentally different continuous-time waveform, termed stepped frequency division multiplexing (SFDM). Unlike conventional approaches that allow continuous frequency variation, SFDM freezes the instantaneous frequency at the midpoint of the underlying chirp trajectory within each Nyquist sampling interval. This design yields a complex envelope that is strictly continuous over the entire symbol duration while preserving exact sample-wise equivalence with discrete AFDM. A unified spectral analysis reveals that the superior OOBE performance of SFDM stems from the absence of internal jump discontinuities, which otherwise dominate the far-out spectral roll-off. Numerical results confirm that SFDM consistently achieves significantly lower OOBE across a wide range of chirp rates.

eess.SP

Sensing-Induced Embodied Communication in the Near Field

Integrated sensing and communication is turning the cellular infrastructure into an active observer of the physical world. When such infrastructure interacts with embodied agents capable of deliberately changing their states and surroundings, the physical world itself can become a communication medium. This paper studies the fundamental communication limits of this sensing-induced embodied communication paradigm in the near field. We consider an agent that maps messages to the positions of a controllable scatterer within a bounded three-dimensional (3D) region, while a base station decodes the selected position from multi-snapshot monostatic sensing echoes. Near-field spherical wavefronts resolve both angle and range, expanding the embodied-symbol space from a 2D plane to a 3D volume. This gain, however, comes with a position-dependent and anisotropic reliability geometry. We characterize this geometry through the pairwise Bhattacharyya distance and derive a local ellipsoidal representation of the resulting 3D confusability regions, whose principal axes quantify directional sensing resolution. The ellipsoid further degenerates into the 2D transverse ellipse in the far-field limit, unifying the two regimes. We then formulate the finite-snapshot $\epsilon$-capacity and translate reliable codebook design into a 3D packing problem. A face-centered cubic construction provides an achievable rate, while a geometric converse yields a complementary upper bound. Numerical results validate the proposed geometry and demonstrate the capacity gain of near-field volumetric packing over far-field planar packing. These results establish a unified geometric and information-theoretic framework for communication through deliberately configured physical states.

cs.IT

AirTF: Over-the-Air Token Fusion for Task-Oriented Multi-Modal Token Communications

In the Internet of Vehicles (IoV), transmitting high-dimensional multi-modal sensory data to edge servers for time-sensitive tasks faces severe spectrum bottlenecks. To address this, we propose a foundation model-driven over-the-air token fusion (AirTF) framework for task-oriented multi-modal token communications. Unlike existing schemes for segmentation that rely on convolutional neural networks (CNNs) with limited local receptive fields, AirTF leverages vision transformer (ViT) encoders to extract globally contextualized semantic tokens from distributed heterogeneous sensors. By concurrently transmitting these spatially aligned tokens over a shared wireless channel, our framework exploits the superposition property of the multiple access channel to inherently fuse complementary multi-modal semantics (e.g., RGB and infrared) directly over the air. This mechanism significantly enhances spectral efficiency compared to orthogonal transmission. Furthermore, the integration of a pre-trained foundation model provides critical visual priors, effectively addressing the data-hungry nature of ViTs on limited, scenario-specific semantic segmentation datasets. Experiments demonstrate that AirTF consistently outperforms orthogonal transmission and CNN-based fusion baselines across AWGN and fading channels. Additional evaluations under a three-user setting, residual synchronization errors, and imperfect channel state information estimation further confirm its robustness. The source code will be made publicly available upon acceptance.

eess.IV

Stepped Frequency Division Multiplexing: A Jump-Free Continuous-Time AFDM Waveform

Affine frequency division multiplexing (AFDM) has emerged as a promising modulation scheme for doubly selective channels, but its canonical continuous-time realization, referred to herein as piecewise continuous AFDM (PC-AFDM), has been observed to exhibit high out-of-band emission (OOBE) whose mechanism has not been analytically characterized. This paper shows that the underlying cause is frequency wrapping, which introduces internal envelope jumps between AFDM sampling instants and generates a high-frequency spectral tail distinct from ordinary block truncation. To eliminate these discontinuities without altering the inverse discrete affine Fourier transform (IDAFT) output sequence, we propose stepped frequency division multiplexing (SFDM). In SFDM, the instantaneous frequency is kept constant at the midpoint of the wrapped chirp within each sampling interval, while the phase is continuously accumulated across interval boundaries. We prove that, under continuous phase accumulation and without additional phase correction, the midpoint choice is the unique sample-preserving choice for arbitrary chirp-rate parameter. The resulting waveform is continuous within each AFDM block, reduces OOBE, and preserves the standard AFDM modulation matrix, guard-interval structure, and receiver processing. Moreover, under fractional-delay propagation, SFDM mitigates the receiver sensitivity that arises when delayed sampling points fall near wrapping-induced discontinuities in PC-AFDM. Numerical results verify the theoretical tail coefficients, demonstrate OOBE reduction, and show improved receiver robustness in the high-percentile and worst-case regimes. These findings establish SFDM as a spectrally cleaner and more reliable physical layer for AFDM systems.

eess.SP

Embodied Communication: Sensing-Induced Reliability Fields and Capacity Bounds

This paper introduces embodied communication, a new wireless communication modality in which information is imprinted onto environmental states and recovered by the receiver through sensing. No dedicated communication transmitter is activated, and no additional communication spectrum is occupied; instead, the sensed environment itself becomes the carrier of information. The key insight is that sensing must be reinterpreted for communication. Rather than asking how accurately an unknown physical state can be estimated, embodied communication asks how reliably two states can be distinguished. We formalize this idea through a multi-snapshot radio frequency (RF) sensing model and derive a sensing-induced reliability field that quantifies the distinguishability between physical states. This field turns embodied symbol design into a geometric packing problem shaped by the sensing resolution of the infrastructure. For this embodied channel, we characterize the finite-snapshot $\epsilon$-capacity through achievable designs and converses. We develop lattice-based codebooks, obtain a closed-form hexagonal design under a main-lobe approximation, and establish information-theoretic and geometric upper bounds. We further reveal an intrinsic sensing-duration tradeoff: more sensing snapshots improve reliability, but also lengthen each embodied symbol, leading to a finite optimal sensing time. These results expose a latent communication pathway in sensing-enabled infrastructure and show how the environment can be transformed from a passive backdrop into an active information carrier.

cs.IT

Delay-Doppler Domain Channel Estimation: What if Sparsity is Unknown?

Sparsity in the delay-Doppler (DD) domain enables efficient channel estimation, but the realization-wise sparsity level is rarely known in advance, and it fluctuates. What if we could estimate the channel without ever knowing how many delays or Dopplers are active? This paper answers that question. We propose a sparsity-agnostic structured estimator that requires no prior knowledge of delay or Doppler sparsity budgets. The key idea is to exploit the Cartesian-product structure of DD support (active delays share a common Doppler set) and to select the support dimensions directly from the data via the Bayesian information criterion. We instantiate the framework on an affine frequency division multiplexing system, where the observation model naturally admits an on-grid DD representation. Numerical results demonstrate that it recovers the exact support with high probability and achieves near-oracle channel reconstruction accuracy, consistently outperforming fixed-budget baselines and sparse Bayesian learning. The approach is waveform-agnostic and offers a practical, adaptive solution for DD-domain channel estimation under unknown and time-varying sparsity.

cs.IT

Accelerating Heterogeneous Agent Collaboration in Dynamic Edge Networks

Deploying large language models (LLMs) at the network edge is hindered by their enormous cost, yet the reasoning quality they provide remains indispensable. Heterogeneous collaboration between edge small models and a server LLM has emerged as a promising direction, but existing methods fail under the dynamic conditions of multi-user contention, autoregressive generation, and time-varying resources. This paper puts forward a process reward model (PRM)-aided two-stage decoupled acceleration (PRADA) framework, which is built on a fundamental change of perspective: instead of querying a PRM online, which cripples multi-user systems with prohibitive latency, we use the PRM solely as an offline teacher. Its reasoning-quality intuition is fully distilled into a lightweight policy that screen each step locally, without any context upload, while a Lagrangian scheduler at the server resolves resource contention through a threshold-structured policy. Across diverse reasoning benchmarks, PRADA retains the vast majority of the LLM's accuracy while substantially reducing end-to-end latency. The results further reveal threshold effects for both server parallel capacity and total bandwidth: performance saturates beyond critical resource levels, after which the system bottleneck shifts from queuing to computation or from communication to contention. These structural findings provide actionable guidance for joint provisioning of computation and communication resources without requiring per-benchmark tuning.

cs.NI

AI Tool Discovery at Scale: All You Need is DNS

The coming era of autonomous AI agents demands a discovery mechanism capable of navigating millions of tools, yet existing solutions buckle under O(N) complexity and centralized governance. Instead of building another fragile overlay, we propose ToolDNS, a radical framework that retrofits semantic tool discovery onto the Internet's most resilient substrate: the Domain Name System (DNS). By embedding functional intent and organizational trust into a hierarchical namespace, ToolDNS transforms an expensive semantic search into a series of lightweight, O(log N) name resolutions. We introduce three protocol-compliant enhancements to enable decentralized governance and semantic pruning: partially unfolded names, EDNS0 intent payloads, and logical subdomains. To rigorously evaluate this approach across the fragmented tooling landscape, we construct and release a large-scale heterogeneous benchmark comprising 33,688 real-world tools spanning MCP, A2A, RESTful, and Skill protocols. On this dataset, ToolDNS slashes the per-query search space by 95.26% while matching state-of-the-art retrieval accuracy. Furthermore, its UDP-native design reduces discovery latency by orders of magnitude compared to HTTP-based registries. Our work demonstrates that scalable AI interoperability requires not more middleware, but a smarter utilization of the infrastructure already beneath our feet.

cs.AI

Cayley Graph Optimization for Scalable Multi-Agent Communication Topologies

Large-scale multi-agent communication has long faced a scalability bottleneck: fully connected networks require quadratic complexity, yet existing sparse topologies rely on hand-crafted rules. This paper treats the communication graph itself as a design variable and proposes CayleyTopo, a family of circulant Cayley graphs whose generator sets are optimized to minimize diameter, directly targeting worst-case information propagation speed. To navigate the enormous search space of possible generator sets, we develop a lightweight reinforcement learning framework that injects a number-theoretic prior to favor structurally rich generators, alongside a message-propagation score that provides dense connectivity feedback during construction. The resulting CayleyTopo consistently outperforms existing hand-crafted topologies, achieving faster information dissemination, greater resilience to link failures, and lower communication load, all while approaching the theoretical Moore bound. Our study opens the door to scalable, robust, and efficient communication foundations for future multi-agent systems, where the graph itself becomes optimizable rather than a fixed constraint.

cs.NI

Multi-Mode Pinching-Antenna Systems: Polarization-Aware Full-Wave Modeling and Optimization

Millimeter-wave and terahertz communications face a fundamental challenge: overcoming severe path loss without sacrificing spectral efficiency. Pinching antenna systems (PASS) address this by bringing radiators physically close to users, yet existing frameworks treat the waveguide as a mere transmission line, overlooking its inherent multi-mode capabilities and the critical role of polarization. This paper develops the first polarization-aware, full-wave electromagnetic model for multi-mode PASS (MMPASS), capturing spatial radiation patterns, modal polarization states, and polarization matching efficiency from first principles. Leveraging this physically grounded model, we reveal fundamental trade-offs among waveguide attenuation, atmospheric absorption, and geometric spreading, yielding closed-form solutions for optimal PA placement and orientation in single-user scenarios. Extending to multi-user settings, we propose a modular optimization framework that integrates fractional programming with closed-form polarization updates, scaling gracefully to arbitrary numbers of waveguides, PAs, and users. Numerical results show that MMPASS achieves up to a 167% increase in spectral efficiency compared with single-mode PASS. Moreover, when comparing MMPASS with its polarization-ignorant counterpart, polarization awareness alone improves the sum rate by up to 23%. By bridging rigorous electromagnetic theory with scalable optimization, MMPASS establishes a physically complete and practically viable foundation for future high-frequency wireless networks.

cs.IT

Unanticipated Adversarial Robustness of Semantic Communication

Semantic communication, enabled by deep joint source-channel coding (DeepJSCC), is widely expected to inherit the vulnerability of deep learning to adversarial perturbations. This paper challenges this prevailing belief and reveals a counterintuitive finding: semantic communication systems exhibit unanticipated adversarial robustness that can exceed that of classical separate source-channel coding systems. On the theoretical front, we establish fundamental bounds on the minimum attack power required to induce a target distortion, overcoming the analytical intractability of highly nonlinear DeepJSCC models by leveraging Lipschitz smoothness. We prove that the implicit regularization from noisy training forces decoder smoothness, a property that inherently provides built-in protection against adversarial attacks. To enable rigorous and fair comparison, we develop two novel attack methodologies that address previously unexplored vulnerabilities: a structure-aware vulnerable set attack that, for the first time, exploits graph-theoretic vulnerabilities in LDPC codes to induce decoding failure with minimal energy, and a progressive gradient ascent attack that leverages the differentiability of DeepJSCC to efficiently find minimum-power perturbations. Designing such attacks is challenging, as classical systems lack gradient information while semantic systems require navigating high-dimensional, non-convex spaces; our methods fill these critical gaps in the literature. Extensive experiments demonstrate that semantic communication requires up to $14$-$16\times$ more attack power to achieve the same distortion as classical systems, empirically substantiating its superior robustness.

cs.IT

Rateless DeepJSCC for Broadcast Channels: a Rate-Distortion-Complexity Tradeoff

In recent years, numerous data-intensive broadcasting applications have emerged at the wireless edge, calling for a flexible tradeoff between distortion, transmission rate, and processing complexity. While deep learning-based joint source-channel coding (DeepJSCC) has been identified as a potential solution to data-intensive communications, most of these schemes are confined to worst-case solutions, lack adaptive complexity, and are inefficient in broadcast settings. To overcome these limitations, this paper introduces nonlinear transform rateless source-channel coding (NTRSCC), a variable-length JSCC framework for broadcast channels based on rateless codes. In particular, we integrate learned source transformations with physical-layer LT codes, develop unequal protection schemes that exploit decoder side information, and devise approximations to enable end-to-end optimization of rateless parameters. Our framework enables heterogeneous receivers to adaptively adjust their received number of rateless symbols and decoding iterations in belief propagation, thereby achieving a controllable tradeoff between distortion, rate, and decoding complexity. Simulation results demonstrate that the proposed method enhances image broadcast quality under stringent communication and processing budgets over heterogeneous edge devices.

cs.IT

From Simulation to Reality: Practical Deep Reinforcement Learning-based Link Adaptation for Cellular Networks

Link Adaptation (LA) that dynamically adjusts the Modulation and Coding Schemes (MCS) to accommodate time-varying channels is crucial and challenging in cellular networks. Deep reinforcement learning (DRL)-based LA that learns to make decision through the interaction with the environment is a promising approach to improve throughput. However, existing DRL-based LA algorithms are typically evaluated in simplified simulation environments, neglecting practical issues such as ACK/NACK feedback delay, retransmission and parallel hybrid automatic repeat request (HARQ). Moreover, these algorithms overlook the impact of DRL execution latency, which can significantly degrade system performance. To address these challenges, we propose Decoupling-DQN (DC-DQN), a new DRL framework that separates traditional DRL's coupled training and inference processes into two modules based on Deep Q Networks (DQN): a real-time inference module and an out-of-decision-loop training module. Based on this framework, we introduce a novel DRL-based LA algorithm, DC-DQN-LA. The algorithm incorporates practical considerations by designing state, action, and reward functions that account for feedback delays, parallel HARQ, and retransmissions. We implemented a prototype using USRP software-defined radios and srsRAN software. Experimental results demonstrate that DC-DQN-LA improves throughput by 40\% to 70\% in mobile scenario compared with baseline LA algorithms, while maintaining comparable block error rates, and can quickly adapt to environment changes in mobile-to-static scenario. These results highlight the efficiency and practicality of the proposed DRL-based LA algorithm.

cs.NI

Deep Variable-Length Feedback Codes

Deep learning has enabled significant advances in feedback-based channel coding, yet existing learned schemes remain fundamentally limited: they employ fixed block lengths, suffer degraded performance at high rates, and cannot fully exploit the adaptive potential of feedback. This paper introduces Deep Variable-Length Feedback (DeepVLF) coding, a flexible coding framework that dynamically adjusts transmission length via learned feedback. We propose two complementary architectures: DeepVLF-R, where termination is receiver-driven, and DeepVLF-T, where the transmitter controls termination. Both architectures leverage bit-group partitioning and transformer-based encoder-decoder networks to enable fine-grained rate adaptation in response to feedback. Evaluations over AWGN and 5G-NR fading channels demonstrate that DeepVLF substantially outperforms state-of-the-art learned feedback codes. It achieves the same block error rate with 20%-55% fewer channel uses and lowers error floors by orders of magnitude, particularly in high-rate regimes. Encoding dynamics analysis further reveals that the models autonomously learn a two-phase strategy analogous to classical Schalkwijk-Kailath coding: an initial information-carrying phase followed by a noise-cancellation refinement phase. This emergent behavior underscores the interpretability and information-theoretic alignment of the learned codes.

cs.IT

Rich-ARQ: From 1-bit Acknowledgment to Rich Neural Coded Feedback

This paper reimagines the foundational feedback mechanism in wireless communication, transforming the prevailing 1-bit binary ACK/NACK with a high-dimensional, information-rich vector to transform passive acknowledgment into an active collaboration. We present Rich-ARQ, a paradigm that introduces neural-coded feedback for collaborative physical-layer channel coding between transmitter and receiver. To realize this vision in practice, we develop a novel asynchronous feedback code that eliminates stalling from feedback delays, adapts dynamically to channel fluctuations, and features a lightweight encoder suitable for on-device deployment. We materialize this concept into the first full-stack, standard-compliant software-defined radio prototype, which decouples AI inference from strict radio timing. Comprehensive over-the-air experiments demonstrate that Rich-ARQ achieves significant SNR gains over conventional 1-bit hybrid ARQ and remarkable latency reduction over prior learning-based feedback codes, moving the promise of intelligent feedback from theory to a practical, high-performance reality for next-generation networks.

cs.IT

How should covariates be handled in randomized trials? Empirical evidence from 50 trials and recommendations for practice

Background and Objective: Covariate adjustment can improve precision and power in randomized clinical trials and is recommended by major regulatory agencies. However, there is limited empirical evidence on how different adjustment strategies perform across diverse real-world trials, leaving uncertainty about which methods and covariates should be prespecified in statistical analysis plans. We aim to address this gap and provide practical recommendations. Methods: We conducted a large-scale empirical study using individual-level data from 50 publicly available randomized trials (29,094 participants; 574 treatment-outcome comparisons). We compared commonly used covariate-adjusted estimators, including analysis of covariance, inverse-probability weighting, g-computation, and machine-learning-based approaches, combined with three covariate-selection strategies. Performance was evaluated using precision gains, changes in point estimates, computational reliability, and the probability that covariate adjustment altered statistical significance relative to an unadjusted analysis. Results: Covariate adjustment improved precision in most settings, with a median variance reduction of 13.3\% for continuous outcomes and 4.6\% for binary outcomes. Parsimonious regression approaches using a small prespecified set of prognostic covariates performed as well as or better than more complex methods, particularly in small to medium samples. Machine-learning-based estimators did not provide additional precision and were more prone to computational failure for binary outcomes. Conclusions: Across trials, parsimonious covariate adjustment provided consistent efficiency gains without introducing systematic bias. These findings support routine covariate adjustment in primary trial analyses. All curated datasets and analysis code are openly released to support future clinical research.

stat.AP

Hierarchical Online-Scheduling for Energy-Efficient Split Inference with Progressive Transmission

Device-edge collaborative inference with Deep Neural Networks (DNNs) faces fundamental trade-offs among accuracy, latency and energy consumption. Current scheduling exhibits two drawbacks: a granularity mismatch between coarse, task-level decisions and fine-grained, packet-level channel dynamics, and insufficient awareness of per-task complexity. Consequently, scheduling solely at the task level leads to inefficient resource utilization. This paper proposes a novel ENergy-ACcuracy Hierarchical optimization framework for split Inference, named ENACHI, that jointly optimizes task- and packet-level scheduling to maximize accuracy under energy and delay constraints. A two-tier Lyapunov-based framework is developed for ENACHI, with a progressive transmission technique further integrated to enhance adaptivity. At the task level, an outer drift-plus-penalty loop makes online decisions for DNN partitioning and bandwidth allocation, and establishes a reference power budget to manage the long-term energy-accuracy trade-off. At the packet level, an uncertainty-aware progressive transmission mechanism is employed to adaptively manage per-sample task complexity. This is integrated with a nested inner control loop implementing a novel reference-tracking policy, which dynamically adjusts per-slot transmit power to adapt to fluctuating channel conditions. Experiments on ImageNet dataset demonstrate that ENACHI outperforms state-of-the-art benchmarks under varying deadlines and bandwidths, achieving a 43.12\% gain in inference accuracy with a 62.13\% reduction in energy consumption under stringent deadlines, and exhibits high scalability by maintaining stable energy consumption in congested multi-user scenarios.

cs.NI