Search arXivSearch

arXiv subjects

Yewen Cao

Publications and source records attributed to Yewen Cao.

7 recordsLinked to original sources

Networked Embodied Communication: From Collective Distinguishability to Communication Reliability

Embodied agents need to convey information to surrounding infrastructure, but their active communication interfaces may be unavailable, constrained, or intentionally inactive. Their ability to manipulate physical states offers a complementary path: messages can be encoded in deliberately selected configurations and recovered through infrastructure sensing. This principle underlies embodied communication. Yet physical differences do not guarantee distinguishable messages: a single sensing viewpoint may leave ambiguities that repeated sensing cannot resolve. This paper develops networked embodied communication, where distributed access points (APs) jointly observe message-bearing scatterer positions under fixed illumination. Under a correlated Gaussian sensing model, we characterize the additional distinguishability supplied by receive APs, establish exact redundancy conditions, and reveal how distinctions absent from individual observations can emerge through cross-AP statistical relationships. We then establish the exact asymptotic optimal maximum-error behavior of a finite alphabet under repeated independent sensing. The largest group of indistinguishable messages determines the error floor; once all messages are distinguishable, the minimum pairwise Chernoff information determines the error exponent. For a given alphabet, receiver cooperation can therefore eliminate an error floor that repetition at any individual AP cannot overcome. Building on these results, we derive finite-budget reliability conditions and jointly design the receive AP set and message-bearing positions. Numerical results show that the proposed search closely approaches exact benchmarks on reduced instances with substantially fewer candidate evaluations than exhaustive enumeration, while receiver cooperation reduces the sensing intervals needed to guarantee reliable decoding.

eess.SP

Resolving the Discontinuity of Continuous-Time AFDM Waveforms

Continuous-time affine frequency division multiplexing (AFDM) waveforms, constructed via frequency wrapping and phase correction, are known to be sample-wise equivalent to the widely adopted discrete AFDM framework. In this paper, we uncover a fundamental and previously overlooked flaw in this construction: its complex envelope is inherently discontinuous for generic chirp parameters. We show that these discontinuities are the direct cause of the high out-of-band emission (OOBE). To resolve this issue, we propose a fundamentally different continuous-time waveform, termed stepped frequency division multiplexing (SFDM). Unlike conventional approaches that allow continuous frequency variation, SFDM freezes the instantaneous frequency at the midpoint of the underlying chirp trajectory within each Nyquist sampling interval. This design yields a complex envelope that is strictly continuous over the entire symbol duration while preserving exact sample-wise equivalence with discrete AFDM. A unified spectral analysis reveals that the superior OOBE performance of SFDM stems from the absence of internal jump discontinuities, which otherwise dominate the far-out spectral roll-off. Numerical results confirm that SFDM consistently achieves significantly lower OOBE across a wide range of chirp rates.

eess.SP

Stepped Frequency Division Multiplexing: A Jump-Free Continuous-Time AFDM Waveform

Affine frequency division multiplexing (AFDM) has emerged as a promising modulation scheme for doubly selective channels, but its canonical continuous-time realization, referred to herein as piecewise continuous AFDM (PC-AFDM), has been observed to exhibit high out-of-band emission (OOBE) whose mechanism has not been analytically characterized. This paper shows that the underlying cause is frequency wrapping, which introduces internal envelope jumps between AFDM sampling instants and generates a high-frequency spectral tail distinct from ordinary block truncation. To eliminate these discontinuities without altering the inverse discrete affine Fourier transform (IDAFT) output sequence, we propose stepped frequency division multiplexing (SFDM). In SFDM, the instantaneous frequency is kept constant at the midpoint of the wrapped chirp within each sampling interval, while the phase is continuously accumulated across interval boundaries. We prove that, under continuous phase accumulation and without additional phase correction, the midpoint choice is the unique sample-preserving choice for arbitrary chirp-rate parameter. The resulting waveform is continuous within each AFDM block, reduces OOBE, and preserves the standard AFDM modulation matrix, guard-interval structure, and receiver processing. Moreover, under fractional-delay propagation, SFDM mitigates the receiver sensitivity that arises when delayed sampling points fall near wrapping-induced discontinuities in PC-AFDM. Numerical results verify the theoretical tail coefficients, demonstrate OOBE reduction, and show improved receiver robustness in the high-percentile and worst-case regimes. These findings establish SFDM as a spectrally cleaner and more reliable physical layer for AFDM systems.

eess.SP

DP^2-VL: Private Photo Dataset Protection by Data Poisoning for Vision-Language Models

Recent advances in visual-language alignment have endowed vision-language models (VLMs) with fine-grained image understanding capabilities. However, this progress also introduces new privacy risks. This paper first proposes a novel privacy threat model named identity-affiliation learning: an attacker fine-tunes a VLM using only a few private photos of a target individual, thereby embedding associations between the target facial identity and their private property and social relationships into the model's internal representations. Once deployed via public APIs, this model enables unauthorized exposure of the target user's private information upon input of their photos. To benchmark VLMs' susceptibility to such identity-affiliation leakage, we introduce the first identity-affiliation dataset comprising seven typical scenarios appearing in private photos. Each scenario is instantiated with multiple identity-centered photo-description pairs. Experimental results demonstrate that mainstream VLMs like LLaVA, Qwen-VL, and MiniGPT-v2, can recognize facial identities and infer identity-affiliation relationships by fine-tuning on small-scale private photographic dataset, and even on synthetically generated datasets. To mitigate this privacy risk, we propose DP2-VL, the first Dataset Protection framework for private photos that leverages Data Poisoning. Though optimizing imperceptible perturbations by pushing the original representations toward an antithetical region, DP2-VL induces a dataset-level shift in the embedding space of VLMs'encoders. This shift separates protected images from clean inference images, causing fine-tuning on the protected set to overfit. Extensive experiments demonstrate that DP2-VL achieves strong generalization across models, robustness to diverse post-processing operations, and consistent effectiveness across varying protection ratios.

cs.CV

Agile Affine Frequency Division Multiplexing

The advancement to 6G calls for waveforms that transcend static robustness to achieve intelligent adaptability. Affine Frequency Division Multiplexing (AFDM), despite its strength in doubly-dispersive channels, has been confined by chirp parameters optimized for worst-case scenarios. This paper shatters this limitation with Agile-AFDM, a novel framework that endows AFDM with dynamic, data-aware intelligence. By redefining chirp parameters as optimizable variables for each transmission block based on real-time channel and data information, Agile-AFDM transforms into an adaptive platform. It can actively reconfigure its waveform to minimize peak-to-average power ratio (PAPR) for power efficiency, suppress inter-carrier interference (ICI) for communication reliability, or reduce Cramer-Rao bound (CRLB) for sensing accuracy. This paradigm shift from a static, one-size-fits-all waveform to a context-aware signal designer is made practical by efficient, tailored optimization algorithms. Comprehensive simulations demonstrate that this capability delivers significant performance gains across all metrics, surpassing conventional OFDM and static AFDM. Agile-AFDM, therefore, offers a crucial step forward in the design of agile waveforms for 6G and beyond.

cs.IT

Fractional Fourier Domain PAPR Reduction

High peak-to-average power ratio (PAPR) has long posed a challenge for multi-carrier systems, impacting amplifier efficiency and overall system performance. This paper introduces dynamic angle fractional Fourier division multiplexing (DA-FrFDM), an innovative multi-carrier system that effectively reduces PAPR for both QAM and Gaussian signals with minimal signaling overhead. DA-FrFDM leverages the fractional Fourier domain to balance PAPR characteristics between the time and frequency domains, achieving significant PAPR reduction while preserving signal quality. Furthermore, DA-FrFDM refines signal processing and enables one-tap equalization in the fractional Fourier domain through the simple multiplication of time-domain signals by a quadratic phase sequence. Our results show that DA-FrFDM not only outperforms existing PAPR reduction techniques but also retains efficient inter-carrier interference (ICI) mitigation capabilities in doubly dispersive channels.

cs.IT

Pyramid Multi-branch Fusion DCNN with Multi-Head Self-Attention for Mandarin Speech Recognition

As one of the major branches of automatic speech recognition, attention-based models greatly improves the feature representation ability of the model. In particular, the multi-head mechanism is employed in the attention, hoping to learn speech features of more aspects in different attention subspaces. For speech recognition of complex languages, on the one hand, a small head size will lead to an obvious shortage of learnable aspects. On the other hand, we need to reduce the dimension of each subspace to keep the size of the overall feature space unchanged when we increase the number of heads, which will significantly weaken the ability to represent the feature of each subspace. Therefore, this paper explores how to use a small attention subspace to represent complete speech features while ensuring many heads. In this work we propose a novel neural network architecture, namely, pyramid multi-branch fusion DCNN with multi-head self-attention. The proposed architecture is inspired by Dilated Convolution Neural Networks (DCNN), it uses multiple branches with DCNN to extract the feature of the input speech under different receptive fields. To reduce the number of parameters, every two branches are merged until all the branches are merged into one. Thus, its shape is like a pyramid rotated 90 degrees. We demonstrate that on Aishell-1, a widely used Mandarin speech dataset, our model achieves a character error rate (CER) of 6.45% on the test sets.

eess.AS