Search arXivSearch

arXiv subjects

Ruijing Liu

Publications and source records attributed to Ruijing Liu.

8 recordsLinked to original sources

WiFi based Multi user Activity Recognition via User Conditioned Spatial Attention

WiFi-based human activity recognition has achieved high accuracy under the single-user scene. Recognizing activities performed by multiple concurrent users remains challenging, because their body-reflected propagation paths superimpose in the channel state information (CSI) measurement. Existing multi-user methods either decompose the signal before classification, suffering from error propagation, or predict activities jointly without reliably associating each prediction with the correct user. Moreover, attention mechanisms that have proven effective for single-user sensing compute a single user-agnostic attention map, blending patterns from different users and providing no mechanism to associate features with specific user slots. This paper proposes a user-conditioned spatial attention (UCSA) module that generates per-user spatial attention maps by conditioning multi-scale spatial attention on learnable user-ID embeddings via feature-wise linear modulation (FiLM). UCSA creates a deep coupling between spatial feature enhancement and user differentiation: the attention module itself becomes user-aware, producing distinct attention maps for each user slot rather than relying on a late-fusion embedding addition. Combined with a shareable multi-semantic spatial attention (SMSA) module for multi-scale spatial enhancement, the proposed framework addresses both feature quality and user differentiation in a unified architecture. Experiments on the WiMANS benchmark across three indoor environments and up to five concurrent users show that the proposed method outperforms nine baselines in all settings, achieving an average accuracy of over 93\% across three environments even with five users. Ablation studies confirm that both SMSA and UCSA contribute substantially to recognition accuracy, with UCSA providing the critical link between shared features and per-user predictions.

eess.SP

Robust Cross-Domain WiFi Fall Detection via Physics-Driven Attention-Enhanced Transformers

Device-free fall detection utilizing WiFi Channel State Information (CSI) has emerged as a promising, privacy-preserving solution for elderly health monitoring in the Internet of Things (IoT) era. However, existing deep learning approaches suffer from severe performance degradation when deployed in unseen environments due to static background overfitting and Non-Line-of-Sight (NLoS) signal attenuation. To address these critical bottlenecks, we propose a robust, domain-generalizable framework featuring a novel Attention-Enhanced CNN-Transformer hybrid architecture. First, we design a physics-driven \textbf{Dynamic Variance Gate (DVG)} to dynamically calculate local temporal variance, acting as a soft-attention mask that eliminates static environmental DC components while amplifying dynamic human motion. Second, we introduce a Physics-Aware Data Augmentation strategy to force the network to learn invariant morphological signatures rather than environment-specific noise. Furthermore, a Convolutional Block Attention Module (CBAM) is integrated to refine spatiotemporal features prior to Transformer-based sequence modeling. Extensive cross-domain evaluations across four distinct indoor environments demonstrate that our method achieves 97.6\% accuracy in NLoS scenarios and 98.8\% in completely unseen environments without target-domain fine-tuning. Finally, we deploy the proposed framework on an edge computing system equipped with commercial WiFi NICs. Real-world live inference field tests confirm the system's robustness against unseen environmental layouts and its capability for continuous, low-latency whole-home safety monitoring.

eess.SP

A BEV-Fusion Based Framework for Sequential Multi-Modal Beam Prediction in mmWave Systems

Beam prediction is critical for reducing beam-training overhead in millimeter-wave (mmWave) systems, especially in high-mobility vehicular scenarios. This paper presents a BEV-Fusion based framework that unifies camera, LiDAR, radar, and GPS modalities in a shared bird's-eye-view (BEV) representation for spatially consistent multi-modal fusion. Unlike priorapproaches that fuse globally pooled one-dimensional features, the proposed method performs fusion in BEV space to preservecross-modal geometric structure and visual semantic density. A learned camera-to-BEV module based on cross-attention is adopted to generate BEV-aligned visual features without relying on precise camera calibration, and a temporal transformer is used to aggregate five-step sequential observations for motion-aware beam prediction. Experiments on the DeepSense 6G benchmark show that BEV-Fusion achieves approximately 87% distance- based accuracy (DBA) on scenarios 32, 33 and 34, outperforming the TransFuser baseline. These results indicate that BEV-space fusion provides an effective spatial abstraction for sensing-assisted beam prediction.

eess.SP

WiFi-based Cross-Domain Gesture Recognition Using Attention Mechanism

While fulfilling communication tasks, wireless signals can also be used to sense the environment. Among various types of sensing media, WiFi signals offer advantages such as widespread availability, low hardware cost, and strong robustness to environmental conditions like light, temperature, and humidity. By analyzing Wi-Fi signals in the environment, it is possible to capture dynamic changes of the human body and accomplish sensing applications such as gesture recognition. Although many existing gesture sensing solutions perform well in-domain but lack cross-domain capabilities (i.e., recognition performance in untrained environments). To address this, we extract Doppler spectra from the channel state information (CSI) received by all receivers and concatenate each Doppler spectrum along the same time axis to generate fused images with multi-angle information as input features. Furthermore, inspired by the convolutional block attention module (CBAM), we propose a gesture recognition network that integrates a multi-semantic spatial attention mechanism with a self-attention-based channel mechanism. This network constructs attention maps to quantify the spatiotemporal features of gestures in images, enabling the extraction of key domain-independent features. Additionally, ResNet18 is employed as the backbone network to further capture deep-level features. To validate the network performance, we evaluate the proposed network on the public Widar3 dataset, and the results show that it not only maintains high in-domain accuracy of 99.72%, but also achieves high performance in cross-domain recognition of 97.61%, significantly outperforming existing best solutions.

cs.CV

Deep Learning-Based CSI Feedback for XL-MIMO Systems in the Near-Field Domain

In this paper, we consider an extremely large-scale massive multiple-input-multiple-output (XL-MIMO) system. As the scale of antenna arrays increases, the range of near-field communications also expands. In this case, the signals no longer exhibit planar wave characteristics but spherical wave characteristics in the near-field channel, which makes the channel state information (CSI) highly complex. Additionally, the increase of the antenna arrays scale also makes the size of the CSI matrix significantly increase. Therefore, CSI feedback in the near-field channel becomes highly challenging. To solve this issue, we propose a deep-learning (DL)-based ExtendNLNet that can compress the CSI, and further reduce the overhead of CSI feedback. In addition, we have introduced the Non-Local block to obtain a larger area of CSI features. Simulation results show that the proposed ExtendNLNet can significantly improve the CSI recovery quality compared to other DL-based methods.

eess.SP

Deep Learning-Based CSI Feedback for RIS-Aided Massive MIMO Systems with Time Correlation

In this paper, we consider an reconfigurable intelligent surface (RIS)-aided frequency division duplex (FDD) massive multiple-input multiple-output (MIMO) downlink system.In the FDD systems, the downlink channel state information (CSI) should be sent to the base station through the feedback link. However, the overhead of CSI feedback occupies substantial uplink bandwidth resources in RIS-aided communication systems. In this work, we propose a deep learning (DL)-based scheme to reduce the overhead of CSI feedback by compressing the cascaded CSI. In the practical RIS-aided communication systems, the cascaded channel at the adjacent slots inevitably has time correlation. We use long short-term memory to learn time correlation, which can help the neural network to improve the recovery quality of the compressed CSI. Moreover, the attention mechanism is introduced to further improve the CSI recovery quality. Simulation results demonstrate that our proposed DLbased scheme can significantly outperform other DL-based methods in terms of the CSI recovery quality

eess.SP

Energy Minimization for Active RIS-Aided UAV-Enabled SWIPT Systems

In this paper, we consider an active reconfigurable intelligent surface (RIS)-aided unmanned aerial vehicle(UAV)-enabled simultaneous wireless information and power transfer(SWIPT) system with multiple ground users. Compared with the conventional passive RIS, the active RIS deploying the internally integrated amplifiers can offset part of the multiplicative fading. In this system, we deal with an optimization problem of minimizing the total energy cost of the UAV. Specifically, we alternately optimize the trajectories, the hovering time, and the reflection vectors at the active RIS by using the successive convex approximation (SCA) method. Simulation results show that the active RIS performs better in energy saving than the conventional passive RIS.

eess.SP

Optimal data placements for triple replication

Given a set $V$ of $v$ servers along with $b$ files (data), each file is replicated (placed) on exactly $k$ servers and thus a file can be represented by a set of $k$ servers. Then we produce a data placement consisting of $b$ subsets of $V$ called blocks, each of size $k$. Each server has some probability to fail and we want to find a placement that minimizes the variance of the number of available files. It was conjectured that there always exists an optimal data placement (with variance better than any other placement for any value of the probability of failure). An optimal data placement for triple replication with $b$ blocks (of size three) on a $v$-set was proved to exist by Wei et al. if $v$ and $b$ are not excluded by two conditions. This article concentrates on the parameters $v, b$ satisfying the two conditions and characterizes the combinatorial properties of the corresponding optimal data placements. Nearly well-balanced triple systems (NWBTSs) are defined to produce optimal data placements. Many constructions for NWBTSs are developed, mainly by constructing candelabra systems with various desirable partitions. The main result of this article is that there always exist optimal data placements for triple replication with $b$ blocks on a $v$-set possibly except when $v\equiv 4$ (mod 24) or $v=50,74$, and $\frac{λv(v-1)}{6}-\frac{v}{6} < b < \frac{λv(v-1)}{6}+\frac{v}{6}$ for an odd integer $λ$.

math.CO