Search arXivSearch

arXiv subjects

Kai Wu

Publications and source records attributed to Kai Wu.

At least 19 recordsLinked to original sources

Enabling High-Bandwidth Flash for Generative Recommendation Serving with Write-Aware KV Cache Policy

Generative recommendation (GR) systems increasingly leverage user-level KV cache reuse to avoid recomputing long user histories. However, the growing KV cache capacity and bandwidth requirements introduce new challenges for memory system. High-Bandwidth Flash (HBF) provides a promising solution by offering substantially higher capacity than HBM while approaching HBM-class read bandwidth, enabling larger scale KV cache retention and improved serving throughput. Yet conventional Least-Recently-Used (LRU) KV cache management tightly couples KV cache writes with cache misses, generating excessive write traffic that rapidly exhausts flash endurance. In this work, we evaluate a write-aware KV cache policy based on admission-controlled LRU-K for HBF-based GR serving. By filtering low-reuse users before cache admission, LRU-K decouples KV cache writes from misses and significantly reduces unnecessary writes. We develop an analytical model to characterize GR serving performance, KV cache write traffic, and HBF lifetime, and evaluate performance across diverse memory systems and GR workloads. Our results show that HBF-based systems achieve 3.8 to 4.7 times higher throughput than HBM-only systems. Moreover, LRU-K extends HBF lifetime from about one year under conventional LRU to over six years with a moderate K=10, while maintaining comparable or even slightly improved throughput. These results highlight the importance of write aware KV cache policy for sustainable HBF-based GR serving.

cs.AR

Multi-UE Networked Sensing: A New Paradigm for 6G Perceptive Mobile Networks

Networked sensing, which jointly exploits observations from multiple distributed nodes, is essential for unlocking the full sensing potential of integrated sensing and communications (ISAC). This article introduces multi-UE sensing, a new networked sensing paradigm for future perceptive mobile networks that exploits the correlated sensing observations naturally arising from distributed user equipment devices (UEs) interacting with common targets. Representative uplink, downlink, and hybrid sensing architectures are presented, together with a multi-view signal processing framework encompassing synchronization, correlation-aware parameter estimation, and sensing fusion. Key open challenges, including correlation modelling, target association, sensing information compression, and communication-sensing co-optimization, are also discussed.

eess.SP

The transverse matter Hamiltonian

Enrico Fermi in 1932 used classical Gauss equation to derive the Coulomb density--density interaction from the longitudinal electromagnetic potential. In this work we extend the Fermi procedure to the transverse component of the vector potential. By using a fully quantum canonical transformation, we replace the transverse vector potential with a current--current and current--current--density interactions. The transformed Hamiltonian is, then, projected in the fermionic space providing a matter--only, totally incoherent and gauge respecting Hamiltonian. The removal of coherence will avoid the breakdown of perturbation theory predicted by Haag's theorem, and make possible to introduce, in the transformed space, the diagrammatic approach. After discussing the implications of the transverse interactions on the current theories based on the Fermi procedure, we conclude by showing how the transverse Hamiltonian provides the quantum origin of the electromagnetic longitudinal--transverse splitting of phonons and other elemental excitations.

math-ph

Rainfall Sensing via Mobile Communication Signals

Rainfall monitoring is important for hydrological observation, disaster warning, and environmental sensing, but conventional rain gauges and weather radars suffer from sparse deployment and high infrastructure costs. This paper proposes PMN-RainSense, a rainfall sensing framework using sub-6-GHz mobile communication signals that supports practical single-antenna deployment. Unlike attenuation-based approaches, which are unreliable at sub-6 GHz because rain-induced attenuation over short mobile access links is only on the order of hundredths of a decibel, the proposed framework exploits fine-grained dynamics. A spectral-temporal channel state information (CSI) compensation method suppresses packet-wise timing and phase distortions while preserving sensing-relevant information. Rainfall-sensitive features are extracted from the delay-Doppler domain to mitigate environmental interference, with angle-domain filtering as an optional extension for multi-antenna receivers. Under bandwidth and antenna constraints, rainfall-correlated Doppler fluctuations serve as the dominant sensing signature, while Doppler-domain normalization improves robustness across links and deployments. Controlled WiFi experiments demonstrate rainfall-associated Doppler broadening and achieve a three-class classification accuracy of 95.48% using a random forest classifier. Long-Term Evolution (LTE) CSI measurements collected from cellular base stations over 11 carrier frequencies from 0.763 to 2.68 GHz yield a mean absolute error (MAE) of 0.25-0.27 mm/h for rainfall intensity estimation using a one-dimensional convolutional network.

eess.SP

MIRA: Medical Image Reflection for Agentic Diagnosis

Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or misleading evidence. Reliable diagnosis therefore requires not only acquiring additional observations, but also verifying whether tool actions are necessary and whether the resulting evidence supports the current hypothesis. We introduce MIRA (Medical Image Reflection for Agentic Diagnosis), a medical visual diagnostic framework for autonomous evidence search and reflective verification. MIRA dynamically invokes image-processing operations, including zooming, grounding, pointing, rotation, and measurement, as well as web search, while evaluating the relevance and consistency of the acquired evidence. We develop MIRA through a two-stage training strategy. First, a tool-augmented Monte Carlo Tree Search data engine explores diverse diagnostic hypotheses and jointly verifies visual grounding accuracy and semantic consistency to construct supervised fine-tuning trajectories. Second, reinforcement learning further improves decision-making through online reflective principle evolution: failure cases are distilled into candidate principles, and only principles that improve held-out rollout rewards are retained. Across nine medical visual reasoning benchmarks, MIRA achieves an average score of 64.73, improving its Qwen3-VL-8B backbone by 7.44 points. It also increases useful tool-use judgments from 56.2% to 73.8% and reduces harmful judgments from 8.9% to 1.6%. Qualitative analyses show that MIRA can re-examine evidence, correct premature conclusions, and adapt its tool-use strategy. Project page: https://MIRA-VL.github.io/

cs.CV

Ambiguity-Resolved Micro-Doppler Construction for Asynchronous Bistatic Sensing

Integrated sensing and communications (ISAC) can turn wireless networks into pervasive sensing platforms, but bistatic deployments are impaired by transceiver clock asynchrony, which induces random phase fluctuations in channel state information (CSI). CSI-ratio sanitization introduces nonlinear distortion that limits multi-target scalability and complicates delay- and angle-of-arrival (AoA)-domain processing. Cross-antenna conjugate multiplication (CACC) preserves a linear structure, but leaves mirror ambiguity and second-order by-products that corrupt motion-induced Doppler signatures. We develop a micro-Doppler construction framework that resolves the ambiguity and suppresses these residual by-products. Cyclic differencing first attenuates the dominant mirror component. We then exploit the facts that residual terms occur at differenced delay coordinates and lack an ordered AoA steering structure, designing a lightweight delay-AoA-Doppler pipeline that isolates the desired kinematic response without coherently accumulating residuals at target bins. The filtered responses are aggregated into a higher-SNR micro-Doppler representation. Ablation studies confirm the value of residual suppression and multidimensional filtering. On a large-scale WiFi gesture dataset, the proposed representation generalizes across four domain factors, attaining 90.1%-98.2% accuracy and outperforming representative baselines by about 18 percentage points on average.

eess.SP

MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors

Current visual generation models are capable of producing high-quality content, yet they lack a coherent perception of the spatial structure. Existing generative novel view synthesis methods typically introduce explicit geometry priors, which enforce spatial consistency but inherently restrict generalization in large view changes. In contrast, recent interactive generative methods favor implicit scene modeling, offering greater flexibility at the cost of precise camera control and geometry consistency. In this paper, we propose MetaView, a diffusion-based monocular novel view synthesis framework that enables rendering under large view changes from a single image. Our key insight is to combine implicit geometry modeling with minimal yet essential explicit 3D cues: we incorporate implicit geometry priors from a feed-forward geometry perception network to regularize structure without imposing restrictive reconstruction pipelines, while leveraging metric depth to anchor the generation to a metric scale. This design allows MetaView to achieve both geometry consistency and precise controllability. Extensive experiments demonstrate that, under challenging monocular large viewpoint changes, MetaView significantly outperforms existing methods and exhibits superior generalization. Our code is publicly available at https://github.com/KlingAIResearch/MetaView.

cs.CV

EVOM: Agentic Meta-Evolution of Actor-Critic Architectures for Reinforcement Learning

In actor-critic reinforcement learning, network architectures are typically manually designed. Automating this design is challenging because each candidate must be trained before evaluation, and the design space is open-ended. To address these challenges, we introduce EVOM, an agentic meta-evolution framework for discovering high-performance actor-critic architectures. We frame architecture search as a bi-level optimization: an inner loop trains weights via the low-fidelity proximal policy optimization (PPO), while an outer loop drives meta-evolution by iteratively refining architecture programs. Crucially, this outer loop is powered by an LLM-based design agent that operates purely as an architecture designer, completely decoupled from policy execution and environment control. Experiments reveal that EVOM outperforms the manually designed baseline, an LLM-guided random search, and the state-of-the-art LLM-guided programmatic policy search method MLES, delivering superior performance on Ant-v4 and HalfCheetah-v4. Ablation studies validate that both the meta-evolution loop and the LLM Design Agent are indispensable for final performance.

cs.LG

Feature to Dynamics: Feature-space to Autoregression strategy for Zero-shot Time Series Forecasting

Zero-shot time series forecasting aims to predict future values for previously unseen series, requiring models to generalize temporal dynamics beyond the training distribution. While recent foundation models achieve strong in-domain performance through large-scale pretraining, their effectiveness often relies on broad data coverage and implicit pattern memorization, which can limit generalization when data are scarce or source and target domains are disjoint. In this work, we propose FSA, a feature-to-strategy framework for controlled zero-shot univariate forecasting. Instead of directly modeling raw sequences in the observation space, FSA learns a structured mapping from an interpretable feature space to an autoregressive strategy space. This design introduces explicit inductive biases that disentangle global trends, periodic components, and local temporal dynamics, enabling the model to capture transferable time-series structure with fewer data assumptions. Empirical results show that, under identical pretraining data, training protocol, and comparable parameter budgets, FSA outperforms Transformer-based architectures in our controlled zero-shot setting.

cs.LG

ReactBench: A Cause-Driven Benchmark for Multimodal Hallucination via Systematic Evaluation

While multimodal large language models (MLLMs) have achieved rapid progress in vision-language understanding, they remain prone to multimodal hallucinations, producing responses that are inconsistent with the visual input. Existing benchmarks predominantly focus on detecting hallucination outcomes rather than evaluating the underlying causes of these failures. Moreover, many benchmarks rely on simplistic scenarios and limited evaluation formats that no longer challenge state-of-the-art models. To address these limitations, we introduce ReactBench, a cause-driven hallucination benchmark featuring multiple tasks and an exam-style evaluation format. By generating adversarial images and hallucination-inducing queries, ReactBench introduces four targeted tasks: Relational Erasure, Counterfactual Attribute, Alteration Tracing, and Dense Counting. These tasks systematically expose co-occurrence bias, language priors, cross-image comparative perception deficiencies, and fine-grained perceptual bottlenecks. Beyond standard accuracy-based evaluation, we leverage Chain-of-Thought reasoning to identify fine-grained sub-causes of hallucination within each task. Extensive evaluations reveal that current MLLMs remain notably vulnerable to cause-specific hallucination triggers, demonstrating the value of ReactBench as a systematic and interpretable testbed for diagnosing and improving multimodal model robustness. The project page is available at https://reactbench.github.io/.

cs.CV

Posterior-Aware Differential Channel Tracking for Reliable Single-Stream DAB+ Passive Radar

Digital audio broadcasting plus (DAB+) is an attractive illuminator for passive radar because it provides persistent, high-power, and geographically widespread very high frequency (VHF) orthogonal frequency-division multiplexing (OFDM) signals. A channel state information (CSI) sensing approach can convert a single received DAB+ stream into a CSI sequence for radar sensing, avoiding the need for a separately received reference signal in conventional passive radars. However, CSI estimation in DAB+ is challenging due to the differentially encoded communication symbols across time. A wrong symbol transition estimation leads to a persistent multiplicative error in the sequential CSI sequence within a DAB+ frame. This paper formulates single-stream DAB+ passive radar as a posterior-probability-aware differential CSI tracking problem. The proposed method uses the previously tracked CSI as a channel prior, performs prediction-aided maximum a posteriori detection of current symbol, converts posterior transition reliability into observation uncertainty, and applies linear minimum mean squared error fusion to obtain a stable tracking CSI. A reliability-informed CSI fusion strategy is also introduced to preserve weak target information. Theoretical analysis is provided, showing guaranteed performance again in symbol and CSI estimation. Simulation results show that the proposed method can reduce CSI estimation error by over 15~dB compared with prior art. It also improves median target-to-background ratio by more than 11~dB in random fading scenes. Experiments in Sydney, Australia demonstrate improved range-Doppler maps for commercial aircraft sensing.

eess.SP

Learning from Medical Entity Trees: An Entity-Centric Medical Data Engineering Framework for MLLMs

Multimodal Large Language Models (MLLMs) have shown transformative potential in medical applications, yet their performance is hindered by conventional data curation strategies that rely on coarse-grained partitioning by modality or department. Such fragmented approaches fail to capture the hierarchical and interconnected nature of clinical medical knowledge, limiting the models' ability to perform fine-grained recognition and complex reasoning. In this paper, we propose a novel Entity-Centric Medical Data Engineering framework. We automatically extract entities from authoritative medical literature to construct a Medical Entity Tree (MET), a hierarchical structure that systematically encodes diseases, anatomical structures, modalities, and symptoms into a unified knowledge repository. Building upon the MET, we propose an advanced data engine that includes: (1) node-guided retrieval to anchor raw data to specific medical concepts, (2) a two-stage hybrid filtering and alignment pipeline to ensure precise visual-semantic correspondence, and (3) knowledge-aware data synthesis to generate enriched captions and targeted reasoning VQA pairs, leveraging structural constraints. Extensive evaluations across six medical benchmarks demonstrate that our approach significantly enhances the medical capabilities of general-purpose MLLMs, improving their ability to handle complex clinical queries and achieve state-of-the-art performance in diverse medical contexts.

cs.CL

Excitons in WSe2 time-resolved ARPES: particle or oscillation?

The time-resolved angle-resolved photoemission spectra of WSe$_2$, a paradigmatic transition metal dichalcogenide, are dominated by a transient signal that, after being initially observed in the gap at the K valley, scatters, on an ultra-fast time scale of $\sim$ 30 fs, to the $\Sigma$ valley. In this work we question the common interpretation of the experimental dynamics in terms of a massive bound electron-hole exciton that scatters with phonons and behaves as a quasi-particle. By using a combined theoretical and experimental investigation, we demonstrate that the observed dynamics can be interpreted as the photo-induced transition from direct to indirect excitonic-insulating order. The features that appear in the experimental spectrum correspond to single-particle levels renormalized by the excitonic spontaneous polarization.

cond-mat.mtrl-sci

Direct N-body simulations of rotating and extremely massive Population III star clusters

Aims. We present eight direct N-body simulations with NBODY6++GPU of extremely massive, initially rotating Population III star clusters with 1.01 x 10^5 stars. Methods. Our models include primordial binaries, a continuous initial mass function, differential rotation, tidal mass loss, updated fitting formulae for extremely massive metal-poor Population III stars, and general-relativistic merger recoil kicks. We assess their impact on cluster dynamics. Results. All runs form black holes below, within, and above the pair-instability gap, with multi-generation growth. Faster-rotating clusters core-collapse earlier; post-collapse clusters host a rotating, axisymmetric subsystem of intermediate-mass black holes (IMBHs) at the centre and an expanding halo of lower-mass objects. Pair-instability supernovae and compact-object formation at ~2-3 Myr sharply reduce total mass and a large fraction of the cluster's angular momentum. All Population III clusters in our simulations have the gravothermal-gravogyro catastrophe phase. Conclusions. We confirm two of the hypothesized formation channels of galactic nucleus seed black holes: gravitational runaway mergers of black holes and of Population III stars, which core-collapse into IMBHs thereafter. A higher initial star cluster bulk rotation correlates with earlier core collapse and, in the event counts reported here, with more coalescences and collisions, as well as lower retained (compact) binary abundances. Initial bulk rotation is a primary control parameter of cluster evolution: faster rotation accelerates early angular-momentum transport, gravothermal collapse, mass segregation, and amplifies post-collapse expansion, which also favours the formation of a compact central IMBH subsystem.

astro-ph.GA

Can LLMs Fool Graph Learning? Exploring Universal Adversarial Attacks on Text-Attributed Graphs

Text-attributed graphs (TAGs) enhance graph learning by integrating rich textual semantics and topological context for each node. While boosting expressiveness, they also expose new vulnerabilities in graph learning through text-based adversarial surfaces. Recent advances leverage diverse backbones, such as graph neural networks (GNNs) and pre-trained language models (PLMs), to capture both structural and textual information in TAGs. This diversity raises a key question: How can we design universal adversarial attacks that generalize across architectures to assess the security of TAG models? The challenge arises from the stark contrast in how different backbones-GNNs and PLMs-perceive and encode graph patterns, coupled with the fact that many PLMs are only accessible via APIs, limiting attacks to black-box settings. To address this, we propose BadGraph, a novel attack framework that deeply elicits large language models (LLMs) understanding of general graph knowledge to jointly perturb both node topology and textual semantics. Specifically, we design a target influencer retrieval module that leverages graph priors to construct cross-modally aligned attack shortcuts, thereby enabling efficient LLM-based perturbation reasoning. Experiments show that BadGraph achieves universal and effective attacks across GNN- and LLM-based reasoners, with up to a 76.3% performance drop, while theoretical and empirical analyses confirm its stealthy yet interpretable nature.

cs.AI

Uplink Networked Sensing via Multiuser Correlation Exploitation

In this correspondence, we investigate networked sensing in perceptive mobile networks under a bistatic multi-transmitter single-receiver uplink topology, where multiple user equipments (UEs) transmit signals over orthogonal frequency-division multiple access (OFDMA) resources and a single base station performs joint sensing. Uplink clock asynchronism introduces offsets that destroy inter-packet coherence and hinder high-resolution sensing, while multi-user observations exhibit exploitable cross-user correlation. We therefore formulate an asynchronous multi-user uplink OFDMA sensing model and exploit common delay-cluster sparsity across UEs. A line-of-sight (LoS)-referenced calibration first suppresses the offsets, after which a shared-private delay-domain sparse Bayesian learning (SBL) model is used for delay support recovery and user grouping. Doppler and angle of arrival are then estimated from temporal and spatial phase differences. Simulation results show that the proposed scheme outperforms per-user processing, particularly under limited subcarrier budgets and in low signal-to-noise ratio (SNR) regimes.

eess.SP

Spherical Latent Motion Prior for Physics-Based Simulated Humanoid Control

Learning motion priors for physics-based humanoid control is an active research topic. Existing approaches mainly include variational autoencoders (VAE) and adversarial motion priors (AMP). VAE introduces information loss, and random latent sampling may sometimes produce invalid behaviors. AMP suffers from mode collapse and struggles to capture diverse motion skills. We present the Spherical Latent Motion Prior (SLMP), a two-stage method for learning motion priors. In the first stage, we train a high-quality motion tracking controller. In the second stage, we distill the tracking controller into a spherical latent space. A combination of distillation, a discriminator, and a discriminator-guided local semantic consistency constraint shapes a structured latent action space, allowing stable random sampling without information loss. To evaluate SLMP, we collect a two-hour human combat motion capture dataset and show that SLMP preserves fine motion detail without information loss, and random sampling yields semantically valid and stable behaviors. When applied to a two-agent physics-based combat task, SLMP produces human-like and physically plausible combat behaviors only using simple rule-based rewards. Furthermore, SLMP generalizes across different humanoid robot morphologies, demonstrating its transferability beyond a single simulated avatar.

cs.RO

Iterative Closed-Loop Motion Synthesis for Scaling the Capabilities of Humanoid Control

Physics-based humanoid control relies on training with motion datasets that have diverse data distributions. However, the fixed difficulty distribution of datasets limits the performance ceiling of the trained control policies. Additionally, the method of acquiring high-quality data through professional motion capture systems is constrained by costs, making it difficult to achieve large-scale scalability. To address these issues, we propose a closed-loop automated motion data generation and iterative framework. It can generate high-quality motion data with rich action semantics, including martial arts, dance, combat, sports, gymnastics, and more. Furthermore, our framework enables difficulty iteration of policies and data through physical metrics and objective evaluations, allowing the trained tracker to break through its original difficulty limits. On the PHC single-primitive tracker, using only approximately 1/10 of the AMASS dataset size, the average failure rate on the test set (2201 clips) is reduced by 45% compared to the baseline. Finally, we conduct comprehensive ablation and comparative experiments to highlight the rationality and advantages of our framework.

cs.RO