Search arXivSearch

arXiv subjects

Jun Wu

Publications and source records attributed to Jun Wu.

At least 19 recordsLinked to original sources

UniPoint: Unified Point-Level Sensor Fusion for Humanoid Locomotion Across Challenging Terrains

Open-world deployment requires humanoid robots to cross highly heterogeneous terrain safely, with perception that simultaneously provides wide coverage, local accuracy, and redundancy against sensor failure. Existing approaches struggle to satisfy all three: one forward depth camera or nearby height sampling covers too little; odometry-corrected elevation maps drift under aggressive motion and miss thin vertical structures; image-level encoding costs grow with camera count. We present UniPoint, a humanoid whole-body locomotion framework built on multi-source point-level sensor fusion. Measurements from a 360° light detection and ranging (LiDAR) sensor and two depth cameras are early-fused into one base-frame point set. Voxelization resamples it to a fixed number of tokens encoded by linear self-attention and proprioception-queried cross-attention, decoupling forward cost from sensor count. The point set retains standing thin barriers; a single-modality failure removes only part of the tokens, so the policy degrades gracefully. A single training run with terrain-aware rewards, perception-degradation injection, and domain randomization produces one policy for all eight terrain types, deployed on an onboard RK3588 without fine-tuning. On a DR02 humanoid, 20 trials at each of nine real-world settings over seven terrain types validate the policy on 70-cm-high platforms, 100-cm gaps, thin barriers, and sparse or narrow footholds; it also generalizes zero-shot outdoors.

cs.RO

SoK: Intent-Oriented Systematization of Multi-Turn LLM Jailbreaks

Large Language Models (LLMs) are increasingly deployed in interactive settings, where user intent commonly unfolds through multi-turn dialogue. Multi-turn jailbreaks exploit this pattern by advancing a harmful intent across turns, so that no single message exposes the full objective. However, existing work treats these attacks as a loose collection of prompt patterns and does not analyze how the adversary organizes and advances harmful intent across an interaction. We develop a four-part, intent-oriented taxonomy that organizes multi-turn jailbreaks by adversarial intent structure. Through controlled ablations, we find that effectiveness is driven by how deliberately intent is organized across turns rather than by context length or query count. We further show that the way intent is organized determines the level at which it becomes detectable, pushing the required detection surface outward from the turn level to the session level to the cross-session level. These findings indicate that turn-local safety mechanisms are structurally insufficient and that single-point evaluation overlooks how intent is organized, motivating evaluation protocols aligned to the level at which harmful intent becomes observable. The code is available at: https://github.com/SiyuanLi00/INTACT.

cs.CR

Connecting heterogeneous dynamics with local entropy

Establishing a robust and physically interpretable link between static structure and heterogeneous relaxation dynamics remains a fundamental challenge in glass physics. Here, we introduce a weighted pair-entropy descriptor based on the conventional two-body excess entropy. For this, we multiply the integrand used to calculate the excess entropy by a weight function that is directly related to the length scale of the pair correlation function. This multiplication does not affect the contribution of the short range order to the local excess entropy, but allows to take into account the structure present on intermediate distances, i.e., the medium-range order. For a canonical two-dimensional Lennard-Jones glass former, the resulting descriptor exhibits a strong correlation with particle-level dynamical propensity at long times (multiples of the alpha-relaxation time), with a maximum structure--dynamics correlation reaching about 0.9, substantially outperforming the predictive power of the conventional local pair excess entropy. These results demonstrate that incorporating a physically motivated structural length scale into entropy-based descriptors markedly enhances their predictive power while preserving physical interpretability. Our findings provide a simple and general framework for investigating structure--dynamics correlations in glass-forming systems.

cond-mat.dis-nn

Asymptotic Entanglement Hiding under Stabilizer Restrictions

Entanglement is central to quantum information processing, while stabilizer operations underpin fault-tolerant quantum computation. We ask how much entanglement remains visible or distillable under stabilizer restrictions. We quantify stabilizer-visible entanglement by restricting the measured relative entropy of entanglement to stabilizer measurements, thereby obtaining converse bounds on entanglement distillation under stabilizer operations. We demonstrate magic-free asymptotic entanglement hiding: we construct explicit convex mixtures of pure stabilizer states on $N$ qutrits per party whose unrestricted visible entanglement and LOCC-distillable entanglement both grow as $Ω(N/\log N)$, while their stabilizer-visible and stabilizer-distillable entanglement vanish as $N\to\infty$. Thus, an unbounded amount of LOCC-distillable entanglement carried by stabilizer states can become asymptotically invisible and undistillable under stabilizer restrictions. We further prove that stabilizer-visible entanglement is $O(1)$ with high probability for Haar-random pure states despite extensive unrestricted visibility, and vanishes uniformly over entangled Werner states as the local dimension grows through odd primes. These results reveal a fundamental separation between entanglement and magic as resources, exposing intrinsic limits on entanglement extraction using stabilizer operations.

quant-ph

Toward Intelligent Skies: Signal Processing and AI Foundations of Low-Altitude Wireless Networks

The rapid growth of low-altitude aerial services and applications, driven by uncrewed aerial vehicles (UAVs), calls for a new class of digital infrastructure beyond conventional terrestrial networks. The low-altitude wireless network (LAWN) has been proposed as dynamically reconfigurable three-dimensional architectures that integrate aerial and ground nodes to provide connectivity, sensing, and control in open, safety-critical airspace. This tutorial presents a comprehensive treatment of LAWNs from the joint perspectives of artificial intelligence (AI) and signal processing. We first review the historical evolution and architectural foundations of LAWNs, introducing altitude-based layers and functional planes, and summarizing the regulatory and standardization landscape. Building on this system view, we then discuss signal processing fundamentals for LAWNs, including 3D channel and system models, performance metrics, waveform and receiver design, localization and tracking, and multi-functionality co-design. Next, we survey AI techniques for LAWNs, covering discriminative and generative models for perception, control, resource management, and security, as well as emerging paradigms such as foundation models, large language models, and digital twins for mission planning and closed-loop optimization. To illustrate AI-signal processing integration in practice, we provide a case study of an AI-driven multi-tier LAWN with hybrid satellite, high-altitude, and ground nodes. The tutorial concludes by outlining key research challenges in architecture design, signal processing-AI co-design, safety and security, experimentation, and standardization, and by highlighting opportunities for LAWNs to evolve into dependable, AI-native infrastructure for the intelligent skies.

eess.SP

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks

As LLMs become increasingly integrated into complex applications, their vulnerability to adversarial attacks has raised significant concerns. However, existing defenses remain reactive in nature. This limitation makes it difficult for them to counter sophisticated threats, as adversaries continuously adjust their strategies across multi-turn interactions. In this paper, we present a proactive defense framework for securing LLMs against evolving multi-turn adversarial attacks that combines disruption, misdirection, and adaptation across successive interaction turns. In particular, it employs a cooperative multi-agent architecture in which specialized agents execute complementary defense strategies. These strategies include controlled response pacing to increase attack costs, strategically ambiguous outputs to mislead adversaries into ineffective strategies, and forensic analysis of interaction logs to identify attack patterns and refine defenses. These agents are coordinated by an adaptive mechanism that dynamically adjusts the defense strategy in response to escalating threats. To facilitate comprehensive evaluation, we present the EMRA dataset designed to simulate evolving strategies across multi-turn attacks, including 5,200 adversarial samples across eight attack types. Experimental results on EMRA across multiple LLM backbones show that the proposed framework reduces ASR by 69% on average relative to evaluated state-of-the-art baselines. Beyond suppressing harmful outputs, it sustains deceptive engagement, achieving an average DR more than six times that of the strongest baselines and increasing attacker-token consumption by 198.83% on average relative to evaluated baselines. Code and dataset are available at https://github.com/SiyuanLi00/CoopGuard.

cs.CR

CReF: Cross-modal and Recurrent Fusion for Depth-conditioned Humanoid Locomotion

Stable traversal over geometrically complex terrain increasingly requires exteroceptive perception, yet prior perceptive humanoid locomotion methods often remain tied to explicit geometric abstractions, either by mediating control through robot-centric 2.5D terrain representations or by shaping depth learning with auxiliary geometry-related targets. While effective, these approaches introduce additional map-construction procedures or multi-stage skill-transfer processes beyond direct depth-to-control learning. We propose CReF (Cross-modal and Recurrent Fusion), a single-stage depth-conditioned humanoid locomotion framework that learns locomotion-relevant features directly from raw forward-facing depth without explicit geometric intermediates. CReF couples proprioception and depth tokens through proprioception-queried cross-modal attention, fuses the resulting representation with a gated residual fusion block, and performs temporal integration with a Gated Recurrent Unit (GRU) regulated by a highway-style output gate for state-dependent blending of recurrent and feedforward features. To further improve terrain interaction, we introduce a terrain-aware foothold placement reward that extracts supportable foothold candidates from foot-end point-cloud samples and rewards touchdown locations that lie close to the nearest supportable candidate. Experiments in simulation and on a physical humanoid demonstrate robust traversal over diverse terrains and effective zero-shot transfer to real-world scenes containing handrails, hollow pallet assemblies, severe reflective interference, and visually cluttered outdoor surroundings.

cs.RO

Evidence-in-the-Loop: Trace-Driven Optimization for Customer-Service LLM Agents

Production customer-service bots must improve answer quality across iterative releases, yet large language models must not bypass evidence boundaries, policy rules, or human-handoff safeguards. We present an \textbf{Evidence-Grounded Customer-Service Agent Workflow} deployed in a real-world customer-service setting. BM25 recall, issue-title-vector recall, issue-description-vector recall, weighted RRF fusion, and cross-encoder reranking construct grounded FAQ evidence for controlled LLM decisions. Policy-guided orchestration then combines this RAG evidence with scenario-specific rule evidence, conversation memory, and clarification state inside a fixed LangGraph DAG~\cite{langgraph2024}. The paper contributes three reusable deployment patterns: \textbf{hybrid RAG evidence construction}, where multi-channel retrieval and reranking produce auditable FAQ candidates; \textbf{evidence-grounded issue/action decision}, where an Evidence-Grounded Decision Module selects an issue/action from typed FAQ evidence and scenario-specific rule evidence; and \textbf{trace-driven RAG and reranker improvement}, where traces diagnose whether failures come from recall, ranking, final candidate selection, clarification, rule-derived evidence, or action policy, and where reranker fine-tuning is evaluated not only for in-domain gain but also for forgetting risk.

cs.IR

Identifiability of Partial-Mastery Cognitive Diagnostic Models

Partial-mastery (PM) cognitive diagnostic models (CDMs) extend traditional CDMs by replacing binary latent attribute mastery indicators with continuous mastery scores for multiple latent attributes. In PM-CDMs, each subject is characterized by a fixed continuous latent mastery vector, from which item-specific binary attribute profiles are independently generated. This formulation provides a bridge between classical CDMs and continuous latent variable models. Despite growing interest in PM-CDMs, their identifiability properties remain unexplored. In this work, we establish the first identifiability results for PM-CDMs. We derive sufficient conditions for identifiability that are direct analogues of established conditions for traditional CDMs. To develop the main argument, we use symbolic computation on a minimal example with five items and two latent attributes to show that the Jacobian of the model parameterization is generically nonzero. Combining tools from real analysis and algebraic statistics, we prove that this local property implies generic finite-to-one identifiability of the item parameters and the marginal distributions of the relevant latent attributes. We further show that if the $Q$-matrix contains such identifiable local structures for all attribute pairs, identifiability extends to the full PM-CDM. These findings provide a rigorous theoretical foundation for estimation and inference in partial-mastery cognitive diagnostic models.

math.ST

DINO-SLAM: DINO-informed RGB-D SLAM for Neural Implicit and Explicit Representations

This paper presents DINO-SLAM, a DINO-informed design strategy to enhance implicit (Neural Radiance Field -- NeRF) and explicit representations (Gaussian Splatting -- GS) in SLAM systems through the more comprehensive semantics understanding enabled by DINO. This latter alone, however, lacks proper 3D geometry understanding, allowing only for marginal improvements. Therefore, we rely on a Scene Geometry Encoder (SGE) to enrich DINO features into geometry-aware DINO features (geoDINO), to better understand those geometric relationships that vanilla DINO features fail to capture. Building upon it, we propose two foundational paradigms for NeRF and GS SLAM systems integrating geoDINO features. Compared to state-of-the-art methods, our DINO-informed pipelines achieve superior performance on the Replica, ScanNet, and TUM datasets.

cs.CV

PUMA: Perception-driven Unified Foothold Prior for Mobility Augmented Quadruped Parkour

Parkour tasks for quadrupeds have emerged as a promising benchmark for agile locomotion. While human athletes can effectively perceive environmental characteristics to select appropriate footholds for obstacle traversal, endowing legged robots with similar perceptual reasoning remains a significant challenge. Existing methods often rely on hierarchical controllers that follow pre-computed footholds, thereby constraining the robot's real-time adaptability and the exploratory potential of reinforcement learning. To overcome these challenges, we present PUMA, an end-to-end learning framework that integrates visual perception and foothold priors into a single-stage training process. This approach leverages terrain features to estimate egocentric polar foothold priors, composed of relative distance and heading, guiding the robot in active posture adaptation for parkour tasks. Extensive experiments conducted in simulation and real-world environments across various discrete complex terrains, demonstrate PUMA's exceptional agility and robustness in challenging scenarios.

cs.RO

SpikeTimer: Exploring Active Copyright Protection in Spiking Neural Networks via Temporal Backdoor Regularization

Spiking Neural Networks (SNN) have emerged as a revolutionary paradigm compared to traditional Deep Neural Networks (DNN) in energy-efficient computing, showcasing exceptional capabilities in processing event-driven sensory data for real-time applications like robotics and edge AI systems. However, unlike extensive studies on DNN copyright solutions, SNN copyright protection remains largely underexplored due to their inherent temporal coding complexities and spike-driven computation. In this study, we propose a novel active copyright protection framework named SpikeTimer for SNNs via temporal backdoor learning. SpikeTimer partitions neuromorphic data into designated timeslices and exclusively embeds authorized tokens within authorized slices. Furthermore, the inherent temporal segmentation characteristic intrinsically enables SpikeTimer to support multi-user authorization mechanisms and accommodates token embedding of arbitrary morphology. Based on this, SpikeTimer precisely responds to authorized data containing a token within the correct timeslice, while producing erroneous responses to unauthorized data. Our key innovation lies in establishing a time-dependent authorization mechanism that protects the SNN copyright by temporal token validity. Additionally, SpikeTimer retains its defensive efficacy even under adversarial attempts. Evaluations on multiple neuromorphic datasets manifest that SpikeTimer achieves around 10% accuracy on unauthorized data with merely around 1.5% degradation on authorized inputs. Moreover, SpikeTimer demonstrates robust resistance against model finetuning and pruning threats.

cs.CR

Integrating Sensing into Covert Communications: Opportunities and Challenges

Covert communications aim to hide the existence of wireless transmissions from unauthorized adversaries. However, conventional designs based on blind interference or passive uncertainty can be ineffective in dynamic propagation environments. This article investigates sensing-empowered covert communications, where adversary and environmental information are used to guide transmission and jamming control. We show how sensing changes covert system design from passive concealment to state-aware decision-making, while also introducing new challenges related to exposure and resource consumption. We further discuss several intelligent sensing paradigms that extract task-relevant information with limited active probing. A case study in low-altitude wireless networks illustrates that sensing-assisted beamforming can improve spatial resource utilization and the reliability of covert data delivery in time-varying channels. Finally, several open issues are discussed to support more adaptive covert wireless systems.

eess.SP

LAWNs Meet SWIPT: Beamforming and Power Splitting Optimization for Predictive Control

Simultaneous wireless information and power transfer (SWIPT) has emerged as a promising paradigm for enabling sustainable connectivity in battery-limited low-altitude wireless networks (LAWNs). This paper investigates a SWIPT-enabled LAWN system in which a multi-antenna base station (BS) simultaneously delivers control information and wireless energy to a fleet of uncrewed aircraft systems (UASs) via power splitting. In particular, the BS remotely guides the UASs to accurately track predefined reference trajectories toward their destinations while avoiding multiple mobile no-fly zones (NFZs). To guarantee collision-free path planning, we first construct smooth and safe reference trajectories using stream function theory. Then, a real-time optimization problem is formulated, which jointly takes into account the wireless control cost and energy sustainability by optimizing control inputs, transmit beamforming vectors, and the power splitting ratios. To address the resultant non-convex problem, a two-stage optimization framework is proposed. First, we develop a model predictive control (MPC)-based method to generate predictive control inputs. Subsequently, we derive a computationally efficient iterative algorithm to optimize the beamforming vectors and power splitting ratios by applying semidefinite relaxation (SDR) and successive convex approximation (SCA) techniques. We further prove that the SDR is tight for our formulation. Extensive numerical results demonstrate that our proposed design significantly outperforms benchmark schemes in terms of tracking accuracy and harvested energy, thereby validating its effectiveness for sustainable implementation in LAWN systems.

eess.SY

GUIDE: Goal-Initialized Directional Understanding for End-to-End Visual Navigation

Learning-based visual navigation for legged robots typically relies on continuous goal updates from hierarchical state estimation to provide a persistent directional reference. This reliance incurs additional sensory and computational overhead and deviates from fully end-to-end mobile autonomy. Furthermore, under partial observability, policies are prone to learn myopic behaviors, easily becoming trapped in dead ends and complex structural layouts. To address these limitations, we investigate a goal-initialized navigation setting, where the target is provided only once at the beginning of an episode, requiring the robot to operate based on intrinsic spatial memory without subsequent goal updates from external modules. In this work, we propose GUIDE, a fully end-to-end reinforcement learning framework designed to cultivate internal directional awareness. Specifically, GUIDE incorporates a spatial anchor predictor that leverages multi-frequency proprioceptive history to extract egomotion representations, thereby maintaining a persistent long-horizon spatial context for navigation. Concurrently, it utilizes raw depth streams to perceive local environmental geometry. We evaluate the proposed framework across both simulation and real-world scenarios on a quadruped robot. Experiments show that GUIDE learns reliable egomotion and directional awareness, enabling a fully end-to-end deployed policy to safely navigate through dense clutter and structured mazes without subsequent goal guidance or prior maps.

cs.RO

Cascaded Rydberg antiblockade: Multi-atom excitation dynamics and entanglement

We propose a cascaded Rydberg antiblockade (RAB) regime via a Floquet modulation in four fully connected interacting atoms, which establishes a new synthetic dimension, Dicke-state lattice (DSL), in the space of collective spin excitations. By applying a global periodic driving, we synthesize an effective Hamiltonian that enables perfect state transfer across the five-site DSL with multiple programmable pathways from stepwise nearest-neighbor jumps to a single-step transition. This DSL platform further allows us to simulate a dynamic Su-Schrieffer-Heeger model, where soft quantum control is employed to achieve topologically inspired full RAB $|0000\rangle \to |1111\rangle$ with enhanced robustness against disorder. Moreover, by incorporating the shortcut to adiabaticity technique, we generate high-fidelity entangled twin-Fock and Greenberger-Horne-Zeilinger states on the four atoms within sub-microsecond timescales, outperforming the speed limits of conventional adiabatic protocols. Our work demonstrates a flexible and programmable synthetic dimension for quantum simulation and multipartite entanglement engineering in Rydberg atom arrays, paving the way for the future development of quantum information processing.

quant-ph

Predictive Control over Low-Altitude Wireless Networks: Joint Trajectory Design and Resource Allocation

Low-altitude wireless networks (LAWNs) have been envisioned as flexible and transformative platforms for enabling delay-sensitive control applications in Internet of Things (IoT) systems. In this work, we investigate the real-time wireless control over LAWNs, where an aerial drone is employed to serve multiple mobile automated guided vehicles (AGVs) via finite blocklength (FBL) transmission. Toward this end, we adopt the model predictive control (MPC) to ensure accurate trajectory tracking, while we analyze the communication reliability using the outage probability. Subsequently, we formulate an optimization problem to jointly determine control policy, transmit power allocation, and drone trajectory by accounting for the maximum travel distance and control input constraints. To address the resultant non-convex optimization problem, we first derive the closed-form expression of the outage probability under FBL transmission. Based on this, we reformulate the original problem as a quadratic programming (QP) problem, followed by developing an alternating optimization (AO) framework. Specifically, we employ the projected gradient descent (PGD) method and the successive convex approximation (SCA) technique to achieve computationally efficient sub-optimal solutions. Furthermore, we thoroughly analyze the convergence and computational complexity of the proposed algorithm. Extensive simulations and AirSim-based experiments are conducted to validate the superiority of our proposed approach compared to the baseline schemes in terms of control performance.

eess.SP

Addressing Imbalance in Multi-Label Data via Label-Specific Distance-based Oversampling

The complex imbalanced label distribution poses a crucial challenge to multi-label classification, as most classifiers are biased towards the majority class and high-frequent labels. Oversampling is an efficient and flexible solution that augments instances to provide a more balanced training dataset for multi-label classifiers. Most existing oversampling methods create synthetic instances in a heuristic way that essentially relies on neighborhood information retrieved using Euclidean distance within the entire feature space. However, they fail to consider the varying semantic relevance of features to different labels, leading to label inconsistency among proximate neighbors and further introducing label confusion and overfitting to synthetic instances. To overcome the above issue, we propose a novel sampling approach called Label-Specific Distance-based Multi-Label Oversampling (LSDMLO) that creates more useful and well-labeled synthetic instances to address the imbalance in multi-label datasets. LSDMLO derives the label-specific distance to identify label-consistent neighbors based on the weighted pertinent feature space, which facilitates selecting seed instances that express more label correlations in boundary areas and generating synthetic instances aligned with the label distribution of original data. The comprehensive experiments verify that the proposed LSDMLO outperforms the state-of-the-art multi-label sampling approaches under various base classifiers.

cs.LG