Search arXivSearch

arXiv subjects

Li Yu

Publications and source records attributed to Li Yu.

At least 19 recordsLinked to original sources

Weighted Homology and Cohomology of Weighted Polyhedra

We define the notion of weighted polyhedron which can be thought of as the geometric realization of a weighted simplicial complex introduced by Dawson. Moreover, we will define a weighted version of singular homology theory for a weighted polyhedron and prove that it is isomorphic to the weighted simplicial homology of the weighted polyhedron. This implies that weighted simplicial homology is an invariant under isomorphisms and more generally under certain type of homotopy equivalences of weighted polyhedra. Moreover, we will generalize the cup product and cap product to weighted singular cohomology. In addition, we will interpret some known theories of orbifolds in terms of our weighted singular homology and cohomology.

math.AT

Electromagnetic World Model for 6G: A Unified Framework for Joint Environment Reconstruction and Channel Prediction

The integration of sensing, communication, and intelligence is becoming a key enabler for sixth generation (6G) wireless systems, where intelligent terminals are expected to simultaneously support efficient link establishment and reliable environmental sensing. However, existing studies mainly exploit sensing information or communication information to address a single task, such as channel prediction or environment reconstruction. Motivated by the shared dependence of optical and radio-frequency signals on the surrounding environment, we propose the electromagnetic world model (EMWM), the first unified framework for joint environment reconstruction and channel prediction. EMWM learns a common electromagnetic representation with the potential to provide a modeling foundation for 6G tasks. Specifically, partial channel state information (CSI) and multi-view red-green-blue (RGB) images are encoded into CSI and visual tokens and jointly processed by a hierarchical world-model backbone with local and global aggregation. Based on the learned representation, a mixture-of-experts (MoE)-based CSI prediction head reconstructs the complete CSI, while a depth prediction head estimates multi-view depth maps that are further converted into three-dimensional (3D) point clouds. Moreover, a large-scale multi-modal dataset is constructed based on a campus digital twin. Experimental results show that EMWM outperforms conventional neural network and large language model (LLM) baselines in both CSI prediction and environment reconstruction, achieving a squared generalized cosine similarity (SGCS) of 0.9699 for CSI prediction while demonstrating robustness across different signal-to-noise ratio (SNR) conditions and zero-shot generalization at 28 GHz.

eess.SP

Adaptive Supervised Anchoring for On-Policy Self-Distillation

On-policy self-distillation (OPSD) adapts a language model by distilling guidance from a frozen teacher on trajectories sampled from the student. Its effectiveness, however, depends critically on the quality of those trajectories. We show that when student rollouts drift from target trajectories, conditioning the teacher on off-target prefixes substantially weakens its task-relevant supervision. Controlled prefix-corruption experiments expose this failure mode, which we term rollout-conditioned signal degradation. To address this problem, we propose a unified training framework that separates two complementary supervision pathways. The first retains rollout-conditioned distribution matching, providing guidance on states the student actually visits. The second applies supervised cross-entropy on canonical ground-truth contexts, avoiding the incompatibility of imposing target tokens on erroneous rollout prefixes. Token-level rollout-target alignment is used to adapt the strength of the canonical-context anchor, emphasizing it during cold start and relaxing it as rollout quality improves. Experiments across multiple model scales, two task families, and general-reasoning benchmarks show that the proposed approach improves task acquisition over OPSD while preserving general capabilities, resulting in a more favorable empirical plasticity-stability trade-off. These findings identify context quality as a central bottleneck in on-policy self-distillation and demonstrate the value of separating rollout-conditioned guidance from canonical supervision.

cs.LG

Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance

The performance bottleneck of agents is increasingly shifting from model capability to the robustness of their execution processes. Tools play a central role as the primary interface through which agents interact with external environments, yet existing methods rarely focus on ensuring robust tool use across diverse runtime conditions. To address this problem, we propose ExpG, a mechanism that builds and refines adaptive guidance capturing each tool's capability boundaries and best practices, thereby enabling agents to use tools more robustly and effectively. ExpG consists of three phases: (1) experience acquisition, which analyzes tool invocation quality from historical execution trajectories, producing structured learnable experiences through multi-aspect attribution; (2) experience distillation, which keeps the experience pool effective by filtering unhelpful experiences, selecting representative ones with an equivalence-class-based method, and summarizing them into generalizable guidance; and (3) experience reuse, which applies the guidance adaptively during future task solving. Extensive experiments show that ExpG brings consistent improvements across the tool selection, tool calling, and response generation tasks, enabling smaller agents to outperform larger ones that do not use ExpG. Moreover, ExpG achieves particularly strong gains in challenging settings, suggesting a promising path toward more robust tool use. Our code, experiments, and results are available.

cs.AI

MVLA-GR: A Phase-Free Multipath-Based Geometry Reconstruction Method via Multi-View Likelihood Accumulation for ISAC

Integrated sensing and communication (ISAC) enables wireless systems to reuse communication signals for environmental sensing, where reconstructing the geometry of surrounding objects is a representative sensing task. However, many conventional methods rely on coherent processing and require accurate phase information, which is often hard to guarantee in practical communication systems, particularly at high carrier frequencies. To address this problem, this paper proposes a Multi-View Likelihood Accumulation Geometry Reconstruction (MVLA-GR) method based on channel impulse response (CIR) measurements, which uses only delay and power observations without requiring phase information. The method extracts dominant multipath components from each observation, and for each candidate spatial location, accumulates components across views whose propagation distances match the location as supporting evidence. A soft distance-matching kernel is introduced to tolerate range estimation errors and viewpoint-dependent scattering migration, and the received power of each component is used as a reliability weight. A joint thresholding strategy combining response magnitude and angular support continuity then converts the continuous support map into a binary geometry estimate. Ray-tracing simulations on canonical and complex targets, as well as real-world vehicle measurements at 36 GHz, demonstrate that MVLA-GR can effectively recover target geometry, providing a low-complexity phase-free solution for ISAC.

eess.SP

DeepRT Engine: A Unified GPU-Parallel Ray-Tracing Framework with Hybrid SBR-IM Path Search for 6G Digital Twin Channel

Digital twin channel (DTC) aims to establish a real-time digital counterpart of physical wireless channels for reproducing and predicting site-specific propagation characteristics. As a high-precision channel computation method for realistic propagation scenarios, ray tracing (RT) serves as a key enabler for DTC construction. However, conventional RT suffers from high complexity under serial path-searching workflows. This letter proposes DeepRT Engine (DeepRT-E), a parallel RT acceleration architecture with a three-stage physically-inspired pipeline for real-time DTC construction. Firstly, DeepRT-E constructs a bounding volume hierarchy (BVH) to partition the scene and reduce redundant ray-surface intersections. Secondly, the shooting and bouncing rays (SBR) algorithm is executed through a ray-level parallel tracing framework to identify candidate surface sequences and prune the search space of the image method (IM). Finally, a parallel batched IM solver refines the retained candidates for accurate propagation-path recovery. Simulation results show that DeepRT-E reduces runtime by 96.3% and achieves a converged error of only 0.001 dB, outperforming Wireless InSite and Sionna in efficiency and accuracy.

eess.SP

The Wigner function for Integer quantum Hall effect

Wigner's quasi-probability distribution function in phase space is a specialized representation of the density matrix, possessing significant physical importance. In this article, we first review the wave function describing electronic motion in an electromagnetic field under the Landau gauge. Next, based on an introduction to the properties of the Wigner function, we calculate the Wigner function for the integer quantum Hall effect using the integral method.

cond-mat.mes-hall

CMSL: Constructive Multi-Sequence Learning for Recommendation Systems

Sequence learning has emerged as the promising paradigm in recommendation systems, surpassing traditional Deep Learning Recommendation Models (DLRM) by capturing the temporal nuances of user behavior. However, current state-of-the-art architectures operate under a limiting analogy: they treat user history as a monolithic chronological sequence like a sentence in a Large Language Model (LLM). We observe a fundamental divergence between natural language and recommendation data: unlike the linear, logical flow of text, user history is inherently multi-faceted. A user's journey is a fragmented reflection of diverse interests, resulting in much weaker coherence between items than is found in LLM training data. This lack of structural unity leads to context pollution. In single-sequence modeling, unrelated behaviors compete for the same attention budget. This "noisy" signal dilutes the model's focus, effectively capping its ability to discern high-intent patterns from background activity. To address this, we propose Constructive Multi-Sequence Learning (CMSL), a paradigm shift from passive sequence ingestion to active "context engineering" that constructs multiple coherent sequences in latent space. CMSL leverages a learnable Sequence Construction Module to disentangle user history into "pure" thematic strands, followed by a linear attention mechanism to efficiently model these strands at scale. CMSL has been deployed across ranking and retrieval tasks and across four major surfaces at Meta.

cs.IR

WiWorld-RealData: A Real-World Multi-Modal Dataset for 6G Wireless World Models

As sixth-generation wireless systems evolve from reliable connectivity toward environment intelligence, wireless world models aim to learn how physical environments and user states affect wireless propagation, requiring real-world data with explicit correspondences between channel responses and environment observations. However, existing channel-environment datasets are predominantly simulation-based or designed for specific communication tasks, limiting their support for general environment-channel relationship learning. To address this gap, we construct WiWorld-RealData, a real-world multi-band channel and multi-modal environment sensing dataset for 6G wireless world model research. It provides synchronized channel impulse responses measured at 3.7 and 6.775 GHz together with multi-view and panoramic images, light detection and ranging point clouds, millimeter-wave radar observations, and global navigation satellite system trajectories. Unified timestamps, sample identifiers, and metadata establish sample-level correspondences across these heterogeneous modalities. The overall measurement campaign produced approximately 10 TB of data, while the current public release provides aligned channel-environment samples from a representative continuous outdoor route. A path-loss prediction case study further validates the dataset using a continuous test route segment, achieving a mean absolute error of 2.02 dB and a root mean square error of 2.69 dB under few-shot adaptation. WiWorld-RealData supports cross-band propagation analysis, environment-aware channel modeling, wireless digital twins, and channel foundation model research. The dataset is available at https://scc.bupt.edu.cn/dataset-manage/datasets/44 and https://doi.org/10.57760/sciencedb.40663.

eess.SP

RankGraph-2: Lifecycle Co-Design for Billion-Node Graph Learning in Recommendation

Graph-based retrieval at billion-node scale requires jointly solving three tightly coupled problems -- graph construction, representation learning, and real-time serving -- yet existing work addresses each in isolation. We present RankGraph-2, a framework deployed at Meta that co-designs all three lifecycle stages for similarity-based retrieval (U2U2I and U2I2I), where each stage's requirements shape the others. Serving requires a co-learned cluster index to avoid expensive online KNN -- this pushes index co-training into the training objective. Training benefits from the observation that similarity-based retrieval tolerates pre-computed neighborhoods, eliminating online graph infrastructure -- this requires construction to produce self-contained data. Construction must also support hour-level refresh for item coverage. Acting on these cascading requirements, RankGraph-2 reduces hundreds of trillions of edges to hundreds of billions via subsampling with popularity bias correction, pre-computes multi-hop neighborhoods via personalized PageRank, and co-learns a residual-quantization cluster index that reduces serving computational cost by 83%. This lifecycle co-design enables a simple architecture to achieve 3.8 x higher recall than a GAT + Deep Graph Infomax model on a bipartite graph and 2.1 x higher than PyTorch-BigGraph on item retrieval. RankGraph-2 delivers up to +0.96% CTR and +2.75% CVR, and has powered 20+ retrieval launches across major surfaces.

cs.IR

MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model

Vision-language-action (VLA) models increasingly condition robot policies on history, depth, or 4D features to resolve ambiguity in long-horizon manipulation. However, more spatiotemporal evidence is not necessarily better: when the injected evidence is not motion-consistent, it can introduce geometric drift, fragmented temporal cues, and unstable action generation. This raises a simple question: should a VLA remember past frames, or remember the motion that connects them? We introduce MotionVLA, a motion-history interface that converts a short past-only video window into compact, time-continuous trajectory-field tokens. Instead of treating history as a sparse set of ndependently lifted frames, MotionVLA represents recent observations as physically coherent motion evidence. Current visual tokens query this history to retrieve task-relevant motion information, which is then recoupled into the VLA stream under trajectory-grounded supervision. Experiments across simulation benchmarks and preliminary real-robot rollouts show that MotionVLA improves long-horizon manipulation while producing smoother and more direct executions. These results suggest that effective VLA memory is not just about providing more 4D context, but about exposing motion-consistent evidence that is usable for control.

cs.RO

Electromagnetic Digital Twin-Enabled Closed-Loop Beam Management in ISAC Systems

Digital twin (DT) is envisioned as a key enabler of sixth-generation (6G) communication systems, evolving from offline descriptive replicas for monitoring and analysis to inthe-loop agents within digital twin networks (DTNs) that couple physical and digital worlds. Recent advances in integrated sensing and communication (ISAC)-driven electromagnetic (EM) scattering methods enable environment twinning by linking channel behaviors to EM properties of the scatterers, supporting interpretable DT states and EM-grounded optimization. However, existing studies primarily focus on DT construction and lack mechanisms for closed-loop control in wireless systems. Moreover, array-geometry mismatch can bias DT reconstruction and degrade control performance, while prior works assume known arrays. To address these gaps, we propose an EM-ISACbased closed-loop DTN framework with a hierarchical design integrating environment twinning, prior injection, and control decision into an end-to-end loop. Leveraging ISAC measurements, the proposed framework jointly reconstructs scatterer information and array-dependent forward operator and employs a low-complexity Bayesian message-passing algorithm to perform contrast inference and array calibration. The reconstructed DT guides codebook preselection to reduce training overhead and narrow candidate beams. Subsequently, downlink beamforming (BF) is performed based on DT-predicted channels, enabling latency-bounded closed-loop control. Simulation results demonstrate improved robustness and control performance under array mismatch.

eess.SP

ChannelAgent-Empowered Electromagnetic Space World Model: A Case Study on Agent-Driven Channel Generation for 6G AI-Native Air Interface

As sixth-generation (6G) wireless networks evolve toward increasingly heterogeneous scenarios, tasks, and service requirements, conventional artificial intelligence (AI) models remain limited in task-aware decision-making and autonomous adaptation. To address this issue, this paper first proposes a ChannelAgent-empowered electromagnetic space world model, in which wireless intelligence is organized into a closed-loop process consisting of multi-modal sensing, ChannelAgent as the intelligent core, and execution with feedback update. As a case study, agent-driven channel generation is instantiated through path loss prediction. Specifically, a task-oriented intelligent feature selection mechanism is designed by integrating reinforcement-learning-inspired policy adaptation with evolutionary search, enabling the agent to iteratively derive compact and task-suitable feature subsets according to the current scenario and performance feedback. Simulation results demonstrate superior performance in both single-scenario and multi-scenario tasks, highlighting the potential of the proposed model for autonomous, adaptive, task-oriented, and closed-loop wireless intelligence.

eess.SP

OP4KSR: One-Step Patch-Free 4K Super-Resolution with Periodic Artifact Suppression

Diffusion-based real-world image super-resolution (Real-ISR) has achieved remarkable perceptual quality; however, directly super-resolving images to 4K remains limited by extreme memory consumption. Consequently, prior methods adopt patch-based inference, sacrificing global context and introducing semantic confusion, spatial inconsistency, and severe latency. We propose OP4KSR, a one-step patch-free 4K SR approach built upon the powerful Flux backbone. By leveraging the extreme-compression F16 VAE, OP4KSR makes 4K SR inference tractable under practical GPU budgets, preserving global spatial-semantic coherence while enabling highly efficient inference. However, adapting this one-step architecture intrinsically triggers severe periodic artifacts. We trace this to a RoPE base frequency allocation mismatch and intra-token spatial ambiguity, both exacerbated by the lack of iterative refinement. To suppress these artifacts, we couple RoPE base frequency rescaling (RFR) with an autocorrelation-based periodicity loss ($\mathcal{L}_\text{AP}$). Furthermore, we curate a dedicated training dataset alongside three benchmarks (one synthetic and two real-world) to advance 4K SR research. Extensive experiments demonstrate that OP4KSR achieves competitive perceptual quality with efficient inference, generating a $4096\times4096$ output in only 5.75 seconds on a single NVIDIA H20 GPU.

cs.CV

Median-of-Means for Nash Equilibrium Seeking in Heavy-Tailed Games

This paper studies the Nash equilibrium seeking problem for stochastic games under heavy-tailed noise. The gradient noise is considered to have a finite $\delta$-th moment ($1<\delta\le 2$), which generalizes the Gaussian noise and covers cases with infinite variance. In this work, we employ the classic method Median-of-Means (MoM) in robust estimation. MoM works by dividing samples into blocks, taking the average of each block, and then taking the median of these block averages, achieving a breakdown point of up to $1/2$. This makes the final estimate reliable even when some samples are very noisy or wrong, and thus is effective to handle the heavy-tailed noise. The method also naturally defends against malicious gradient attacks. Compared with gradient clipping, which is the most popular method to deal with the heavy-tailed noise, MoM requires no preset clipping threshold and is insensitive to the tail behavior of the noise. Under standard assumptions, we prove the almost sure convergence of the algorithm and derive its almost sure convergence rate. To address the systematic bias caused by asymmetric noise, we further design an online bias correction strategy. Simulation results show the effectiveness and efficiency of the proposed algorithms.

math.OC

Paradigm Shift from Statistical Channel Modeling to Digital Twin Prediction: An Environment-Generalizable ChannelLM for 6G AI-enabled Air Interface

As 6G advances, ubiquitous connectivity and higher capacity requirements of the air interface pose substantial challenges for accurate and real-time wireless channel acquisition in diverse environments. Conventional statistical channel modeling relies on offline measurement data from limited environments, struggling to support online applications facing diverse environments. To this end, the digital twin channel (DTC) has emerged as a novel paradigm that constructs a digital replica of the physical environment through high-fidelity sensing and predicts corresponding channel in real time utilizing artificial intelligence (AI) models. As the engine of DTC, existing AI models struggle to simultaneously achieve strong environmental generalization in real-world and end-to-end channel prediction for real time tasks. Therefore, this paper proposes a channel large model (ChannelLM)-driven DTC architecture comprising three modules: low-complexity and high-accuracy environment reconstruction based on dynamic object detection and multimodal alignment of image and point cloud data, physically interpretable environment feature extraction, and a ChannelLM core to mapping these features into generalized environment representations for multi-task channel prediction. Simulation results demonstrate that, in unseen test environments, compared with small-scale AI models, ChannelLM reduces prediction errors by 4.23 dB in channel state information prediction while achieving an end-to-end inference latency of 70 milliseconds in the real world.

eess.SP

TF-SSD: A Strong Pipeline via Synergic Mask Filter for Training-free Co-salient Object Detection

Co-salient Object Detection (CoSOD) aims to segment salient objects that consistently appear across a group of related images. Despite the notable progress achieved by recent training-based approaches, they still remain constrained by the closed-set datasets and exhibit limited generalization. However, few studies explore the potential of Vision Foundation Models (VFMs) to address CoSOD, which demonstrate a strong generalized ability and robust saliency understanding. In this paper, we investigate and leverage VFMs for CoSOD, and further propose a novel training-free method, TF-SSD, through the synergy between SAM and DINO. Specifically, we first utilize SAM to generate comprehensive raw proposals, which serve as a candidate mask pool. Then, we introduce a quality mask generator to filter out redundant masks, thereby acquiring a refined mask set. Since this generator is built upon SAM, it inherently lacks semantic understanding of saliency. To this end, we adopt an intra-image saliency filter that employs DINO's attention maps to identify visually salient masks within individual images. Moreover, to extend saliency understanding across group images, we propose an inter-image prototype selector, which computes similarity scores among cross-image prototypes to select masks with the highest score. These selected masks serve as final predictions for CoSOD. Extensive experiments show that our TF-SSD outperforms existing methods (e.g., 13.7\% gains over the recent training-free method). Codes are available at https://github.com/hzz-yy/TF-SSD.

cs.CV

On Stanley-Reisner Rings with Minimal Betti Numbers

We classify simplicial complexes with a given number of vertices whose Stanley-Reisner ring has the minimal sum of Betti numbers. The Betti numbers of the Stanley-Reisner rings of such kind of simplicial complexes are given by the binomial coefficients. We obtain a topological characterization of these simplicial complexes K by showing that any full subcomplex of K is homotopy equivalent either to a point or to a sphere. Moreover, we prove that such kind of simplicial complexes coincide with simplicial complexes having a minimal Taylor resolution. This allows us to characterize these simplicial complexes in a purely combinatorial way.

math.AC