Search arXivSearch

arXiv subjects

Wenjun Yu

Publications and source records attributed to Wenjun Yu.

At least 19 recordsLinked to original sources

Quantum-classical crossover in fault-tolerant quantum dynamics simulation

While quantum computers promise to solve classically intractable problems, identifying the point at which fault-tolerant quantum computation outperforms the best classical algorithms for practical applications remains an outstanding challenge. Here we establish a concrete quantum-classical crossover for quantum many-body dynamics under realistic hardware conditions. We introduce a scalable fault-tolerant framework that combines coherent observable estimation with a space-time-efficient implementation of non-Clifford rotations, suppressing the residual logical errors that limit existing partially fault-tolerant approaches. A benchmark against state-of-the-art tensor-network and variational Monte Carlo algorithms reveals a concrete crossover for mixed-field Ising dynamics at modest system sizes. For a physical error rate of $p=10^{-3}$, fault-tolerant simulation requires approximately 2 hours and $3.7 \times 10^5$ physical qubits for a 100-site 1D system, whereas tensor network approaches would require about 100 years. For 2D models, where rapid entanglement growth limits the classical evolution time, we project quantum runtimes within minutes. A physical error rate of $p=10^{-4}$ leads to at least an order of magnitude reduction in qubit count ($3.1 \times 10^4$ physical qubits) and runtime (minutes for 1D and seconds for 2D). The reduction in quantum runtime arises from our improved rotation-state injection and co-design of quantum error correction and observable-estimation protocols, which jointly suppress logical-error accumulation and reduce sampling overhead. Our results establish a scalable route towards practical quantum advantage and identify quantitative engineering targets for future fault-tolerant architectures.

quant-ph

CFM-Bench: A Unified Multi-Domain, Multi-Task Benchmark for Channel Foundation Models

Channel foundation models (CFMs) are commonly evaluated in model-specific pipelines that differ in data, radio configurations, partitions, adaptation procedures, task definitions, and metrics, preventing reproducible comparison across CFMs and against task-specific networks. We release CFM-Bench, a unified multi-domain, multi-task benchmark comprising 157,900 official single-frame examples from six domains spanning 3GPP statistical simulation, two ray-tracing pipelines, terrestrial and aerial measurements, and synchronized vehicular multimodal simulation. Source-specific interfaces preserve complex channel state information (CSI) and the physical metadata available in each domain while allowing documented model-specific preprocessing. To reduce spatio-temporal leakage, official partitions isolate complete trajectories, measurement sessions, flights, vehicle links, simulation realizations, or buffered spatial regions. CFM-Bench excludes all benchmark splits from foundation-model pretraining, reserves the official training split for downstream fine-tuning, and reports the data used during model development. Six task groups across PHY, RAN, and ISAC applications cover CSI feedback, frequency and temporal channel extrapolation, propagation-state classification, current- and future-beam prediction, and single-frame and temporal localization. Representative experiments on CSI feedback, channel extrapolation, current-beam prediction, and wireless positioning provide reproducible reference results for pretrained channel-prediction models and task-specific neural networks. These results show that relative model performance can vary across data domains, highlighting the importance of using common data partitions, task definitions, and evaluation metrics. CFM-Bench provides a common substrate for evaluating the transferability of channel representations across models, domains, and tasks.

cs.AI

All simplices exhibit canonical Ramsey property

We prove that all nondegenerate simplices have the canonical Ramsey property, thereby resolving a central open problem in canonical Euclidean Ramsey theory and providing a canonical counterpart to the celebrated simplex Ramsey theorem of Frankl and R\"{o}dl~[JAMS, 1990].

math.CO

SQuaD-SQL: Efficient Text-to-SQL with Small Language Models via LLM-Guided Knowledge Distillation

Text-to-SQL is a fundamental task in natural language processing that enables users to interact with structured databases using natural language. While large language models (LLMs) have demonstrated remarkable performance on this task, their substantial computational requirements hinder deployment in resource-constrained settings. In this paper, we introduce SQuaD-SQL (Small-Qualified and Distilled for SQL), a novel approach that empowers small language models (SLMs) to approach the performance of LLMs on the Text-to-SQL task while significantly improving efficiency through knowledge distillation and synthetic data generation. Our method comprises three key components: (1) LLM-based synthetic data generation, where structured knowledge is extracted from LLMs via carefully designed prompting strategies; (2) parameter-efficient fine-tuning, enabling full model training on a single consumer-grade GPU; and (3) domain-adaptive fine-tuning, where domain-specific synthetic data further enhances performance in targeted domains. Experiments on the WikiSQL dataset demonstrate that SQuaD-SQL achieves an execution accuracy of 86.9% on the test set, approaching the performance of LLMs while offering faster inference and lower memory usage. These results suggest that, with proper training strategies, SLMs can serve as practical and efficient alternatives for Text-to-SQL applications in resource-limited environments.

cs.CL

CSI-CLIP++: A Scalable Channel Foundation Model for Wireless Communication via CIR-CSI Consistency

Self-supervised learning can exploit large-scale unlabeled channel data to improve the transferability of wireless AI models. Existing channel foundation models are often built on single-domain representations or reconstruction-oriented objectives, which may not explicitly capture the physical correspondence between frequency- and delay-domain channel views. This paper proposes CSI-CLIP++, a scalable channel foundation model for MIMO wireless channels. CSI-CLIP++ treats frequency-domain channel state information (CSI) and delay-domain channel impulse response (CIR) as paired views of the same propagation process and learns transferable representations through CSI-CIR contrastive alignment. The pretrained CSI encoder is adapted to channel identification, beam prediction, and positioning, representing PHY, RAN, and ISAC applications. Experiments on large-scale DeepMIMO scenarios show consistent gains over supervised baselines across environments, carrier frequencies, and data scales. CSI-CLIP++ improves beam prediction Top-1 accuracy by up to 19.31 percentage points and achieves competitive positioning performance, including cross-simulator transfer on a Sionna RT dataset. Backbone scaling results further show that the proposed objective remains effective across encoder architectures and benefits from larger model capacity.

eess.SP

When KV Meets Embeddings: Dynamic GPU Memory Allocation for Accelerating Generative Recommender Serving

Generative Recommender (GR) inference places embedding hot caches (EMB) and KV caches in direct competition for limited GPU HBM: allocating more memory to one improves its efficiency but degrades the other. Existing systems optimize them in isolation, overlooking that the optimal EMB-KV allocation ratio can shift by up to 0.35 across workload regimes, leaving 20-30\% latency improvement unrealized. While online reallocation is required to close this gap, naive approaches introduce H2D refill traffic on the critical path, causing P99 SLO violations. To address this, we present RACER, which jointly manages HBM allocation and request routing at runtime through two key components: (1) Adaptive Memory Allocation, a three-layer PPO-based controller (frozen base policy, online residual adapter, and burst-aware recovery controller) that achieves $32\,\mathrm{\mu s}$ decision latency while staying within 0.024-0.029 of the offline-optimal ratio; and (2) EMB-KV-Aware Scheduling, which routes requests by jointly considering KV residency, embedding locality, and node load to avoid routing inefficiencies under heterogeneous allocations. Evaluations on three production-scale datasets over a 32-node A100 cluster show that RACER reduces P99 latency by 24-38\% over the best static policy and achieves 93.5-99.6\% SLO satisfaction across Steady, Trend, and Burst workloads, significantly outperforming state-of-the-art baselines without sacrificing throughput.

cs.DC

On the Capacity of Sequences of Coloring Channels

A single coloring channel is defined by a subset of letters it allows to pass through, while deleting all others. A sequence of coloring channels provides multiple views of the same transmitted letter sequence, forming a type of sequence-reconstruction problem useful for protein identification and information storage at the molecular level. We provide exact capacities of several sequences of coloring channels: uniform sunflowers, two arbitrary intersecting sets, and paths. We also show how this capacity depends solely on a related graph we define, called the pairs graph. Using this equivalence, we prove lower and upper bounds on the capacity, and a tailored bound for a coloring-channel sequence forming a cycle. In particular, for an alphabet of size $4$, these results give the exact capacity of all coloring-channel sequences except for a cycle of length $4$, for which we only provide bounds.

cs.IT

Quantum state determinability from local marginals is universally robust

A fundamental problem in quantum physics is to establish whether a multiparticle quantum state can be uniquely determined from its local marginals. In theory, this problem has been addressed in the exact case where the marginals are perfectly known. In practice, however, experiments only have access to finite statistics and therefore can only determine the marginals of a quantum state up to an error. In this Letter, we prove that unique determinability universally survives such local imperfections: specifically, for every uniquely determined state, we show that deviations of local marginals propagate to global states strictly bounded by a power law with exponent $\alpha\in(0,1]$. This result induces a classification of multipartite quantum states by their power-law exponents, with linear scaling $\alpha=1$ as the most favorable regime. We derive a necessary and sufficient criterion for linear robustness and translate it into an executable semidefinite-programming certification. Applying our theory, we prove that stabilizer states are inherently square-root robust and provide a complete robustness classification for the Dicke family. Finally, we exploit these results to construct a scalable two-local genuine multipartite entanglement witness, demonstrating the viability of this framework for broad practical applications.

quant-ph

Entanglement-Induced Resilience of Quantum Dynamics

Quantum many-body devices suffer from imperfections that destabilize dynamics and limit scalability. We show that the dynamical growth of entanglement can intrinsically protect generic quantum dynamics against coherent and perturbative noise. Through rigorous theoretical analysis of general quantum dynamics and numerical simulations of spin chains and fermionic lattices, we prove that entanglement-entropy growth confines the influence of local Hamiltonian perturbations, thereby suppressing errors in dynamical errors. The degree of protection correlates quantitatively with the entanglement entropy of subsystems on which the perturbations act, and applies broadly to both analog quantum simulators and real-time control protocols. This entanglement-induced resilience is conceptually distinct from quantum error correction or dynamical decoupling: it passively leverages native many-body correlations without additional qubits, measurements, or control overhead. Our results reveal a generic mechanism linking entanglement growth to dynamical stability and provide practical guidelines for designing noise-resilient quantum devices.

quant-ph

AI-Driven Channel State Information (CSI) Extrapolation for 6G: Current Situations, Challenges and Future Research

CSI extrapolation is an effective method for acquiring channel state information (CSI), essential for optimizing performance of sixth-generation (6G) communication systems. Traditional channel estimation methods face scalability challenges due to the surging overhead in emerging high-mobility, extremely large-scale multiple-input multiple-output (EL-MIMO), and multi-band systems. CSI extrapolation techniques mitigate these challenges by using partial CSI to infer complete CSI, significantly reducing overhead. Despite growing interest, a comprehensive review of state-of-the-art (SOTA) CSI extrapolation techniques is lacking. This paper addresses this gap by comprehensively reviewing the current status, challenges, and future directions of CSI extrapolation for the first time. Firstly, we analyze the performance metrics specific to CSI extrapolation in 6G, including extrapolation accuracy, adaption to dynamic scenarios and algorithm costs. We then review both model-driven and artificial intelligence (AI)-driven approaches for time, frequency, antenna, and multi-domain CSI extrapolation. Key insights and takeaways from these methods are summarized. Given the promise of AI-driven methods in meeting performance requirements, we also examine the open-source channel datasets and simulators that could be used to train high-performance AI-driven CSI extrapolation models. Finally, we discuss the critical challenges of the existing research and propose perspective research opportunities.

eess.SP

Near-Zero-Overhead Freshness for Recommendation Systems via Inference-Side Model Updates

Deep Learning Recommendation Models (DLRMs) underpin personalized services but face a critical freshness-accuracy tradeoff due to massive parameter synchronization overheads. Production DLRMs deploy decoupled training/inference clusters, where synchronizing petabyte-scale embedding tables (EMTs) causes multi-minute staleness, degrading recommendation quality and revenue. We observe that (1) inference nodes exhibit sustained CPU underutilization (peak <= 20%), and (2) EMT gradients possess intrinsic low-rank structure, enabling compact update representation. We present LiveUpdate, a system that eliminates inter-cluster synchronization by colocating Low-Rank Adaptation (LoRA) trainers within inference nodes. LiveUpdate addresses two core challenges: (1) dynamic rank adaptation via singular value monitoring to constrain memory overhead (<2% of EMTs), and (2) NUMA-aware resource scheduling with hardware-enforced QoS to eliminate update inference contention (P99 latency impact <20ms). Evaluations show LiveUpdate reduces update costs by 2x versus delta-update baselines while achieving higher accuracy within 1-hour windows. By transforming idle inference resources into freshness engines, LiveUpdate delivers online model updates while outperforming state-of-the-art delta-update methods by 0.04% to 0.24% in accuracy.

cs.DC

Vision-Language Semantic Aggregation Leveraging Foundation Model for Generalizable Medical Image Segmentation

Multimodal models have achieved remarkable success in natural image segmentation, yet they often underperform when applied to the medical domain. Through extensive study, we attribute this performance gap to the challenges of multimodal fusion, primarily the significant semantic gap between abstract textual prompts and fine-grained medical visual features, as well as the resulting feature dispersion. To address these issues, we revisit the problem from the perspective of semantic aggregation. Specifically, we propose an Expectation-Maximization (EM) Aggregation mechanism and a Text-Guided Pixel Decoder. The former mitigates feature dispersion by dynamically clustering features into compact semantic centers to enhance cross-modal correspondence. The latter is designed to bridge the semantic gap by leveraging domain-invariant textual knowledge to effectively guide deep visual representations. The synergy between these two mechanisms significantly improves the model's generalization ability. Extensive experiments on public cardiac and fundus datasets demonstrate that our method consistently outperforms existing SOTA approaches across multiple domain generalization benchmarks.

cs.CV

Enhanced Fingerprint-based Positioning With Practical Imperfections: Deep learning-based approaches

High-precision positioning is vital for cellular networks to support innovative applications such as extended reality, unmanned aerial vehicles (UAVs), and industrial Internet of Things (IoT) systems. Existing positioning algorithms using deep learning techniques require vast amounts of labeled data, which are difficult to obtain in real-world cellular environments, and these models often struggle to generalize effectively. To advance cellular positioning techniques, the 2024 Wireless Communication Algorithm Elite Competition as conducted, which provided a dataset from a three-sector outdoor cellular system, incorporating practical challenges such as limited labeled-dataset, dynamic wireless environments within the target and unevenly-spaced anchors, Our team developed three innovative positioning frameworks that swept the top three awards of this competition, namely the semi-supervised framework with consistency, ensemble learning-based algorithm and decoupled mapping heads-based algorithm. Specifically, the semi-supervised framework with consistency effectively generates high-quality pseudo-labels, enlarging the labeled-dataset for model training. The ensemble learning-based algorithm amalgamates the positioning coordinates from models trained under different strategies, effectively combating the dynamic positioning environments. The decoupled mapping heads-based algorithm utilized sector rotation scheme to resolve the uneven-spaced anchor issue. Simulation results demonstrate the superior performance of our proposed positioning algorithms compared to existing benchmarks in terms of the {90%, 80%, 67%, 50%} percentile and mean distance error.

eess.SP

Optimal Reconstruction Codes with Given Reads in Multiple Burst-Substitutions Channels

We study optimal reconstruction codes over the multiple-burst substitution channel. Our main contribution is establishing a trade-off between the error-correction capability of the code, the number of reads used in the reconstruction process, and the decoding list size. We show that over a channel that introduces at most $t$ bursts, we can use a length-$n$ code capable of correcting $\epsilon$ errors, with $\Theta(n^\rho)$ reads, and decoding with a list of size $O(n^\lambda)$, where $t-1=\epsilon+\rho+\lambda$. In the process of proving this, we establish sharp asymptotic bounds on the size of error balls in the burst metric. More precisely, we prove a Johnson-type lower bound via Kahn's Theorem on large matchings in hypergraphs, and an upper bound via a novel variant of Kleitman's Theorem under the burst metric, which might be of independent interest. Beyond this main trade-off, we derive several related results using a variety of combinatorial techniques. In particular, along with tools from recent advances in discrete geometry, we improve the classical Gilbert-Varshamov bound in the asymptotic regime for multiple bursts, and determine the minimum redundancy required for reconstruction codes with polynomially many reads. We also propose an efficient list-reconstruction algorithm that achieves the above guarantees, based on a majority-with-threshold decoding scheme.

cs.IT

Quantum Hamiltonian Certification

We formalize and study the Hamiltonian certification problem. Given access to $e^{-\mathrm{i} Ht}$ for an unknown Hamiltonian $H$, the goal of the problem is to determine whether $H$ is $\varepsilon_1$-close to or $\varepsilon_2$-far from a target Hamiltonian $H_0$. While Hamiltonian learning methods have been extensively studied, they often require restrictive assumptions and suffer from inefficiencies when adapted for certification tasks. This work introduces a direct and efficient framework for Hamiltonian certification. Our approach achieves \textit{optimal} total evolution time $\Theta((\varepsilon_2-\varepsilon_1)^{-1})$ for certification under the normalized Frobenius norm, without prior structural assumptions. This approach also extends to certify Hamiltonians with respect to all Pauli norms and normalized Schatten $p$-norms for $1\leq p\leq2$ in the one-sided error setting ($\varepsilon_1=0$). Notably, the result in Pauli $1$-norm suggests a quadratic advantage of our approach over all possible Hamiltonian learning approaches. We also establish matching lower bounds to show the optimality of our approach across all the above norms. We complement our result by showing that the certification problem with respect to normalized Schatten $\infty$-norm is $\mathsf{coQMA}$-hard, and therefore unlikely to have efficient solutions. This hardness result provides strong evidence that our focus on the above metrics is not merely a technical choice but a requirement for efficient certification. To enhance practical applicability, we develop an ancilla-free certification method that maintains the inverse precision scaling while eliminating the need for auxiliary qubits, making our approach immediately accessible for near-term quantum devices with limited resources.

quant-ph

Bra-ket entanglement, an indicator bridging entanglement, magic, and coherence

Understanding the intricate interplay between distinct quantum resources is a fundamental prerequisite for rigorously characterizing the boundary between classical and quantum technologies. Among the vast landscape of quantum resources, entanglement, magic, and coherence have arguably attracted the most intense investigation. However, while universally recognized as the core drivers of quantum advantage, our understanding of their structural interplay remains fragmented and compartmentalized. In this work, we introduce an indicator called {\em bra-ket entanglement} (BKE) defined in the operator vectorization space to bridge all three quantum resources. Specifically, we show that BKE governs a resource dependence transition in the generation of entanglement: in the low-BKE regime, the growth of entanglement is dominated by coherence, largely independent of magic. However, as BKE increases, the dependence on coherence will gradually be replaced by a dependence on magic. Consequently, in the high-BKE regime, entanglement generation becomes dominated by magic, largely independent of coherence. These results are built on a series of new entropy-theoretic relations and are verified through numerical experiments. We also discuss implications of our results for the resource transitions in classical simulations of mixed states and marginal probabilities and for relating different classical simulation methods.

quant-ph

Sequence Reconstruction under Channels with Multiple Bursts of Insertions or Deletions

The sequence reconstruction problem involves a model where a sequence is transmitted over several identical channels. This model investigates the minimum number of channels required for the unique reconstruction of the transmitted sequence. Levenshtein established that this number exceeds the maximum size of the intersection between the error balls of any two distinct transmitted sequences by one. In this paper, we consider channels subject to multiple bursts of insertions and multiple bursts of deletions, respectively, where each burst has an exact length of value b. We provide a complete solution for the insertion case while partially addressing the deletion case.

cs.IT

DConAD: A Differencing-based Contrastive Representation Learning Framework for Time Series Anomaly Detection

Time series anomaly detection holds notable importance for risk identification and fault detection across diverse application domains. Unsupervised learning methods have become popular because they have no requirement for labels. However, due to the challenges posed by the multiplicity of abnormal patterns, the sparsity of anomalies, and the growth of data scale and complexity, these methods often fail to capture robust and representative dependencies within the time series for identifying anomalies. To enhance the ability of models to capture normal patterns of time series and avoid the retrogression of modeling ability triggered by the dependencies on high-quality prior knowledge, we propose a differencing-based contrastive representation learning framework for time series anomaly detection (DConAD). Specifically, DConAD generates differential data to provide additional information about time series and utilizes transformer-based architecture to capture spatiotemporal dependencies, which enhances the robustness of unbiased representation learning ability. Furthermore, DConAD implements a novel KL divergence-based contrastive learning paradigm that only uses positive samples to avoid deviation from reconstruction and deploys the stop-gradient strategy to compel convergence. Extensive experiments on five public datasets show the superiority and effectiveness of DConAD compared with nine baselines. The code is available at https://github.com/shaieesss/DConAD.

cs.LG