Search arXiv⌕ Search

arXiv subjects

Zhaoyi Li

Publications and source records attributed to Zhaoyi Li.

At least 19 recordsLinked to original sources

Claim-Gated Source-Risk Auditing for Generative Search

A generative search answer can cite a supported passage yet omit a source relationship that changes its interpretation. We specify a claim-gated audit of the query-source-answer tuple. An omission is resolved only when relationship evidence, answer adoption, materiality, and disclosure are all observed; incomplete evidence remains unresolved rather than being treated as independence. The specification separates this endpoint from citation support and review priority, and binds decisions to versioned evidence spans. A reference checker makes the record contract executable. On an exhaustive synthetic suite, it reproduces all 81 three-state predicate combinations and rejects 192 deliberately malformed records. Common-guard baselines and predicate ablations isolate endpoint logic from missing-evidence handling, while controlled transitions check support separation and evidence removal. These are finite contract-conformance results, not detector accuracy or evidence of improved user outcomes. We define the independent annotation, held-out evaluation, and paired utility tests still required to establish semantic validity and deployment benefit.

cs.AI↗

IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives

An external record may contain a procedure to apply or text to read, depending on the user's request. IBBench-Light tests both uses against the same record. Twelve semantic bases yield 144 matched pairs per model; four quantized instruction models produced 1,152 archived greedy responses. Paired exact-contract accuracy (PECA) requires both members to satisfy their output contracts. Qwen succeeds on 132 execute and 109 process prompts, but only 97 complete pairs, showing what marginal averages omit. We audit literal-target exposure and case normalization, then add 1,722 logged CPU generations to test directive-absent controls, twelve additional semantic bases, within-base wording changes, and generation stopping. In the pinned Phi rerun, changing the end-of-sequence (EOS) set changes exact paired success from 0/144 to 62/144. A bounded IHEval comparison uses the same SmolLM2 checkpoint and output budget while preserving its published instruction roles and scorer. The benchmark measures conditional task and output-contract success. Its task margins and paired count need to be read together with the stopping policy.

cs.AI↗

A Kagome-Derived Mosaic Lattice Family A3V9Te13 (A = Cs, Rb) with Tunable Strong Electronic Correlations

The pursuit of geometrically frustrated lattices beyond conventional paradigms remains a central challenge in the design of quantum materials. Herein, we report the discovery of the A3V9Te13 (A = Cs, Rb) family of vanadium-based intermetallic compounds, which host a unique two-dimensional Mosaic lattice derived from the Kagome network, composed of an ordered tessellation of triangles, squares, and pentagons. The Cs compound (CVT) exhibits strong electronic correlations, characterized by non-Fermi liquid behavior at low temperatures, an exceptionally large Sommerfeld coefficient, and a bulk phase transition at T* $\approx$ 47 K with possible charge- or spin-related origin. Inspired by pressure-tuning in related Kagome systems, we demonstrate that the electronic ground state of this lattice is exquisitely tunable via chemical pressure. Systematic substitution of Cs with smaller Rb ions suppresses the T* phase transition and the correlated electronic response, ultimately driving the system into a highly frustrated semiconducting ground state without long-range magnetic order down to 60 mK. This work unveils a new structural platform for exploring the interplay between geometric frustration and strong electron correlations, providing a chemically controllable platform for exploring the phase space between distinct correlated electronic states.

cond-mat.mtrl-sci↗

Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models

On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data. We present a controlled study that varies one generalization factor at a time, from in-domain distribution shifts to cross-domain transfer and the multi-teacher setting. We find that OPD transfers a teacher's reasoning behavior rather than its answers to particular problems: training difficulty barely matters, and even problems the teacher never solves are useful. Transfer depends strongly on the origin relationship between teacher and student: same-origin pairs bring the student close to the teacher across languages, reasoning horizons, and even other domains, whereas cross-origin pairs mostly fit the trained distribution. This broad reach is a double-edged sword: since routing prompts to domain experts cannot confine each teacher's influence, combining them yields a mixture-dependent seesaw among their capabilities. These results clarify when OPD generalizes and offer a useful perspective for diagnosing multi-teacher OPD.

cs.CL↗

Dolomite Mineral-Inspired Equilateral Triangular-Lattice Magnets for Quantum Magnetism

Equilateral triangular lattice magnets provide a versatile materials platform for exploring exotic quantum spin phenomena, while their field-tunable magnetic entropy offers opportunities for low-temperature adiabatic demagnetization refrigeration. Inspired by the natural mineral, we proposed a chemical strategy to achieve equilateral TL magnets, leveraging the high crystal symmetry of a large family of dolomite-type materials. As typical examples, the dolomite-type materials SnM(BO3)2 (M = Co, Mn) were synthesized, and structural analysis reveals that Co2+ and Mn2+ ions form equilateral triangular lattices with an A-B-C stacking fashion. The magnetic susceptibilities and specific heat measurements reveal dominant antiferromagnetic interactions, with Neel temperatures of 0.49K for SnCo(BO3)2 and 0.96K for SnMn(BO3)2, respectively. Our results establish the dolomite-type M'M(X)2 (M'and M sites allow various valence states, e.g., +4/+2 or +3/+3; X = CO32- or BO33-) system as a chemically flexible and structurally perfect material platform for exploring frustrated magnetism and low-temperature magnetocaloric applications.

cond-mat.mtrl-sci↗

Scaling-optimal purification of noisy qubit unitary channels

We consider the problem of purifying noisy qubit unitary channels. Given the ability to apply an unknown qubit unitary channel followed by depolarizing noise, we aim to construct a superchannel that purifies the noisy unitary back to the original unknown unitary. We first provide numerical evidence that sequential strategies can strictly outperform parallel strategies when the number of channel uses is finite, highlighting the fundamental distinction from state purification. We then provide a concrete $\mathrm{U}(2)$-covariant parallel protocol based on a novel entanglement-assisted quantum error-correcting code that suppresses the first-order noise strength as $O(1/n)$ with $n$ channel uses and show this scaling is asymptotically optimal in the low-noise regime, even when sequential strategies are allowed.

quant-ph↗

Field-Induced Up-Up-Down State and Frustrated Magnetism in a Non-Kramers Triangular Antiferromagnet

A previously unreported triangular lattice (TL) antiferromagnet, TmZnGaO4, was synthesized as single crystals, and its crystal structure, magnetic susceptibilities, and specific heat were reported. Its crystal structure is isomorphic to that of the transverse-field Ising antiferromagnet TmMgGaO4, with Tm3+ ions located in the TLs, separated by a nonmagnetic bilayer composed mainly of Ga3+ and Zn2+ ions. The magnetic susceptibilities indicate the dominating antiferromagnetic interactions. The magnetization curves (M-H) exhibit strong easy-c-axis anisotropy, with a clear one-third magnetic plateau emerging, consistent with a field-induced up-up-down spin configuration. Instead of forming a conventional long-range magnetic order, the system exhibits two broad anomalies at 0.11 K and 2.81 K in zero-field specific heat measurements, highlighting the persistence of strong spin fluctuations and the potential for exotic quantum spin states. The above results reveal its future interest in exploring exotic quantum spin states in TmZnGaO4.

cond-mat.str-el↗

LLMSurgeon: Diagnosing Data Mixture of Large Language Models

The pretraining data mixture of Large Language Models (LLMs) constitutes their "digital DNA", shaping model behaviors, capabilities, and failure modes. Yet this composition is rarely disclosed, making post-hoc auditing of data combination or provenance difficult. In this work, we formalize $\textbf{Data Mixture Surgery (DMS)}$: given only generated text from a target LLM, estimate the domain-level distribution of its pretraining corpus under a predefined taxonomy. We propose $\textbf{LLMSurgeon}$, a strong framework that casts DMS as an inverse problem under the label-shift assumption. Rather than directly aggregating classifier outputs, LLMSurgeon estimates a calibrated $\textit{soft}$ confusion matrix and solves a constrained inverse problem to correct systematic domain confusion and recover the latent mixture prior. To evaluate, we introduce $\textbf{LLMScan}$, a recipe-verifiable evaluation suite built from open-source LLMs with transparent pretraining mixtures. Across LLMScan, LLMSurgeon recovers domain mixtures with high fidelity under fixed protocols. Our work presents a practical, post-hoc approach for auditing the digital DNA of foundation models without access to their training data.

cs.CL↗

An Exponential Sample-Complexity Advantage for Coherent Quantum Inference

Standard quantum inference converts quantum data into classical outputs. We study an alternative inference setting in which the desired output is quantum, preserving coherence. Such settings include quantum purity amplification (QPA), mixed-state approximate purification or cloning, and density matrix exponentiation. We show that such protocols can achieve exponentially lower sample complexity than incoherent, measurement-mediated protocols. For QPA with principal eigenstate targets and $d$-dimensional inputs, coherent processing achieves error $\varepsilon$ using $O(1/\varepsilon)$ copies, versus the $Ω(d/\varepsilon)$ copies required by any incoherent protocol. Together, these sharp coherent-incoherent separations seed a theory of coherent quantum inference, with an entanglement-breaking limit identifying the optimal incoherent counterpart of each coherent protocol.

quant-ph↗

Quantum Purity Amplification for Arbitrary Eigenstates and Multiple Outputs

Quantum purity amplification (QPA) is the task of coherently transforming $n$ copies of a mixed state into high-fidelity copies of a chosen eigenstate. We solve QPA in the general setting of $n$ input copies, $m$ output copies, arbitrary target eigenstates, arbitrary local dimension $d$, and generic input spectra. We characterize the optimal channel and derive its all-site and one-site performance laws across output regimes. For the asymptotic analysis, we use a path-graph parametrization to show that, when the target eigenvalue has a constant spectral gap $D_{k,\mathrm{min}}$, achieving all-site error $\varepsilon$ requires a number of input copies independent of $d$ and scaling as $O(m/(\varepsilon D_{k,\mathrm{min}}^2))$. When $m/n$ approaches a constant, the performance exhibits phase-like regimes, which we characterize explicitly. For the nonasymptotic analysis, we develop a theory of generalized Young diagrams that yields tight sample complexity bounds and provides the first dimension-uniform guarantee for optimal QPA. We also provide asymptotically efficient implementations of the optimal protocol. Together, these results establish QPA as a rigorous example of coherent quantum information processing with dimension-uniform sample complexity, supplying the technical foundation for the coherent-incoherent separation developed in the companion work.

quant-ph↗

Scaling Reasoning Hop Exposes Weaknesses: Demystifying and Improving Hop Generalization in Large Language Models

Chain-of-thought (CoT) reasoning has become the standard paradigm for enabling Large Language Models (LLMs) to solve complex problems. However, recent studies reveal a sharp performance drop in reasoning hop generalization scenarios, where the required number of reasoning steps exceeds training distributions while the underlying algorithm remains unchanged. The internal mechanisms driving this failure remain poorly understood. In this work, we conduct a systematic study on tasks from multiple domains, and find that errors concentrate at token positions of a few critical error types, rather than being uniformly distributed. Closer inspection reveals that these token-level erroneous predictions stem from internal competition mechanisms: certain attention heads, termed erroneous processing heads (ep heads), tip the balance by amplifying incorrect reasoning trajectories while suppressing correct ones. Notably, removing individual ep heads during inference can often restore the correct predictions. Motivated by these insights, we propose test-time correction of reasoning, a lightweight intervention method that dynamically identifies and deactivates ep heads in the reasoning process. Extensive experiments across different tasks and LLMs show that it consistently improves reasoning hop generalization, highlighting both its effectiveness and potential.

cs.CL↗

BiT-MCTS: A Theme-based Bidirectional MCTS Approach to Chinese Fiction Generation

Generating long-form linear fiction from open-ended themes remains a major challenge for large language models, which frequently fail to guarantee global structure and narrative diversity when using premise-based or linear outlining approaches. We present BiT-MCTS, a theme-driven framework that operationalizes a "climax-first, bidirectional expansion" strategy motivated by Freytag's Pyramid. Given a theme, our method extracts a core dramatic conflict and generates an explicit climax, then employs a bidirectional Monte Carlo Tree Search (MCTS) to expand the plot backward (rising action, exposition) and forward (falling action, resolution) to produce a structured outline. A final generation stage realizes a complete narrative from the refined outline. We construct a Chinese theme corpus for evaluation and conduct extensive experiments across three contemporary LLM backbones. Results show that BiT-MCTS improves narrative coherence, plot structure, and thematic depth relative to strong baselines, while enabling substantially longer, more coherent stories according to automatic metrics and human judgments.

cs.CL↗

On the Role of Reasoning Patterns in the Generalization Discrepancy of Long Chain-of-Thought Supervised Fine-Tuning

Supervised Fine-Tuning (SFT) on long Chain-of-Thought (CoT) trajectories has become a pivotal phase in building large reasoning models. However, how CoT trajectories from different sources influence the generalization performance of models remains an open question. In this paper, we conduct a comparative study using two sources of verified CoT trajectories generated by two competing models, \texttt{DeepSeek-R1-0528} and \texttt{gpt-oss-120b}, with their problem sets controlled to be identical. Despite their comparable performance, we uncover a striking paradox: lower training loss does not translate to better generalization. SFT on \texttt{DeepSeek-R1-0528} data achieves remarkably lower training loss, yet exhibits significantly worse generalization performance on reasoning benchmarks compared to those trained on \texttt{gpt-oss-120b}. To understand this paradox, we perform a multi-faceted analysis probing token-level SFT loss and step-level reasoning behaviors. Our analysis reveals a difference in reasoning patterns. \texttt{gpt-oss-120b} exhibits highly convergent and deductive trajectories, whereas \texttt{DeepSeek-R1-0528} favors a divergent and branch-heavy exploration pattern. Consequently, models trained with \texttt{DeepSeek-R1} data inherit inefficient exploration behaviors, often getting trapped in redundant exploratory branches that hinder them from reaching correct solutions. Building upon this insight, we propose a simple yet effective remedy of filtering out frequently branching trajectories to improve the generalization of SFT. Experiments show that training on selected \texttt{DeepSeek-R1-0528} subsets surprisingly improves reasoning performance by up to 5.1% on AIME25, 5.5% on BeyondAIME, and on average 3.6% on five benchmarks.

cs.CL↗

GTM: A General Time-series Model for Enhanced Representation Learning of Time-Series Data

Despite recent progress in time-series foundation models, challenges persist in improving representation learning and adapting to diverse downstream tasks. We introduce a General Time-series Model (GTM), which advances representation learning via a novel frequency-domain attention mechanism that captures time-granularity-aware features, an aspect underexplored in prior research. We further propose a novel pre-training strategy that unifies reconstruction and autoregressive objectives through a hybrid masking mechanism. Our pre-training strategy, combined with 2D positional encoding and span shuffling, enhances the robustness and generalization of representations. GTM is established as the first generative-task-agnostic model for time-series analysis, enabling seamless adaptation to various generative tasks without any task-specific modifications. Extensive experiments demonstrate that GTM consistently outperforms SOTA models on various generative tasks and achieves strong classification results with minimal adaptation. Furthermore, GTM exhibits clear scaling behavior, with accuracy improving as model size and pre-training data increase.

cs.LG↗

HieraMAS: Optimizing Intra-Node LLM Mixtures and Inter-Node Topology for Multi-Agent Systems

Multi-agent systems (MAS) built on large language models (LLMs) have shown strong performance across many tasks. Most existing approaches improve only one aspect at a time, such as the communication topology, role assignment, or LLM routing, while treating each agent as a single, indivisible unit. This misses the opportunity to use mixtures of LLMs within an agent to strengthen role-specific abilities. We propose HieraMAS, a hierarchical collaboration framework that combines intra-node LLM mixtures with an inter-node communication topology. HieraMAS introduces supernodes, where each functional role is implemented by multiple heterogeneous LLMs using a propose-synthesis structure. Optimizing HieraMAS creates unique credit-assignment challenges: final task performance depends heavily on the underlying LLMs' capabilities, which can lead reinforcement methods to incorrectly reward suboptimal configurations. To address this, we use a two-stage algorithm: (1) multi-level reward attribution, which provides fine-grained feedback at both the node level and the overall system level; (2) graph classification for topology selection, which treats choosing the communication structure as a holistic decision rather than optimizing edges one by one. Experiments on reasoning and coding benchmarks show that HieraMAS substantially outperforms existing methods while also delivering better cost-performance trade-offs.

cs.MA↗

Pushing the Frontier of Black-Box LVLM Attacks via Fine-Grained Detail Targeting

Black-box adversarial attacks on Large Vision-Language Models (LVLMs) are challenging due to missing gradients and complex multimodal boundaries. While prior state-of-the-art transfer-based approaches like M-Attack perform well using local crop-level matching between source and target images, we find this induces high-variance, nearly orthogonal gradients across iterations, violating coherent local alignment and destabilizing optimization. We attribute this to (i) ViT translation sensitivity that yields spike-like gradients and (ii) structural asymmetry between source and target crops. We reformulate local matching as an asymmetric expectation over source transformations and target semantics, and build a gradient-denoising upgrade to M-Attack. On the source side, Multi-Crop Alignment (MCA) averages gradients from multiple independently sampled local views per iteration to reduce variance. On the target side, Auxiliary Target Alignment (ATA) replaces aggressive target augmentation with a small auxiliary set from a semantically correlated distribution, producing a smoother, lower-variance target manifold. We further reinterpret momentum as Patch Momentum, replaying historical crop gradients; combined with a refined patch-size ensemble (PE+), this strengthens transferable directions. Together these modules form M-Attack-V2, a simple, modular enhancement over M-Attack that substantially improves transfer-based black-box attacks on frontier LVLMs: boosting success rates on Claude-4.0 from 8% to 30%, Gemini-2.5-Pro from 83% to 97%, and GPT-5 from 98% to 100%, outperforming prior black-box LVLM attacks. Code and data are publicly available at: https://github.com/vila-lab/M-Attack-V2.

cs.LG↗

Dataset Distillation via Committee Voting

Dataset distillation aims to synthesize a compact yet representative dataset that preserves the essential characteristics of the original data for efficient model training. Existing methods mainly focus on improving data-synthetic alignment or scaling distillation to large datasets. In this work, we propose $\textbf{C}$ommittee $\textbf{V}$oting for $\textbf{D}$ataset $\textbf{D}$istillation ($\textbf{CV-DD}$), an orthogonal approach that leverages the collective knowledge of multiple models to produce higher-quality distilled data. We first establish a strong baseline that achieves state-of-the-art performance through modern architectural and optimization choices. By integrating distributions and predictions from multiple models and generating high-quality soft labels, our method captures a broader range of data characteristics, reduces model-specific bias and the impact of distribution shifts, and significantly improves generalization. This voting-based strategy enhances diversity and robustness, alleviates overfitting, and improves post-evaluation performance. Extensive experiments across multiple datasets and IPC settings demonstrate that CV-DD consistently outperforms single- and multi-model distillation methods and generalizes well to non-training-based frameworks and challenging synthetic-to-real transfer tasks. Code is available at: https://github.com/Jiacheng8/CV-DD.

cs.CV↗

Fine-Tuning Flow Matching via Maximum Likelihood Estimation of Reconstructions

Flow Matching (FM) models achieve remarkable results in generative tasks. Building upon diffusion models, FM's simulation-free training paradigm enables simplicity and efficiency but introduces a train-inference gap: model outputs cannot be assessed during training. Moreover, the straight flow assumption suffers from some inherent limitations. To address this, we propose to fine-tune FM via Maximum Likelihood Estimation (MLE) of reconstructions -- enabled by FM's smooth ODE formulation, unlike the stochastic differential equations (SDEs) in diffusion models. We first theoretically analyze the relationship between training loss and inference error in FM under numerical precision constraints. We then propose an easy-to-implement fine-tuning framework based on MLE of reconstructions, with flexibility for sophisticated extensions. Building on this, we incorporate a generalized artificial viscosity term that enhances flow stability and robustness, accompanied by a direct parameterization method and rigorous theoretical guarantees. Experiments demonstrate our method's effectiveness across diverse settings: a toy example provides mechanistic insights into the fine-tuning process, while large-scale evaluations on meteorological forecasting and robotic manipulation policies validate reliable performance improvements.

cs.LG↗