Search arXivSearch

arXiv subjects

Yang Xiao

Publications and source records attributed to Yang Xiao.

At least 19 recordsLinked to original sources

Selective Quantum Mpemba Dynamics in Chaotic Spin Chains

In a symmetry-restoration quantum Mpemba effect, an initial state with stronger local symmetry breaking can lose that memory faster than a state that starts closer to the symmetric manifold. We study how this local ordering reversal depends on the spatial profile of a longitudinal field in a clean U(1)-conserving spin chain, comparing spectral level statistics with crossings of the entanglement asymmetry for the same Hamiltonians. We find that Gaussian orthogonal ensemble (GOE)-like level statistics alone do not determine whether Mpemba crossings occur. Across field textures, GOE-like spectra can occur with or without entanglement-asymmetry crossings at the same field strength. In a near-staggered detuned field, the total charge-sector coherence reverses its ordering while the entanglement asymmetry retains its initial order. This contrast links the crossing response to the distribution of local charge-sector coherence and its entropic weighting in the reduced density matrix.

quant-ph

Overcoming the Quasi-Static Bottleneck: A Finite-Time Quantum Otto Information Engine Achieving Near-Unity Efficiency

Generally, quantum heat engines driven by Gibbs reservoirs achieve their maximum work and efficiency only in quasi-static processes, a constraint that hampers practical applications due to the resulting vanishing power output. To address this fundamental limitation, we propose a finite-time quantum Otto information engine (OIE) driven by a Maxwell's demon paired with a single Gibbs reservoir. We demonstrate that the demon's measurement and feedback control can preserve quantum coherence to extract coherent work, while harnessing quantum internal friction as a work source. Consequently, the OIE can produce work by only modulating the eigenstates of the Hamiltonian, and surpass the quasi-static limits of its conventional counterpart in the work output and the efficiency, which accounts for the energetic cost of the demon. Notably, the OIE can achieve near-perfect efficiency alongside positive work extraction in the rapid-driving regime.

quant-ph

The Filtering Demon: Beyond Standard Thermodynamic Bounds and Harnessing Measurement Error and Quantum Friction

Standard stochastic heat engines operate blindly, enforcing work extraction protocols indiscriminately on microstates, and consequently suppressing work output, stability, efficiency. To overcome this, we propose an Otto information engine (OIE) that employs a Maxwell's demon to filter out detrimental stochastic trajectories, and demonstrate that the OIE can provide enhanced work output and stability compared to the corresponding standard Otto engine. Remarkably, even after accounting for the energetic costs of the demon, the OIE efficiency surpasses both the standard Otto limit and Carnot bound. Furthermore, by applying the fluctuation theorem of information dissipation, we derive upper and lower bounds on efficiency, confirming that our results adhere to the second law of thermodynamics. Finally, and counterintuitively, measurement errors and quantum inner friction, traditionally considered deleterious, can be harnessed as resources, which improves robustness and enables us to ignore the adiabatic strokes time.

quant-ph

Unbounded Work Extraction and Zero Work Fluctuations in a Super-Carnot Otto Information Engine

Generally, the work output of stochastic heat engines is governed by the stochastic trajectory distribution and the energy spectrum. Because the trajectory distribution depends on the thermal reservoir temperatures, the work output is not only constrained by temperature but is also susceptible to thermal fluctuations. To address these limitations, we introduce two continuous Maxwell's demons into a quantum Otto cycle, forming an Otto information engine (OIE). We demonstrate that the work output of the OIE depends solely on the energy level gap, enabling arbitrary work extraction while eliminating work fluctuations, thereby ensuring cycle-to-cycle identical work output. Furthermore, we show that the engine's efficiency can surpass the standard Carnot efficiency even after accounting for the energy cost of demon's memory erasure. Finally, we show that the OIE can deliver superior output power even when the demon's measurement time is taken into account, and the corresponding Monte Carlo simulation has been executed.

quant-ph

EasyScan_HEP 2: LLM-Agent Parameter-Scan Workflows in High Energy Physics

Large-language-model (LLM) agents are beginning to reshape the preparation and steering of computational workflows in high-energy physics phenomenology. To accommodate this change, we upgrade EasyScan_HEP to make the construction of parameter-scan configuration files more accessible to LLM-agent assistance. EasyScan_HEP 2 exposes command-line and machine-readable interfaces for LLM-agent workflows, allowing an assistant to translate natural language requests into an explicit .ini configuration that defines the scan method, external-program workflow, constraints, and outputs. The resulting configuration can be inspected through a local Web interface. The framework also supports LLM-agent-guided extension to new scan methods, as illustrated by the integration of BESTFIT, EMCEE, and DYNESTY. In this way, EasyScan_HEP 2 adapts parameter-scan workflows to LLM-agent use while preserving reproducibility, transparency, and user control.

hep-ph

M2Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models

Recent advancements have successfully adapted autoregressive language models to process multimodal signals, such as images and actions. Since raw action signals are continuous, effective tokenization is essential to map high-dimensional inputs into compact discrete tokens for autoregressive processing. However, existing discrete action tokenizers often suffer from high reconstruction loss, failing to preserve the fine-grained dynamics required for precise control. This "discretization bottleneck" significantly limits the performance ceiling of downstream Vision-Language-Action (VLA) models. To address this, we propose ${M}^2$Tok, a Multi-head Multi-codebook Action Tokenizer designed to minimize reconstruction error and enhance policy performance. Our approach introduces two key structural innovations: (1) we decompose the latent action features into multiple heads, enabling the model to implicitly align specific heads with distinct action dimensions; (2) we assign independent codebooks to each head for quantization. By leveraging the combinatorial nature of multiple codebooks, we significantly expand the representational expressivity of the tokenizer, leading to substantially lower reconstruction loss compared to previous methods. We evaluate the ${M}^2$Tok-based VLA on the RoboTwin, Simpler-Env, and 3 zero-shot real-world tasks. Experimental results demonstrate our method not only achieves superior reconstruction fidelity but also significantly boosts the success rate of VLA models. Comprehensive ablation studies further confirm the effectiveness of the multi-head and multi-codebook mechanisms. Code is available at https://github.com/cpaaax/M2Tok.

cs.RO

Beyond the Name: Demographic Leakage in De-Identified Résumés and Evaluation Artifacts in LLM Bias Audits

De-identified résumé screening assumes that redacting explicit fields prevents ethnocultural inference; however, recent audits attribute residual leakage to declared languages. We investigate whether eliminating language fields resolves this leakage across nine open-weight models and 620 counterfactual résumés. By holding language attributes strictly identical, we isolate unstructured prose across five ethnocultural conditions and three cue-salience tiers. Target-group recovery averages 0.757 overall and saturates at 1.000 under high salience, demonstrating that non-language prose sustains demographic inference. Crucially, models diverge only under faint cues (0.086-0.690), establishing salience as an essential evaluation axis. Furthermore, pairwise LLM-as-a-judge outcomes are highly sensitive to evaluation design: forbidding ties yields an apparent selection-rate ratio of 0.39 alongside strong position and content effects, whereas permitting ties produces near-universal ties for most models ($\ge94\%$). Downstream scoring shows only very small between-condition differences, highlighting the need to distinguish demographic signals recoverable from résumé content from effects introduced by the evaluation protocol.

cs.CL

Competence-Preserving Resume Perturbations Expose Presentation Sensitivity in LLM Screening

Resume screeners must infer job-relevant competence from resumes whose presentation can vary substantially in wording, structure, stylistic polish, and document extraction quality. Ideally, such surface variation should not change decisions when the underlying qualification evidence is unchanged. We introduce a controlled audit of this property, constructing occupation-grounded candidate profiles at controlled competence levels and rendering each profile into multiple resume presentations. A deterministic validation gate excludes variants that alter the underlying evidence before scoring. Across six open instruction-tuned LLM conditions, we find a clear disconnect between screening validity and presentation stability. Llama-3.1-8B with its native chat template achieves the strongest validity ($0.781$) yet reverses $29.6\%$ of matched pairwise decisions under competence-preserving presentation changes; Mistral-7B-v0.3 reaches validity $0.644$ with a $41.4\%$ flip rate. Native chat formatting improves validity for several chat-tuned models but does not remove this instability. These results show that resume-screening evaluations should assess not only whether a system identifies stronger candidates, but also whether those decisions remain stable when the same competence evidence is presented differently.

cs.CL

RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems

Existing emotional support conversation systems mainly focus on one-on-one seeker-supporter interactions and individual emotional states, leaving interpersonal relations in multi-party scenarios underexplored. In this work, we introduce relation-aware emotional support conversation, a new task that evaluates whether LLMs can capture and utilize the evolving dynamics of relationships to offer more effective emotional support. We construct RESCUE (Relation-aware Emotional Support Conversation Understanding and Evaluation Benchmark) from real couple and family interview conversations, containing 191 samples, 7,079 annotated turns, and 1,064.8 minutes of video. Based on rich annotations of socio-emotional and support-related dynamics, RESCUE defines six tasks that evaluate two core capabilities required for relation-aware emotional support: Relational Understanding and Relation-Sensitive Support. Experiments with ten LLMs show that current models perform relatively well on tasks relying on local emotional or intervention cues, but struggle with relation-intensive tasks such as relation pattern prediction, viewpoint prediction, and support strategy prediction. These findings reveal the limitations of current LLMs in modeling interpersonal relations and making relation-sensitive support decisions.

cs.AI

The Inert Doublet Model of Dark Matter and the LUX-ZEPLIN High-Recoil Event

The recent LUX-ZEPLIN (LZ) search reported a nuclear-recoil event near $248~{\rm keV}$. Such a high-recoil event can be interpreted in terms of endothermic inelastic dark matter scattering, which typically requires a mass splitting of order a few hundred keV between the initial and final dark states. The inert doublet model (IDM) provides a natural realization of this scenario through the $Z$-mediated transition $H+N\to A+N$. In this work, we systematically examine whether this interpretation can be consistently realized in the IDM under the relevant theoretical constraints and existing experimental bounds. We perform a detailed profile-likelihood analysis of the LZ event, and find a viable high-mass IDM region with a profile best fit at $m_H=1080~{\rm GeV}$ and $m_A-m_H=369~{\rm keV}$.

hep-ph

CWF: A Collaborative Writing Framework for Personalized and Reliable Popular Science Writing

We introduce Personalized and Reliable Popular Science Writing, a novel task that requires adapting scientific explanations to audiences with different cognitive levels while preserving factual accuracy. However, improving personalization often introduces simplifications that increase the risk of hallucination and factual distortion. To address these challenges, we first construct a dataset of 39,134 entries and a reader-centric Personalized Science Communication Benchmark (PSCB) that jointly evaluates audience adaptation and factual accuracy. To reduce data and computational requirements while improving generalization across domains and audiences, we introduce DA-MoE, which explicitly decouples audience adaptation from domain knowledge through separate modeling. To enable robust verification and revision in evidence-scarce scenarios, a multi-agent fact-checking mechanism that augments limited evidence with role-specific agent debate and propagates confidence over a graph is proposed. Experiments on PSCB show that our approach achieves state-of-the-art performance. Our code is open-sourced at https://github.com/DPInnovationWorks/CWF.

cs.AI

PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents

Long-horizon agent runs generate experience that can improve both the current run and future work. Most self-improvement methods process this experience only after execution ends, so they cannot redirect the active run or immediately apply and validate lessons learned from it. We argue that self-improvement should instead be live, using emerging experience both to redirect the active run and to update the persistent harness. Existing agent architectures do not fully support this goal. Single-agent self-correction combines task execution and trajectory assessment within one context, while subagent delegation separates execution but typically cannot redirect an active subagent. We present PILOT, a supervisor-worker harness for live self-improvement through two coupled mechanisms: (1) live steering lets a separate supervisor redirect or abort the active worker during execution; and (2) live self-evolution distils procedures and failure modes revealed during execution into reusable skills and memory. Across two frozen backbones and three benchmarks, PILOT ranks first in five of six configurations. On Terminal-Bench 2.0, PILOT outperforms counterpart harnesses by up to 9.8 percentage points. In the self-improvement setting, PILOT gains 14.6 points with GLM-5.1 and 12.4 points with Kimi-K2.6. Mean output tokens fall by 42.9% and 47.4%, while successful evaluations per million output tokens rise by 110.3% and 134.0%, respectively.

cs.AI

Model-Consistent Byzantine-Resilient Decentralized Federated Learning for Collaborative Missions

Decentralized federated learning (DFL) is a promising paradigm for autonomous nodes to collaboratively train AI models without relying on a central server. However, existing DFL solutions do not guarantee global model consistency, a critical requirement for collaborative mission-critical scenarios where model divergence undermines decision uniformity and safety. This lack of consistency also amplifies vulnerability to Byzantine adversaries, who exploit the decentralized network topology and weak synchrony to perform equivocation and model poisoning attacks against individual victims. This paper introduces DFL-C, a novel Byzantine-resilient DFL architecture that enables decentralized nodes to perform collaborative training with global model consistency. At its core, DFL-C integrates an asynchronous common subset (ACS) consensus protocol into the DFL workflow to ensure all nodes aggregate a uniform set of model updates to establish global model consistency, despite individual Byzantine equivocation. DFL-C further implements a dual-domain trust scoring mechanism to provide resilience against data-domain Byzantine manipulations including model poisoning attacks. This mechanism complements the consensus protocol, significantly reducing the latter's runtime. Our experimental results demonstrate that DFL-C maintains model accuracy while achieving global model consistency under Byzantine behaviors with moderate consensus overhead. Notably, when compared with the state-of-the-art DFL solution BALANCE (Fang et al.) that does not provide model consistency, DFL-C achieves better model accuracy against untargeted model poisoning attacks and comparable resilience against backdoor attacks, with the advantage widened under non-IID scenarios.

cs.DC

Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation

Recent multimodal large language models can convert visual designs directly into executable code, but real mobile products require multiple screenshots to become a buildable codebase with shared components and working navigation. This project-level setting exposes three limits of existing design-to-code benchmarks: they focus on single-page generation rather than complete codebases, cannot evaluate cross-page navigation, and do not measure project-wide maintainability. We introduce MobileForge, the first benchmark for project-level multi-screen mobile app generation, comprising real mobile apps, human-reviewed screens, structured page-relationship annotations, and navigation test specifications. MobileForge supports five-axis evaluation of build, navigation, visual fidelity, code maintainability, and efficiency. We also propose state-isolated navigation testing to avoid cascading failures in navigation evaluation and an anchor-referenced list-wise visual evaluation protocol to improve visual-judge reliability. Across end-to-end runs on six frontier multimodal LLMs, current models can build mobile-app projects that compile and reach the correct pages, but interactive navigation remains unreliable and visual fidelity and maintainability still lag. The benchmark and supporting materials are available at https://github.com/anoa12159-hue/mobileforge_eval.

cs.HC

Renormalization Group Analysis of Pairing Instabilities in Nuclear Fermi Liquids

A nuclear Fermi liquid exhibits competing pairing instabilities in different spin, isospin, and orbital channels. In a Fermi-surface renormalization group (RG) treatment, the channel that develops a pole first is determined not only by its tree-level attraction but also by its one-loop RG coefficient. We illustrate this mechanism in a minimal $S/P$-wave model. A spherical Fermi surface establishes the reference competition between the lowest even- and odd-parity interactions. Axial deformation changes the relevant Fermi-surface integrals and lifts the degeneracy between longitudinal and transverse $P$-wave components. In isospin-asymmetric matter, neutron--proton Fermi-momentum splitting restricts the simultaneous low-energy contribution of the two species and can terminate the $np$ running at finite threshold scales. Our calculations are intended as controlled one-loop RG illustrations rather than as quantitative nuclear-matter calculations. We show how the Fermi-surface geometry and composition can change the ordering of competing pairing instabilities.

nucl-th

Partial FC: Training 10 Million Identities on a Single Machine

Training face recognition models with millions of identities is challenging because classifier storage, logit memory, and computation grow linearly with the number of classes, eventually making full softmax impractical even when the backbone itself fits comfortably in memory. We present Partial FC (PFC), a scalable approximation to large-class softmax that preserves every positive class center while activating only a sampled subset of negative centers in each mini-batch. This asymmetric treatment retains every target term while avoiding exhaustive interaction with millions of mostly uninformative negatives. Our distributed implementation partitions the classifier across GPUs and samples within each owned shard, so sampling reduces local matrix multiplication and logit storage while sharding eliminates class-gradient synchronization across workers. Together, these properties reduce GPU-resident classifier memory, computation, and class-dependent communication without feature-based hard-negative retrieval. End-to-end system benchmarks demonstrate efficient scaling to massive class spaces, including 64 million classes on a single eight-GPU machine. Across large-scale face-recognition datasets, moderate sampling rates maintain competitive recognition accuracy while substantially improving training efficiency. Our best PFC configurations achieve 97.2\% TAR on IJB-C at FAR $=10^{-4}$ and 94.0\% TAR on ICCV21-MFR at FAR $=10^{-6}$. Beyond clean training data, PFC is robust to inter-class conflicts, label noise, and long-tailed identity distributions: under 40\% label noise, PFC with conflict filtering raises ICCV21-MFR TAR from 43.9\% to 80.2\%, while on long-tailed data PFC improves TAR from 87.4\% to 92.0\%. These results establish positive-preserving negative sampling as an effective foundation for scalable, accurate, and robust identity classification.

cs.CV

Wrong Design Intent Is Worse Than Never Conditioning: A Derangement-Control Diagnosis of Header Conditioning in CAD Program Completion

Fine-tuned code LLMs are routinely conditioned on a design-intent specification, but the correctness axis of such a signal -- a wrong intent rather than an absent one -- has not been tested, and the benefit of conditioning is usually scored with the same detector that defines the signal. We study CADCON, a five-feature design-intent header prepended to CadQuery-style programs during LoRA fine-tuning of Qwen2.5-Coder-1.5B, scoring adherence with executable geometric assertions that share no code with the header-defining extractor. On a pre-registered sample of 400 deduplicated held-out programs stratified over eleven intent profiles, at 40% prefix and three seeds, a semantically wrong header degrades adherence below the never-header-trained baseline on 3/3 seeds under both tokenizations at the program level, and on 3/3 token and 2/3 text seeds at the 298 distinct model inputs they present. Wrong-header executability is not depressed relative to that baseline. A derangement control, retrained so every program receives another program's header -- holding the header marginal fixed while destroying its correlation with the program -- saw the same programs, indices and wrong headers. Its correct-to-wrong change is -0.006/+0.016/-0.003 against 0.124/0.241/0.230 for the standard model, and the interaction is significant on 3/3 seeds (p <= 5.9e-7), so the model's sensitivity to whether the header is right or wrong requires the learned mapping. The control sits below the baseline by the same margin under a correct as under a wrong header, so we claim that sensitivity and not the below-baseline level. On features the true intent lacks, the standard model realizes a feature far more often when the wrong header names it; the control does not. Ground truth itself scores only 0.567 here, the scale on which arm levels should be read. Wrong design intent is not inert: it actively misdirects generation.

cs.LG

MARCO: Click-Intent Decomposition for Calibrated Ads Conversion Prediction

Not all clicks are equal. Industrial ads ranking decouples conversion probability into click-through rate (CTR) and post-click conversion rate (CVR), yet treats every click as the same event. In reality, users provide a free, self-generated signal of intent through their physical UI interactions. Different click types on the same ad exhibit a 4-fold difference in actual conversion rates. By conflating these signals, the standard CVR model under-predicts high-intent clicks and over-predicts low-intent ones, which is a bias masked by near-perfect aggregate calibration. We propose MARCO (Multi-intent Ads Ranking Composition Optimization), a framework that resolves this bias by decomposing each click by intent. Using the logged click type as a free behavioral label, MARCO trains per-intent CVR heads on homogeneous populations, and at serving time composes their per-intent CVR estimates under a predicted distribution over intents. Theoretically, we prove that decomposition never raises population risk, give the exact headroom under squared loss and non-negativity under the deployed loss, and show through a routing-efficiency dial how much of it reaches serving. Because the population-optimal score is unchanged, any gain is a finite-capacity estimation and calibration effect that we validated both offline and online. For deployment at scale, we further cast multi-impression, multi-click attribution as credit assignment with a bias-variance tradeoff analogous to RL return estimation, showing last-impression, first-click attribution is the low-bias, low-variance, deterministic choice under production constraints, and derive three consistency conditions enforced end-to-end at scale. Deployed at binary intent granularity, MARCO corrects per-intent calibration to approximately 100%, lifts conversions per click by +2.80%, and drives +0.98% cumulative improvement in topline metrics.

cs.LG