Search arXiv⌕ Search

arXiv subjects

Zhiyuan Li

Publications and source records attributed to Zhiyuan Li.

At least 19 recordsLinked to original sources

Provable Benefit of SignGD: A Minimal Model Under Heavy-Tailed Class Imbalance

Adaptive and non-Euclidean optimizers often outperform Euclidean methods such as stochastic gradient descent (SGD) in language modeling by a large margin. Existing theory usually explains this gap by assuming favorable smoothness geometry or noise structure tailored to the specific optimizer. We instead ask whether such geometry can be induced from a concrete learning setting. Starting from an optimizer gap that persists across realistic language-modeling experiments, we progressively remove sequence dependence, architectural complexity, and stochasticity. We find that the gap exists in a minimal setting: the softmax unigram model with heavy-tailed data. This model exposes a simple deterministic mechanism under heavy-tailed class imbalance. We prove that GD learns rare tokens slowly because the corresponding logits receive only tiny updates, while SignGD removes this magnitude dependence and moves rare and common coordinates on a more comparable scale. We make this precise with upper and lower bounds for the convergence rate of GD and upper bounds for the convergence of SignGD. Our stochastic bounds contain additional noise-dependent terms that can obscure this advantage in the convergence guarantees and can be reduced by increasing the batch size

cs.LG↗

Supersingular Tate conjecture for irreducible symplectic varieties of known types: I

We prove that for a supersingular irreducible symplectic variety admitting suitable lifting to characteristic zero has Tate Chow motive if it is of deformation type $K3^{[n]}$, OG6 (with Artin invariant $\neq$ 4), or OG10 (with Artin invariant $\neq$ 12), and has supersingular abelian Chow motive if it is of $\mathrm{Kum}^n$-type (with Artin invariant $\neq$ 3). In particular, any product of those irreducible symplectic varieties satisfies the supersingular Tate conjecture for the whole $\ell$-adic or crystalline cohomology ring.

math.AG↗

More Choices, Fewer Decisions: Ordinal-Scale Bias in JEV-like Direct-Decision Models

Direct-decision models turn text into low-latency structured labels and scores, making them attractive for classification and automatic evaluation. Yet reliability requires more than accuracy: a model must also use the ordinal decision scale supplied by the user faithfully. We analyze JEV~1.13 and three open KEV models. Our investigation begins with ANLI, where JEV assigns 38.8\% of all predictions and 51.3\% of errors to Neutral despite 74.95\% accuracy, nearly balanced gold labels, and balanced candidate positions. Across 36 ordinal datasets, final decisions use only 67--76\% of the effective gold support, versus 87--102\% on four nominal tasks. Randomizing candidate order weakens but does not remove this compression. Holding items and source scores fixed while balancing gold support and positions, we refine scales from $K=2$ to $14$; utilization falls for every model and reaches 26--75\% at $K=14$, although candidate probabilities remain broad for most models. Targeted BA-LoRA post-training raises gold-relative utilization from roughly 47\% to 86\% on eight supervised scales at both KEV sizes, showing that the compression is learned and modifiable rather than an immutable architectural limit. We call this ordinal scale-utilization bias: decision-stage candidate-space compression distinct from accuracy, gold imbalance, fixed position, and candidate count alone. The code and data are available at https://github.com/Glax147/jev_ordinal_scale_bia

cs.AI↗

Controlling Switching Evolution in Lead-Free Perovskite-Inspired Chalcogenide Memristors for Neuromorphic Computing

Memristors have emerged as key building blocks of neuromorphic computing architectures due to their ability to integrate data storage and processing. While metal halide perovskites have recently shown significant promise owing to their mixed electronic-ionic conduction and low-cost solution processability, their reliance on toxic lead and limited stability presents critical challenges. Here, we report environmentally friendly, low-toxicity AgBiS2-based solution-processable memristors exhibiting an ultra-low SET voltage of ~0.08 V and a high ON/OFF ratio of >104. First-principles calculations identify Ag interstitials as the energetically most favourable native defect and reveal low migration barriers within the Ag sublattice, for both interstitials and vacancies, facilitating ionic transport in the AgBiS2 lattice. Through interface and thickness engineering, the resistive switching behaviour can be systematically tuned from abrupt digital to gradual analog modes. Notably, thicker switching layers promote the evolution of stable conductive pathways through intermediate metastable states, revealing a controllable filament evolution process. Electrochemical impedance spectroscopy reveals pronounced negative capacitance (inductive) behaviour at low bias voltages, arising from coupled electronic-ionic dynamics. Consistent with this behaviour, pulse measurements demonstrate gradual conductance modulation under pulse trains, emulating synaptic responses relevant for neuromorphic computing. Finally, post-operando structural analysis reveals substantial morphological evolution of the switching layer driven by repeated filament formation and rupture. Linking structural dynamics to switching variability provides important design principles for achieving reliable and durable sustainable memristors.

cond-mat.mtrl-sci↗

H-WM: Robotic Task and Motion Planning Guided by Hierarchical World Model

World models are becoming central to robotic planning and control by predicting future state transitions. Existing approaches mainly rely on visual, latent, or language prediction, which can be difficult to ground in executable robot actions and prone to compounding errors over long horizons. In contrast, traditional robotic task and motion planning enables structured long-horizon reasoning through compact symbolic representations of world transitions, but typically lacks synchronized visual prediction. We propose Hierarchical World Model (H-WM), which jointly predicts logical and visual state transitions by combining a high-level logical world model with a low-level visual world model. The predicted logical actions and latent visual state transitions are jointly incorporated into Vision-Language-Action (VLA) models as intermediate state guidance for long-horizon task execution. Experiments on three long-horizon benchmarks and real robots show that H-WM consistently improves VLA's performance by stabilizing long-horizon execution and mitigating error accumulation. We also construct LIBERO-Logic, a frame-level aligned dataset that pairs visual observations and continuous robot states with logical actions and predicate-based logical states.

cs.RO↗

RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning

Long-horizon robotic tasks require a breadth of capabilities beyond what any single existing robot control policy can reliably provide. Combining heterogeneous policies with complementary strengths offers a promising solution, but introduces two key challenges: uncertain capability boundaries and distribution mismatches during policy handoffs. These challenges remain largely unaddressed by existing planning methods, which typically assume homogeneous, predefined skills with fixed applicability. We propose RoboHarness, a unified framework that encapsulates independently developed heterogeneous policies, including vision-language-action models (VLAs), world-action models (WAMs), reinforcement learning (RL) policies, and task and motion planners (TAMP), as reusable agentic skills. RoboHarness integrates understanding, memory, and evolution skills to reason about policy capabilities and support capability-aware task decomposition and policy routing. To mitigate distribution mismatches during policy handoffs, we introduce Memory Bridge, a plug-in policy-chaining mechanism that enables reliable transitions between heterogeneous policies without joint retraining. Extensive experiments across five public benchmarks, 500 customized tasks across 10 classes, and 135 real-robot trials demonstrate substantial gains in long-horizon and memory-dependent tasks, as well as robustness to out-of-distribution conditions.

cs.RO↗

Experimental demonstration of broadband-laser suppression of cross-beam energy transfer

In direct-drive inertial confinement fusion (ICF), cross-beam energy transfer (CBET) redirects a significant fraction of the incident laser energy out of the plasma. We report the experimental demonstration that broadband lasers reduce CBET-enhanced reflected-light return, performed at the low-coherence Kunwu laser facility with two crossed beams of 0.6% bandwidth and up to 550J at $\sim$2.6$\times$10$^{14}$ W/cm$^{2}$. A coupled ray-tracing model shows that the strong CBET amplification of narrowband reflected light is much weaker under broadband illumination. In symmetric incidence condition of two orthogonal beams, the total fractional scattered energy decreases from 7.57% to 4.64%; in asymmetric incidence condition, which separates stimulated Brillouin scattering (SBS) from specular reflection, the SBS fraction decreases from 3.75% to 0.46% and the specular-reflection fraction decreases from approximately 2.2% to 1.3%. These results establish broadband lasers as an effective, experimentally validated approach to reducing energy escape through CBET-enhanced reflection in direct-drive ICF.

physics.plasm-ph↗

Beyond Semantic Narrowing: Robust and Efficient LLM Watermarking with Hamming Neighborhoods

Semantic watermarking improves robustness against watermark removal attacks by embedding detectable signals into sentence-level representations. However, existing watermarking methods typically impose watermark-specific semantic preferences on generated sentences without explicitly accounting for the highly non-uniform and context-dependent semantic preference of LLM generation. When these two preferences are poorly aligned, many natural continuations become incompatible with the watermark, causing semantic narrowing: reduced semantic freedom, increased resampling cost, and potential degradation on tasks with strict semantic requirements. To alleviate this problem, we propose HammingMark, which uses the semantic hash of the preceding sentence as a dynamic center and accepts candidates whose hashes fall within its Hamming neighborhood. Defining watermark validity over a Hamming neighborhood in compact hash space retains a larger fraction of naturally likely semantic continuations. The coarse many-to-one hash mapping further allows diverse semantic realizations to remain watermark-valid. Experiments on C4 and BookSum show that HammingMark achieves strong robustness, high detectability, and near-unwatermarked generation quality, requiring only 2.2 sampled candidates per accepted sentence,a 72.8% reduction compared with the most sampling-efficient existing method. On more complex tasks with strict semantic constraints, HammingMark achieves the highest detection rates with the highest or tied-highest ROUGE-L scores, demonstrating its effectiveness in balancing watermark detectability and generation quality under constrained generation settings.

cs.CR↗

Why Cross-Skeleton Retargeting Is Non-Identifiable: Structural Limits of Generative Motion Models

Cross-skeleton motion generation trains generative models to carry action structure and motion intention from one body to another. Yet a target motion that shows the right action has two explanations that the training data cannot tell apart: the model transferred the source clip, or it recovered a typical motion for the requested action. We show that this ambiguity is structural rather than incidental: under standard generative objectives, the source-conditioned retargeting map is non-identifiable in sparse heterogeneous motion domains. Unpaired distribution matching yields gauge non-identifiability: the latent spaces of different skeletons can be transformed relative to one another without changing the training evidence, so different source-conditioned maps fit it equally well. Sparse paired supervision admits the complementary failure mode, \emph{conditional-mean degeneration}: when clips are paired only by action, squared-error training converges to an average target motion that ignores the source clip. To make the missing evidence observable, we introduce Source-Instance Fidelity (SIF), a diagnostic that tests whether outputs differ from one another the way their source clips do, with the target skeleton and action held fixed. Under this diagnostic, methods that succeed at the standard action-level test on animal motion data often sit at the source-blind floor, while the methods that rise above it retain only a partial relational signal. Retargeting therefore needs objectives and evaluations that can identify the source-conditioned map it claims to learn. Project page: https://cross-skeleton-retargeting.netlify.app/.

cs.CV↗

Spatial Grafting: Grounding 3D Features for Flow-Matching Robot Policies

Pretrained robot manipulation policies such as vision-language-action models (VLAs) or world-action models (WAMs) leave interaction-relevant metric geometry implicit. Recent breakthroughs in spatial reconstruction can supply the necessary geometry reliably, but their features describe local shape without stating where it lies with respect to the robot. How best to deliver these features to a pretrained policy remains unresolved. We propose Spatial Grafting, a versatile, lightweight spatial module that binds frozen reconstruction features to metric, robot-relative geometry. Spatial Grafting constructs metric-grounded spatial tokens and injects them into the flow-matching action expert through cross-attention, without modifying the host's perceptual pathway, so the host retains the full benefit of its pretraining. We evaluate it more broadly than any geometry-aware policy we compare against: one graft architecture, with no per-host redesign, on two VLAs and two WAMs, across four simulation benchmarks that span short-horizon manipulation, visual robustness, clutter and long-horizon mobile manipulation, and on three real-robot platforms with single- and dual-arm configurations. On RoboTwin 2.0, a dual-arm manipulation benchmark, the graft improves every host across VLAs and WAMs. Grafted $π_{0.5}$ gains 11.3% and 15.6% on clean and randomized scenes, reaching 94.0% and 92.4%, above the strongest published 3D-conditioned policy, WAM4D (93.8% and 89.9%). The margin widens as the horizon lengthens: on tasks from BEHAVIOR-1K, a dual-arm mobile manipulation challenge scored by average task progress, it surpasses the 2025 challenge winner on five of six tasks,by up to 0.47 Q-score, and exceeds a map-conditioned spatial policy on average across the three tasks both report.

cs.RO↗

INFUSER: Influence-Guided Self-Evolution Improves Reasoning

Self-evolution offers a scalable path to stronger reasoning: a pretrained language model improves itself with only minimal external supervision. Yet existing methods either depend on extensively curated or teacher-generated training data, or, when the generator runs unsupervised, reward it by a difficulty heuristic that need not improve the solver. We introduce INFUSER, an iterative co-training framework with two co-evolving roles: a Generator that drafts questions and reference golden answers from a pool of unstructured, automatically collected documents, and a Solver that improves by training on them. The solver is trained with standard correctness rewards against the generator-provided answers, while the generator is rewarded by an optimizer-aware influence score that measures whether each proposed question would actually improve the solver on the target distribution. Because this continuous, noisy influence score is poorly served by standard GRPO, we propose DuGRPO, a dual-normalized variant of GRPO, for generator training. Together, these turn the document pool into an adaptive curriculum that favors questions useful to the current solver, not just hard ones. On Qwen3-8B-Base, INFUSER outperforms strong self-evolution baselines with over 20% relative improvement on Olympiad and SuperGPQA benchmarks, and an 8B INFUSER co-evolving generator outperforms a frozen 32B thinking generator on math and coding. Ablations confirm each design choice is necessary, and two extensions, applying INFUSER to an instruction-finetuned anchor and augmenting it with rule-verifiable RLVR data, further demonstrate the flexibility and generalizability of the framework. Code is available at https://github.com/FFishy-git/INFUSER.

cs.LG↗

Towards Understanding Momentum Acceleration in River-Valley Loss Landscape

The empirical success of pretraining large language models has inspired a deeper investigation into the underlying loss landscapes and the optimization dynamics. Recent empirical and theoretical study suggest that the training loss landscape often exhibits a "river-valley" structure, which features a low-loss manifold (river) flanked by sharp orthogonal directions with higher loss (mountains). In the long term, the optimization progress is determined primarily by the progress along the river. Within such a landscape, gradient descent with large learning rates can move faster along the river despite high apparent loss due to vertical oscillations, while a subsequent sharp decay in the learning rate suppresses these oscillations, revealing genuine optimization progress. This explains the recent success of warmup-stable-decay (WSD) learning rate scheduler which, unlike cosine scheduling, keeps stable high learning rate and decays before producing intermediate checkpoints. Building on this foundation, in this work we take a step further and study the role of momentum within such a loss landscape. We establish theoretical analysis that characterizes how momentum accelerates optimization by stabilizing large learning rates that can not be tolerated by vanilla GD without deviating significantly from the river. The enabled large learning rate in-turn gives greater speed along the river and makes faster essential progress in the long run. Another intriguing observation from theory is that for a river-valley landscape with very flat and slow-spinning river, the momentum itself does not contribute directly to acceleration in terms of the speed of tracking the river, while the main acceleration comes from the admissible larger learning rate.

cs.LG↗

Simultaneous recovery of multiple parameters in nonlocal diffusion equations from internal measurements

This paper is devoted to simultaneously recovering multiple parameters from internal measurements for nonlocal diffusion equations. The uniqueness of the inverse problem is established by employing the asymptotic behavior of solutions, analytic continuation, the Laplace transform, and properties of analytic functions. For numerical reconstruction, we apply the Levenberg-Marquardt method to obtain a stable approximate solution of the inverse problem. Numerical examples are provided to demonstrate the efficiency of the proposed algorithm and to validate our theoretical findings.

math.NA↗

Rethinking Pairwise Token Interaction in Spiking Transformers

Spiking Transformers inherit token interaction mechanisms from conventional Transformers, yet their sparse binary representations fundamentally alter how token-to-token communication is established. In particular, spike-based query-key matching produces highly sparse and input-dependent interaction patterns, coupling information propagation to the instantaneous availability of matching spike events. This motivates a different interaction paradigm in which long-range communication does not rely solely on pairwise spike coincidence. We therefore propose Gated Spike Axial Propagation (GSAP), a spike-native token interaction mechanism that decouples information propagation from context selection. Instead of directly determining communication through query-key matching, GSAP first propagates spike-based context along the horizontal and vertical axes, allowing information to reach distant tokens through structured sequential propagation. A receiver-conditioned gate then determines how much of the propagated context is incorporated at each token, while a lightweight local pathway preserves fine-grained neighborhood information. In this way, GSAP reformulates token interaction as a propagate-then-select process, enabling structured long-range communication while retaining the sparse event-driven nature of spiking representations. Code is available at https://github.com/Fancyssc/GSAP.

cs.NE↗

Constraints on the Hot Circumgalactic Medium around Nearby L* Galaxies from SRG/eROSITA All Sky Survey

The circumgalactic medium (CGM), a multi-phase gas acting as the dynamic interface between the main body of a normal galaxy and the intergalactic medium, is a key ingredient of the galactic ecosystem and provides crucial diagnostics to a wide array of physical processes affecting galaxy evolution. However, direct evidence for a predominantly hot (at a characteristic temperature of a few million Kelvin) CGM around present-day $L^*$ galaxies remains elusive, despite extensive observational searches and strong physical reasons that it should exist. Here, we present a systematic search for the hot CGM around a representative sample of nearby $L^*$ galaxies selected from the 50 Mpc Galaxy Catalog (50MGC), by stacking their X-ray images and spectra from the SRG/eROSITA all-sky survey. Significant ($7.4σ$) diffuse X-ray emission is detected out to a galactocentric radius $\gtrsim$ 50 kpc, the cumulative spectrum of which requires a substantial thermal component, consistent with a hot plasma with log-normal temperature distribution, and thus argues against a predominantly non-thermal origin. The diffuse emission exhibits a relatively steep intensity profile ($β=0.52_{-0.06}^{+0.08}$), inferring a moderate amount ($2\times10^{10}~M_{\odot}$) of hot gas within 10--200 kpc. The radial distribution and total amount of the hot CGM gas are broadly in agreement with prediction by IllustrisTNG simulations for present-day Milky Way-like galaxies, while also showing dependencies on both star formation and nuclear activity. The constraints on the hot CGM derived in this study hold promise for calibrating key physical processes, such as stellar feedback and active galactic nucleus feedback, in next-generation cosmological simulations.

astro-ph.GA↗

StageGuard: Learning Stage Transitions for Long-Horizon Robot Tasks via Agentic Distillation

Hierarchical planning frameworks combine skills from multiple robot control policies for long-horizon task execution, where determining when to terminate the current skill and advance to the next subtask is essential. Existing approaches often rely on pre-designed completion signal checkers that are hard to obtain in real-world execution. Large-scale vision-language models (VLMs) offer strong reasoning capabilities, but their decision boundaries are not inherently aligned with task completion criteria, while cloud deployment and lengthy reasoning introduce substantial latency, limiting real-time monitoring. We propose StageGuard, an agentic distillation framework for accurate and efficient stage-transition decisions. StageGuard combines teacher-model reasoning with demonstration trajectories to generate structured explanations of subtask completion and policy switching. A lightweight student VLM uses these explanations to generate compact self-explanations, which are used for supervised fine-tuning. We evaluate stage-transition prediction on trajectories from two benchmarks and assess closed-loop task success through integration into hierarchical robot control on BEHAVIOR-1K, with further validation on real robots. Results show substantial improvements in stage-transition prediction while supporting efficient online monitoring.

cs.RO↗

Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation

Recent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have made them powerful tools in embodied navigation, enabling agents to leverage commonsense and spatial reasoning for efficient exploration in unfamiliar environments. Existing LLM-based approaches convert global memory, such as semantic or topological maps, into language descriptions to guide navigation. While this improves efficiency and reduces redundant exploration, the loss of geometric information in language-based representations hinders spatial reasoning, especially in intricate environments. To address this, VLM-based approaches directly process ego-centric visual inputs to select optimal directions for exploration. However, relying solely on a first-person perspective makes navigation a partially observed decision-making problem, leading to suboptimal decisions in complex environments. In this paper, we present a novel vision-language model (VLM)-based navigation framework that addresses these challenges by adaptively retrieving task-relevant cues from a global memory module and integrating them with the agent's egocentric observations. By dynamically aligning global contextual information with local perception, our approach enhances spatial reasoning and decision-making in long-horizon tasks. The proposed method surpasses previous state-of-the-art approaches by a significant margin on both the HSSD and HM3D benchmarks and demonstrates strong performance on a real robot.

cs.RO↗

A Scenario-Knowledge-Driven Pipeline for Just-in-Time Assistance

Detecting a silently struggling kiosk user is only the first step; deciding whether, when, and how to help depends on scenario knowledge usually buried in model weights and thresholds. We propose a scenario-knowledge-driven pipeline: a single scenario knowledge document, human-authored and version-controlled, configures sensing, constrains LLM reasoning, and shapes a graded intervention proposal. Narration, assistance-need assessment, and proposal are kept separate for independent audit. As proof of concept, we replay two recorded kiosk sessions offline, chosen before the runs for their struggle evidence and retrospective detail. Both cases support what the design promises: checkable reporting and measured escalation. Across 95 updates, every sentence of the append-only narration cites the primitive events underlying it, and the rule layer detects 12 of 13 and 7 of 7 annotated struggle episodes under a strict criterion. The assessor de-escalates on recovery and reaches the top rung exactly once, under maximally converging evidence. At the decisive help-seeking turn, narration, assessment, and the participants' retrospective accounts converge. The appropriateness of these interventions, the pipeline's restraint on sessions without struggle, and the document's transfer to a new scenario frame the agenda.

cs.HC↗