Search arXivSearch

arXiv subjects

Zirui Li

Publications and source records attributed to Zirui Li.

At least 19 recordsLinked to original sources

EmphTTS: an emphasis-control TTS with reinforcement learning

Generating controllable and human-like emphasis remains an open challenge in text-to-speech, even when explicit emphasis control signals are provided in the text input, limiting the communicative accuracy of synthetic speech in real-world applications. Reinforcement learning has recently shown promise for post-training TTS systems to align with human preference, yet existing methods have not been applied to word-level prosodic control. We present EmphTTS, a non-autoregressive TTS system that applies Group Relative Policy Optimization (GRPO) to the duration predictor with an emphasis localization reward, enabling direct optimization for word-level emphasis. Evaluations show that EmphTTS achieves the best emphasis controllability and performs the best in emphasis objective evaluation. In subjective preference tests, EmphTTS is significantly preferred over synthetic groundtruth and most baselines. Ablation studies show that GRPO improves emphasis realization beyond supervised-finetuning-based duration modeling and simple speaking-rate adjustment, while alleviating the mismatch between the independently trained duration predictor and TTS model.

eess.AS

Beyond Final-Token Classification: Heterogeneous Readouts for Evidence-Grounded Suicide Risk Detection

The IEEE BigData Cup benchmark combines three prediction problems with different output structures: ordinal suicide-risk classification, multi-label psychosocial factor detection, and extraction of supporting phrases. We introduce heterogeneous readout decomposition (HRD), which separates semantic verification from output realization. A locally deployed Qwen3.8-27B model, adapted with task-specific QLoRA adapters, produces both answer-token margins and layer-63 answer states for card-conditioned queries. HRD compares four latent scores for ordinal risk, retains token margins for most factors while routing seven labels through one shared latent probe, and constructs evidence sets from verbatim span candidates with calibrated, risk-conditional constraints. On two held-out user-grouped confirmation folds, the latent risk readout improves weighted F1 from 0.8237 to 0.8372 and macro F1 from 0.7965 to 0.8185. Selective factor routing improves macro F1 by 0.0105 and tail-label macro F1 by 0.0189; in contrast, global latent replacement and independent label-specific probes fail. The constrained evidence decoder raises pooled phrase F1 from 0.7488 to 0.7609 in row-level out-of-fold evaluation. Lenormand's best public result is 0.8052 on Subtask 1 and 0.6636 on Subtask 2, giving a composite score of 0.7627. These results identify the answer-token boundary, rather than semantic representation alone, as a measurable source of error in this benchmark.

cs.CL

NLPCC 2026 Task 10: Citation-Level Faithfulness Verification with DeBERTa Ensembles and Class-Wise Calibration

This paper presents our system for Track 2 of the NLPCC 2026 Shared Task 10 on citation-level faithfulness in AI-assisted scientific reporting. Given an atomic scientific claim and the structured full text of its cited paper, the task requires both a four-way relation label and up to three evidence paragraph identifiers. The label head ensembles a paragraph-aware cross-encoder with a document-level DeBERTa-large classifier, followed by class-wise decision calibration. Probability-level fusion is motivated by an out-of-fold tendency to over-predict Topical Match. The evidence head combines paragraph scores from top-20 and top-30 joint models with BM25 scores. The system runs fully offline without external retrieval or LLM prompting. On the final leaderboard, our system achieved 82.9898 overall (89.5491 Macro-F1 and 76.4305 Joint@3), ranking second in Track 2. Ablations and error analysis show that model complementarity and calibration drive the label gains. Gold-evidence inference changes label Macro-F1 negligibly, whereas evidence ranking remains important for Joint@3.

cs.CL

Ego-Dynamics-Augmented World Model for Autonomous Driving with Zero-Shot Cross-Embodiment Adaptation

End-to-end autonomous driving requires generalization ability across platforms with dissimilar physical characteristics. The chassis defines the physical embodiment of each platform, and real-world fleets span sub-tonne microcars to bus-class vehicles. Consequently, the driving stack must either be retrained per platform or adapt to the underlying chassis dynamics online. World model (WM)-based reinforcement learning offers a sample-efficient path toward end-to-end autonomous driving on egocentric bird's-eye-view (BEV) representations, but its effectiveness hinges on how faithfully the WM captures the ego vehicle's dynamics. This work identifies a structural bottleneck in BEV-based WMs: observation transitions entangle ego-motion with scene dynamics, consuming modeling capacity at the cost of imagination accuracy. This burden is embodiment-dependent: dissimilar chassis produce different observation warps under the same control input. The proposed DynaDreamer addresses this bottleneck by conditioning the WM's latent distributions on a physics-informed ego-dynamics context derived from a lateral dynamics model with a neural tire force formulation. This context is extracted online via a neural-ODE encoder-decoder that simultaneously identifies the underlying chassis parameters. Information-theoretic analysis confirms that this conditioning removes the ego-motion terms from both the WM's transition entropy and its prior-posterior KL divergence. The identified physical parameterization enables zero-shot cross-embodiment adaptation across a dynamically diverse fleet without per-platform retraining. Simulation results show 28% and 43% improvements in driving task success rates over the strongest baseline in urban and highway scenarios, and the advantage over the base Transformer WM reaches up to 73% when extrapolating to unseen chassis.

cs.RO

SSP-DMGTimeNet: Physics-Constrained Learning for Spatiotemporal Trajectory Prediction of Vehicle Platoons

Existing car-following prediction methods mainly optimize trajectory accuracy, while rarely considering whether predicted disturbances propagate realistically along a vehicle platoon. This limitation may lead to accurate but string-unstable predictions. We propose SSP-DMGTimeNet, a physics-constrained learning framework for spatiotemporal trajectory prediction of vehicle platoons. The model combines multi-scale temporal representations with cross-vehicle interaction features to capture complex and time-varying platoon dynamics. A propagation-delay-aware causal attention mechanism explicitly models upstream-to-downstream disturbance propagation by learning response delays between adjacent vehicles and accumulating them along the platoon. In addition, time- and frequency-domain string-stability losses relieve disturbance amplification across both adjacent vehicles and arbitrary sub-platoons during training. Experiments on HighD show that SSP-DMGTimeNet achieves an unstable-window rate of 0.65\% for five-vehicle platoons and a maximum head-to-tail amplification of 0.898 on the ground-truth excitation subset, while maintaining competitive trajectory prediction performance. In zero-shot evaluation on NGSIM US-101 and I-80, the model achieves velocity MAEs of 1.316~m/s and 1.252~m/s, with unstable-window rates of 3.90\% and 4.10\%, respectively. These results demonstrate that incorporating platoon-level physical constraints can effectively balance trajectory prediction accuracy and disturbance propagation stability.

cs.AI

Modulational spectrum of infinite-depth hydroelastic Stokes waves

We determine the complete local Bloch spectrum bifurcating from the origin for small-amplitude periodic hydroelastic Stokes waves in infinite depth, under the combined effects of gravity, surface tension, and elastic bending. Away from the Wilton-type resonance set, we construct a real-analytic Stokes-wave branch and analyze the four eigenvalues emerging from the defective zero eigenvalue of the linearized hydroelastic Euler system. Using analytic spectral perturbation theory and Hamiltonian-reversible reductions, we decouple them into a Benjamin--Feir pair and a long-wave pair. The long-wave pair remains purely imaginary and has the singular scale $\cO(\sqrt{|μ|})$, whereas the Benjamin--Feir pair is governed by an explicit discriminant whose leading sign yields a sharp criterion for modulational stability and instability. We derive the exact non-resonant phase diagram in the surface-tension-bending parameter plane and identify a bounded stability island generated by elastic bending. In the unstable region, and away from a drift degeneracy, the Benjamin--Feir branches form a local figure-eight curve. In the zero-bending limit, the reduced coefficients recover the known deep-water gravity and gravity-capillary results, while the change from the finite-depth $\cO(|μ|)$ long-wave scale to $\cO(\sqrt{|μ|})$ shows that the infinite-depth problem is singular.

math.AP

Universal mapping of drop impact spreading from wetting to Leidenfrost regimes

We establish a universal mapping between the maximum spreading of a drop under wetting and Leidenfrost impact conditions from an energy-dissipation perspective, with the latter featuring a stable vapor film between the drop and the substrate. Experiments and direct numerical simulations demonstrate that the mapping remains valid over two to three orders of magnitude in the Ohnesorge and Weber numbers. This framework provides a simple route for converting established predictions for wetting impacts into their Leidenfrost counterparts, with implications for predicting and controlling droplet impact.

physics.flu-dyn

LLM-Assisted Semantic Alignment and Integration in Collaborative Model-Based Systems Engineering Using SysML v2

Cross-organizational collaboration in Model-Based Systems Engineering (MBSE) faces many challenges in achieving semantic alignment across independently developed system models. SysML v2 introduces enhanced structural modularity and formal semantics, offering a stronger foundation for interoperable modeling. Meanwhile, GPT-based Large Language Models (LLMs) provide new capabilities for assisting model understanding and integration. This paper proposes a structured, prompt-driven approach for LLM-assisted semantic alignment of SysML v2 models. The core contribution lies in the iterative development of an alignment approach and interaction prompts, incorporating model extraction, semantic matching, and verification. The approach leverages SysML v2 constructs such as alias, import, and metadata extensions to support traceable, soft alignment integration. It is demonstrated with a GPT-based LLM through an example of a measurement system. Benefits and limitations are discussed.

cs.SE

Learning from Mistakes: Rollout-Retrieval Lifelong Policy Learning for Autonomous Driving

Autonomous driving policies should be able to improve continually as deployment exposes them to increasingly diverse and long-tail traffic situations. However, most learning-based policies are trained or fine-tuned on expert demonstrations and then rely largely on generalization to handle challenging closed-loop scenarios, lacking an explicit mechanism to correct and retain the mistakes exposed in these scenarios. This paper studies autonomous driving policy improvement from a lifelong learning perspective: Can a pretrained policy improve continually by accumulating corrective knowledge derived from its own mistakes, while retaining previously acquired driving competence? To answer this question, we propose Rollout-Retrieval Lifelong Policy Learning (R$^2$LPL), a policy learning framework that retrieves corrective targets from recoverable policy-induced mistakes and retains the resulting knowledge through lifelong policy learning. R^2LPL addresses a key bottleneck in continual policy improvement: closed-loop mistakes reveal where the policy is weak, but do not directly specify what the policy should learn. By filtering recoverable mistake-related states and retrieving feasible corrective targets, R$^2$LPL turns sparse failure evidence into compact supervised knowledge for stable and sample-efficient policy improvement. We evaluate R$^2$LPL on large-scale closed-loop nuPlan benchmarks. With only a few rollout and continual-learning cycles, R$^2$LPL elevates a learning-based planner with moderate initial performance to state-of-the-art performance across the evaluated benchmarks, especially on the challenging and long-tail Test14-hard split. These results demonstrate the effectiveness of R$^2$LPL in converting recoverable closed-loop mistakes into corrective knowledge for sustained policy improvement.

cs.RO

Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout

Safe L2/L3 driving automation requires anticipating human-in-the-loop reactions during shared-control transitions. While most driving world models forecast the external environment, in-cabin intelligence remains strictly recognition-oriented and lacks multi-step rollout capabilities for driver dynamics. We introduce Driver-WM, a driver-centric latent world model that rolls out in-cabin dynamics causally conditioned on out-cabin traffic context. This formulation unifies physical kinematics forecasting with auxiliary behavioral and emotional semantic recognition. Operating in a compact latent space constructed from frozen vision-language features, Driver-WM adopts a dual-stream architecture to separately encode external traffic and internal driver states. These streams are directionally coupled via a gated causal injection mechanism, which uses a learned vector gate to modulate external contextual perturbations while strictly enforcing temporal causality. Experiments on AIDE show robust long-horizon forecasting on reactive high-motion clips, improved driver/traffic semantic alignment, and controlled interventions that expose the external-to-internal mechanism.

cs.RO

DCP-Prune: Ultra-Low Token Pruning with Distribution Consistency Preservation

Recent vision token pruning methods effectively preserve model performance under moderate token budgets but become unstable under ultra-low token budget. Our analysis shows that as the pruning budget decreases, accuracy degradation is often accompanied by larger feature distribution shifts. Critically, the degree of this distribution shift strongly correlates with performance degradation. To better characterize this phenomenon, we introduce a lightweight distribution consistency metric to estimate the distribution shift between retained and full tokens. Motivated by these observations, we propose a two-stage pruning framework consisting of Anchor-Context Graph Recovery (ACGR) and Text-Aware Token Cluster Selection (TATCS). Specifically, ACGR transfers contextual information before token removal, while TATCS dynamically re-selects representative tokens when severe distribution shift is detected. Extensive experiments demonstrate that our method achieves superior and more stable performance under ultra-low token budget. Notably, it retains 92.1% of the upper-bound average performance on LLaVA-1.5-7B with only 16 visual tokens.

cs.CV

Benjamin-Feir spectrum of hydroelastic Stokes waves

We determine the complete Benjamin-Feir spectrum near the origin for small-amplitude hydroelastic Stokes waves of the two-dimensional finite-depth irrotational Euler equations with surface tension and elastic bending. For the non-resonant Stokes branch and away from an intrinsic characteristic-collision surface $\mathfrak D$, we resolve all four Bloch eigenvalues bifurcating from the origin in the long-wave Floquet regime. Exploiting the Hamiltonian and reversible structure of the problem, we reduce the linearized Bloch operator to the four-dimensional spectral subspace bifurcating from the generalized kernel at the origin and conjugate the resulting matrix to the direct sum of a Benjamin-Feir block and a long-wave block. The long-wave pair remains purely imaginary, whereas the Benjamin-Feir pair is governed by an explicit closed-form instability index $\operatorname{Ind}(\mathtt{h},κ,b)$: a positive index produces a local figure-eight spectral curve with nonzero real part, while a negative index implies that all four small eigenvalues remain purely imaginary. Together with the Wilton-type resonance loci and the characteristic-collision surface $\mathfrak D$, this index yields a three-parameter spectral-stability diagram in the depth $\mathtt{h}$, surface tension $κ$, and bending rigidity $b$. The diagram recovers the classical pure-gravity critical-depth limit and, on the zero-bending boundary, the gravity--capillary stability diagram. It also reveals a genuinely hydroelastic phenomenon: all Wilton-type resonances disappear whenever $b\geq 1/14$ or $κ\geq 1/2$. This provides the first complete rigorous characterization of the local Benjamin-Feir spectrum for a hydroelastic free-boundary problem.

math.AP

IntentNav: Learning Spatial-Visual Object Navigation from Human Demonstrations

Object navigation requires a robot to search for an unobserved target in an unknown environment by deciding where to explore next under partial observability. Effective search resembles human-like exploration: selectively probing visually promising frontiers while relying on spatial memory to avoid redundant revisits. We propose IntentNav, a spatial-visual imitation framework that learns human-like ObjectNav policies from human demonstrations. To infer high-level search intent from low-level human actions, we introduce Frontier-based Human-Intent Labeling, which looks ahead in human demonstrations and labels the frontier that best explains the demonstrator's future search direction. We construct a spatial-visual candidate space, where BEV memory tracks explored regions, unexplored frontiers, and trajectory history, while egocentric visual memory provides semantic cues for each candidate. A VLM policy is trained to select among these grounded candidates, using Intent-Aligned Objective to encourage consistent and human-like exploration. IntentNav achieves state-of-the-art performance on the MP3D, HM3D-v1 and HM3D-v2 ObjectNav benchmarks. The proposed candidate-level navigation interface transfers zero-shot to wheeled, quadruped, and humanoid robots without further VLM fine-tuning. \href{https://anonymous.4open.science/w/IntentNav/}{Project page}.

cs.RO

Linear Complexity Fermionic Simulation on Quantum Devices with Hardware Connectivity Constraints

Simulating fermionic systems on quantum hardware requires compiling fermionic Hamiltonians into executable quantum circuits. Existing approaches treat each compilation stage independently, applying heuristics with localized objectives that produce circuits with superquartic gate count and depth scaling and compilation times reaching several hours for large instances. We present Accordion, an end-to-end framework that co-designs the fermion-to-qubit mapping with circuit synthesis and hardware routing. Accordion fixes the Jordan Wigner mapping, which despite its higher Pauli weight produces Pauli operators with structural regularity that enables provably efficient circuit generation. For full-rank all-to-all electronic structure Hamiltonians, we prove O(N^4) gate count and circuit depth, matching the information-theoretic lower bound imposed by the Theta(N^4) second excitation terms. On linear, IBM heavy-hex, and square-grid architectures, Accordion reduces gate count by up to 79% and circuit depth by up to 77% relative to the best baseline.

cs.AR

Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal Structure

Latent or continuous chain-of-thought methods replace explicit textual rationales with a number of internal latent steps, but these intermediate computations are difficult to evaluate beyond correlation-based probes. In this paper, we view latent chain-of-thought as a manipulable causal process in representation space by modeling latent steps as variables in a structural causal model (SCM) and analyzing their effects through step-wise do-interventions. We study two representative paradigms (i.e., Coconut and CODI) on both mathematical and general reasoning tasks to investigate three key questions: (1) which steps are causally necessary for correctness and when answers become decodable early; (2) how influence propagates across steps and how this structure compares to explicit CoT; and (3) whether intermediate trajectories retain competing answer modes and how output-level commitment differs from representational commitment across steps. We find that latent-step budgets behave less like homogeneous extra depth and more like staged functionality with non-local routing, and we identify a persistent gap between early output bias and late representational commitment. These results motivate mode-conditional and stability-aware analyses, together with corresponding training/decoding objectives, as more reliable tools for interpreting and improving latent reasoning systems. Code is available at https://github.com/J1mL1/causal-latent-cot.

cs.AI

JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation

We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF materials released by the Japanese Ministry of Health, Labour and Welfare, JMed48k contains 48,862 exam questions and 20,142 images from 11 national licensing examinations between 2005 and 2025, with visual content annotated under an 8-type taxonomy. From this corpus, we derive JMed48k-Eval, a recent five-year evaluation subset with 12,484 scored questions, including 9,905 text-only questions and 2,579 questions with images. We evaluate 21 proprietary, open-source, and medical-specific models, reporting text-only and with-image performance separately. Because these subsets contain different questions, we further introduce a paired image-removal audit that evaluates questions with images before and after removing visual content to explore four answer-transition states. The audit shows that proprietary and open source models gain substantially from images, whereas medical-specific systems show limited observable use of visual evidence, with many correct answers persisting after image removal. Even among proprietary models, the net image-removal effect varies sevenfold across professions, from +5.7 points on Physician questions to +39.8 points on Public Health Nurse questions. We release JMed48k to support reproducible, profession-stratified evaluation of vision-language models in medical licensing settings.

cs.CV

JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models

As Large Language Models (LLMs) are increasingly deployed in healthcare field, it becomes essential to carefully evaluate their medical safety before clinical use. However, existing safety benchmarks remain predominantly English-centric, and test with only single-turn prompts despite multi-turn clinical consultations. To address these gaps, we introduce JMedEthicBench, the first multi-turn conversational benchmark for evaluating medical safety of LLMs for Japanese healthcare. Our benchmark is based on 67 guidelines from the Japan Medical Association and contains over 50,000 adversarial conversations generated using seven automatically discovered jailbreak strategies. Using a dual-LLM scoring protocol, we evaluate 27 models and find that commercial models maintain robust safety while medical-specialized models exhibit increased vulnerability. Furthermore, safety scores decline significantly across conversation turns (median: 9.5 to 5.0, $p < 0.001$). Cross-lingual evaluation on both Japanese and English versions of our benchmark reveals that medical model vulnerabilities persist across languages, indicating inherent alignment limitations rather than language-specific factors. These findings suggest that domain-specific fine-tuning may accidentally weaken safety mechanisms and that multi-turn interactions represent a distinct threat surface requiring dedicated alignment strategies.

cs.CL

PathCal: State-Aware Reflection-Marker Calibration for Efficient Reasoning

The emergence of Large Reasoning Language Models (LRMs) has paved the way for tackling complex reasoning tasks through test-time scaling by generating long-form Chain-of-Thought (CoT) trajectories during inference. Meanwhile, these trajectories often contain explicit reflection markers such as ``wait'', ``but'', and ``alternatively'', signaling hesitation, revision, and the consideration of alternative explorations, respectively. Recent studies on test-time control leverage such markers as lightweight handles for steering reasoning, typically treating them as a single coarse-grained category rather than distinguishing their distinct functional roles. In this paper, we conduct type-wise suppression and fixed-prefix intervention, revealing that reflection markers differ not only in their functional roles but also in when they exert the greatest influence. Specifically, different marker classes affect accuracy and generation length in distinct ways, and marker choices are most consequential before the model settles into a stable reasoning trajectory. Motivated by these findings, we introduce PathCal, a novel training-free decoding controller that calibrates reasoning paths by distinguishing marker types and intervening only at locally uncertain states. At each decoding step, PathCal utilizes the distribution over reflection-markers to estimate local competition between maintaining the current reasoning trajectory and initiating a competing branch, and softly rebalances marker logits when competing-branch evidence becomes excessive. Experiments across six reasoning benchmarks demonstrate that PathCal achieves a better efficiency--performance trade-off, improving or preserving accuracy while reducing generation length, without relying on external verifiers or additional sampling.

cs.AI