Search arXivSearch

arXiv subjects

Hong Zhang

Publications and source records attributed to Hong Zhang.

At least 19 recordsLinked to original sources

Hapi: A Multivariable Land-Surface Transformer for Medium-Range Hydrological Forecasting at Continental Scale

Accurate flood forecasts several days in advance are essential for flood control, water-resource management, and emergency response. Producing them at high resolution over a continental domain calls for local hydrological detail together with spatial context extending from river basins to synoptic weather systems. We developed Hapi, a U-Net Swin Transformer that uses fine three-dimensional patches and hierarchical shifted-window attention to forecast discharge, surface runoff, snow water equivalent, and soil wetness across the contiguous United States. The model produces 24--72-hour forecasts at $0.05^{\circ}$ resolution, with learned Laplacian task weights adjusting each variable's contribution to training. On 2024 test data using reconstructed weather and land-surface inputs from ERA5-Land, Hapi outperformed an operational physics-based model and a state-of-the-art AI model in flood detection. Independent validation against 3,881 U.S. Geological Survey gauges and a Hurricane Helene case study supported its advantage over the physics-based model in reproducing daily discharge. Controlled experiments showed that learned task weighting strengthens rare-flood detection, which is particularly sensitive to changes in precipitation inputs. Hapi produced a four-variable, 72-hour forecast across the contiguous United States with an average inference time of 0.11 seconds on a single A100 GPU.

cs.AI

ROVER: Robust Loop Closure Verification with Trajectory Prior in Repetitive Environments

Loop closure detection is important for simultaneous localization and mapping (SLAM), which associates current observations with historical keyframes, achieving drift correction and global relocalization. However, a falsely detected loop can be fatal, and this is especially difficult in repetitive environments where appearance-based features fail due to the high similarity. Therefore, verifying a loop closure is a critical step to avoid false-positive detections. Existing works in loop closure verification predominantly focus on learning invariant appearance features, neglecting the prior knowledge of the robot's spatial-temporal motion cue, i.e., trajectory. In this article, we propose ROVER, a loop closure verification method that leverages the historical trajectory as a prior constraint to reject false loops in challenging repetitive environments. For each loop candidate, it is first used to estimate the robot trajectory with pose-graph optimization. This trajectory is then submitted to a scoring scheme that assesses its compliance with the trajectory without the loop, which we refer to as the trajectory prior constraint (TPC), to determine if the loop candidate should be accepted. Benchmark comparisons and real-world experiments demonstrate the effectiveness of the proposed method. Furthermore, we integrate ROVER into state-of-the-art SLAM systems to verify its robustness and efficiency. Our source code and self-collected dataset will be made available online at https://rover-lcv.github.io/ upon publication of this article.

cs.RO

DiffSynth-Music: Audio-Conditioned KV-Cache Adapters for Controllable Music Generation

Text and lyrics specify broad musical characteristics and sung content but offer limited control over musical timing, melody, and reference-based style. We introduce DiffSynth-Music (https://modelscope.cn/models/DiffSynth-Studio/DiffSynth-Music), a framework that adds composable audio conditioning to a music synthesis backbone through layer-wise key-value injection. The three template models, Control, Prosody, and Reference, are initialized from the backbone diffusion transformer and trained with conditional flow matching. They support five control types: beats, vocals, accompaniment, prosody, and reference audio. A shared variational autoencoder maps conditioning waveforms into a common latent space, enabling their attention memories to be combined. With the template timestep fixed at the clean-data endpoint and other inputs held constant, each control cache is computed once and reused throughout sampling. Training pairs are derived from music recordings using beat extraction, source separation, vocal resynthesis, and reference-excerpt selection. Single-control evaluations on Mandarin and English songs demonstrate improved adherence across all five control types and better lyric fidelity under vocal conditioning relative to the backbone. Automatic music-quality and instruction-following scores remain broadly comparable to those of the evaluated base models, with metric-specific trade-offs. We release the three template models to support research and creative applications in controllable music generation.

cs.SD

An Agentic Evaluation Framework for AI-Generated Scientific Code in PETSc

While LLMs have accelerated scientific code generation, comprehensively evaluating generated code remains challenging. Many benchmarks emphasize functional correctness or task completion, which is insufficient for code built on production HPC libraries, where solver selection, API conventions, memory management, parallel awareness, and performance also matter. We introduce PETSCAgent-Bench, a multidimensional benchmark and agent-based framework for assessing whether AI-generated scientific code uses a production HPC library as an expert would. A tool-augmented evaluator compiles, executes, and measures code and combines deterministic checks with LLM-based assessments in a 14-evaluator pipeline spanning five categories: correctness, performance, code quality, algorithmic appropriateness, and library-specific conventions. A2A and MCP enable black-box evaluation of compatible coding agents. Across realistic PETSc problems, frontier models generate readable, well-structured code but struggle with correctness on challenging problems and with library-specific conventions even when code compiles and runs---limitations that conventional pass/fail evaluation does not capture.

cs.AI

Renewal process's guide to fractional Navier-Stokes equations

The Navier-Stokes equations, which remain unsolved, are crucial equations in fluid mechanics. Discovering the solutions to the Navier-Stokes equations is one of the challenging Millennium problems. In 1900, Hilbert proposed a potential approach to tackle this problem by establishing the relationship between microscopic dynamics and the macroscopic continuum equations. The key bridge is the derivation of Boltzmann equation and the theory of probability. In this paper, we shall use the collision renewal process with arbitrarily distributed waiting times to derive the Boltzmann equation for the time evolution of the probability of the velocity and the displacement of the particle, based on which we prove that the renewal process with exponential collision waiting time is equivalent to the classical Navier-Stokes equations, and that with power-law waiting time is equivalent to the fractional Navier-Stokes equations. Since the collision renewal process with arbitrarily distributed waiting times is a random process and is easy to perform the stochastic simulations of trajectories to obtain the corresponding solution, we actually find a stochastic approach to solve the classical and fractional Navier-Stokes equations.

cond-mat.stat-mech

Intermittent continuous-time random walks under renewal reset mechanism

Stochastic resetting as a practical and efficient search strategy in complex and disordered environments has long been a topic of interest to researchers. Based on the competition between jumping and resetting, this article proposes and investigates intermittent continuous-time random walks (CTRWs) under stochastic resetting, using the smaller waiting time for jump and reset as the renewal time, where the waiting times for both jump and reset can have arbitrary distributions. After each renewal event, the system will proceed with new waiting times for jump and reset regardless of their previous histories. We study the governing equation and Montroll-Weiss equation with renewal resetting, as well as the Markovian resetting for intermittent CTRWs. We prove the existence of non-equilibrium stationary states within the renewal reset mechanism when the jump and reset waiting times follow any exponential and power law distributions. For exponential and Gaussian distributed jump lengths, we examine the mean square displacements (MSDs) of particles to determine their monotonicity and asymptotic stability. Moreover, we calculate the first-arrival time to quantify search efficiency, and validate the intermittent CTRWs under renewal resetting lead to a finite mean first-arrival time (MFAT) to any fixed position for exponential jump and reset waiting time distributions (WTDs), power-law jump and exponential reset WTDs, as well as exponential jump and power-law reset WTDs. However, the MFAT diverges for power-law jump and reset WTDs. The intermittent CTRW model, which is based on the competition mechanism, can be applied to many physical scenarios, such as the foraging strategy of animals that return to their nests after an unsuccessful foraging attempt, or the work planning of intelligent robots that return to energy replenishment points after prolonged operation.

cond-mat.stat-mech

The Dynamical Instability of Rotating Boson Stars

We investigate the dynamical instability of rotating boson stars described by the Gross--Pitaevskii--Poisson equations with contact self-interactions. Through three-dimensional simulations, we confirm that the rotating boson star undergoes a quasiperiodic conversion between ring-like and twin-star-like density configurations in the early nonlinear stage. We develop a systematic linear stability analysis to identify the modes driving this instability and show that repulsive self-interactions could significantly increase the lifetime of the rotating boson stars. We further construct a three-mode Hamiltonian to describe the early nonlinear stage, which explains the quasiperiodic conversion. This analytical framework agrees reasonably well with the simulation results and provides a clear picture to understand the dynamics of rotating boson stars.

gr-qc

Retrospective Causal Attribution under Case-Control Sampling

The probability of necessity PN quantifies the probability that an exposed individual who experienced an outcome would not have experienced it in the absence of exposure. Case-control studies are an important resource for investigating etiologic questions, but their sampling design can introduce selection bias and complicate the statistical inference for PN.In this paper, we develop a nonparametric framework for identification and efficient estimation of PN under case-control sampling. With an externally supplied population outcome prevalence, we derive an exact identification formula that identifies PN under standard causal assumptions and monotonicity and yields a valid lower bound without monotonicity. For rare outcomes, we derive a more tractable approximation that requires no external prevalence information and prove that its approximation error vanishes at the order of the population outcome prevalence. We further establish the semiparametric efficiency theory for the exact and approximate functionals, propose asymptotically efficient estimators, and construct confidence intervals for the corresponding targets. The proposed approach has potential applications in biomedical and epidemiological studies where causal attribution is investigated using retrospective data.

stat.ME

Traj-VLN: Learning Pixel-Space Interaction via Autoregressive Trajectory Generation

Benefiting from the powerful priors embedded in large-scale pre-training data and the emerging commonsense reasoning ability, large language models (LLMs) have shown unprecedented generalization capabilities in many research fields. Recently, projecting visual embeddings into the language space via vision-language models (VLMs) to achieve sim-toreal and cross-scene generalization has become a prevailing paradigm in the field of Vision-and-Language Navigation in Continuous Environments (VLN-CE). VLN requires an embodied agent to navigate through unseen environments following natural linguistic instructions. We emphasize that a VLN task can be decomposed into a sequence of sub-tasks, each corresponding to a process of 3D spatial interaction with the environments described by instructions such as "walk to the end of the sofa and turn left." However, such spatial interactions involving moving into the image along the direction of depth sensing are puzzling for VLMs as they were predominantly trained on conversations with RGB images. Rather than incorporating depth or 3D geometric information-which VLMs rarely encounter during pretrainingwe propose an alternative approach: fine-tuning VLMs to learn navigation interactions directly in 2D pixel space through autoregressive trajectory generation. Given a linguistic instruction and historical observations, our model sequentially predicts a series of pixel coordinates, drawing a trajectory from the bottom center of the current observation. While prior work has proved that pixel-goal supervision outperforms learning of discrete actions, our experiments further verify that the supervision of pixel-space trajectory significantly enhances VLN performance. Moreover, we demonstrate that our flagship model achieves state-of-the-art level performance with relatively limited computational resources and training data.

cs.CV

Breakdown of Aharonov-Bohm cage in Rydberg synthetic lattices: the roles of inhomogeneity and long-range exchange

While the interaction-induced breakdown of Aharonov-Bohm (AB) cage is typically attributed to uniform bound-pair transport, systems with inhomogeneous exchange interactions realized with Rydberg synthetic lattices exhibit more complex dynamics. Employing the evolution-path symmetry (EPS) framework developed recently, we analyze the two-particle dynamics via path interference in Fock space. We find that a homogeneous nearest-neighbor exchange interaction cannot break the AB cage, regardless of whether the long-range exchange interaction is present or not. In contrast, we demonstrate that inhomogeneous nearest-neighbor exchange interaction breaks the destructive-interference EPS, and lifts the degeneracy of many-body compact localized states, thereby generating non-local dispersive eigenstates. Consequently, the initial state gains a non-zero overlap with these dispersive states, enabling delocalized transport. Furthermore, while long-range exchange interaction alone preserves the AB cage, its coupling with nearest-neighbor inhomogeneous exchange interaction opens non-canceling pathways that alter the diffusion profile. Our work connects microscopic path interference with macroscopic spectral reorganization, offering an analytical understanding of the mechanism underlying exchange-interaction-induced transport in Rydberg synthetic lattices.

cond-mat.quant-gas

VLAConf: Calibrated Task-Success Confidence for Vision-Language-Action Models

Task-success confidence estimation for Vision-Language-Action (VLA) models provides a crucial task-level signal for monitoring manipulation in open-world environments and supporting downstream decision-making. Existing methods typically construct task-success confidence from action-token probabilities. However, such probabilities are not naturally available in flow-matching policies, limiting their applicability to mainstream flow-matching VLAs. To address this issue, we propose VLAConf, a two-stage representation-level confidence framework that operates on frozen pretrained VLA representations. A step-conditioned Coin-Flip Network learns an uncalibrated inverse success-support score from successful demonstrations, while a low-capacity calibrator fitted on outcome-labeled successful and failed rollouts maps the aggregated score to task-success probability. Experimental results on the LIBERO benchmark demonstrate that VLAConf improves online task-success confidence estimation over alternative approaches. We further demonstrate its utility in selective expert assistance, where confidence-triggered handoffs improve task success over no intervention. Its applicability is also evaluated in real-robot experiments. To access the source code and supplementary videos, visit https://sites.google.com/view/vlaconf.

cs.RO

A fully discrete LBRFD-IPDG method for linear fourth-order parabolic equations

We propose a fully discrete method for linear fourth-order parabolic equations with Dirichlet boundary conditions, combining an implicit LBRFD multistep scheme in time with a mixed interior penalty discontinuous Galerkin (IPDG) method in space. The temporal discretization employs equispaced linear barycentric rational interpolants and incorporates a startup procedure. To facilitate the spatial discretization, the original problem is reformulated through an auxiliary variable. For certain parameter pairs $(n,d)$, the LBRFD method is shown to be $A(α)$-stable and to possess a wider stability angle than the corresponding BDF$p$ method of the same order. Stability and a priori error estimates are established via a $G$-energy technique and the discrete Grönwall lemma. The theoretical analysis yields a total $L^2$ error estimate of order $h^{k-1}+τ^p$, where $p=d$ if $n-d$ is even and $p=d+1$ if $n-d$ is odd. The reduced spatial convergence rate is attributed to boundary contributions on $\partialΩ$. Despite this theoretical prediction, numerical experiments confirm the stability and demonstrate optimal convergence of order $h^{k+1}+τ^p$.

math.NA

Matching Supervision to the Student's Learning Capacity: A Unified Framework for On-Policy Self-Distillation

On-policy self-distillation (OPSD) improves the reasoning abilities of LLMs by internalizing privileged context into model parameters through self-distillation. Two recent research lines promote vanilla OPSD by choosing which tokens to learn from and by controlling how much privileged information the teacher receives, respectively. However, we show that each line optimizes one variable while holding the other fixed, which leads to a suboptimal solution. We argue that the two variables are coupled through the student's learning capacity: the privileged information sets the per-token divergence the teacher prescribes, while token weighting selects which of these the student must absorb. We formalize the two lines of work into a unified optimization framework, which maximizes the aggregate teacher--student divergence, subject to a budget on the aggregate learning difficulty the student can absorb. Under this modelling, we propose Unified On-Policy Self-Distillation (USD), a lightweight online algorithm to solve the Lagrangian. USD reveals that a single dual variable governs both decisions: at one price for learning difficulty, it simultaneously sets the token-selection threshold and the direction of privileged-information adjustment, keeping supervision matched to the student's evolving capacity. Through extensive experiments, USD consistently demonstrates superior performance over OPSD and token- and PI-side baselines across various model scales on various reasoning benchmarks. Code is available at https://github.com/lauvlalala/USD.

cs.AI

HAM-VLN: Harnessing Hierarchical Agentic Memory for Zero-Shot Vision-and-Language Navigation

Vision-and-language navigation (VLN) enables robots to follow instructions in previously unseen environments. Recently, a training-free paradigm has emerged: the robot queries a multimodal LLM to understand its observations and plan the next action. However, long-horizon navigation based on either image streams or dense map inevitably introduces a growing memory and reasoning bottleneck. We present HAM-VLN, a decision-coupled, agent-authored memory that equips the robot with a persistent, depth-grounded world graph. In the same model call used to select the next action, HAM-VLN also records semantic and reflective information---including room type, objects, navigation progress, and failure notes. Recent waypoints remain verbatim within a bounded window, while older history re-enters the context only through retrieval scored by relevance, recency, and salience, together with one-hop topological expansion. This design requires no additional LLM calls beyond the per-waypoint decision. Compared to previous methods, HAM-VLN not only improves various navigation metrics but also reduces the context length by more than 65%. Specifically, HAM-VLN achieves 61.0% Success Rate (SR) on VLN-CE R2R, 52.7% SR on VLN-CE RxR, and 79.7% SR on HM3D-v2 ObjectNav without any training.

cs.RO

A Portable and Versatile Limited-Memory BFGS Implementation in PETSc/TAO

The limited-memory BFGS (L-BFGS) Hessian update scheme is the critical kernel in many quasi-Newton optimization algorithms. The most common approach to implementing L-BFGS uses $2m$ sequential rank-1 updates as part of solving a linear system when there are $m$ history steps. The performance of this approach suffers when the latency of synchronization is significant, and its poor temporal locality increases the memory traffic when vectors do not fit in cache. The compact dense representation of L-BFGS results in an approach that has minimal synchronization latency and better temporal locality, but it requires an additional pass over the basis vectors and an additional basis that must be recomputed when the $B_0$ matrix changes as in variable-metric methods. In the Portable Extensible Toolkit for Scientific Computation and the Toolkit for Advanced Optimization (PETSc/TAO), we have implemented an intermediate dense formulation of BFGS that retains most of the good characteristics of both the recursive and compact dense approaches. We report single-node performance tests of these implementations on the U.S. Department of Energy's Polaris and Frontier machines, testing both GPU-based and CPU-based computations.

cs.DC

AnchorMark: Robust Diffusion Watermarking via Latent-Space Rotation Synchrony

Inversion-based watermarking embeds watermark payloads directly into the generative process, avoiding a separate post-hoc image-domain embedding stage while preserving the native visual fidelity of synthesized images. However, existing methods remain vulnerable to compound lossy post-processing, particularly when rotation is involved, as it disrupts the spatial correspondence required for latent-space decoding. To overcome this limitation, we introduce AnchorMark, a training-free, robust inversion-based watermarking. We uncover a latent-space property termed Rotation Synchrony: image-domain rotations and their counterparts in the recovered initial latent share the same angle. Building on this property, AnchorMark embeds a synchronization anchor in the central region of the initial latent, enabling accurate estimation and correction of the rotation angle during extraction. Experiments show that AnchorMark substantially improves bit accuracy under rotation and combined attacks, with limited impact on image quality.

cs.CR

BlindPSNR: A No-Reference Fidelity Predictor for Low-Light Image Enhancement

Low-light image enhancement (LLIE) methods involve tunable parameters that are typically fixed, often leading to performance degradation when applied across scenes. Manually selecting the best configuration, however, can be time-consuming and not always practical. Peak signal-to-noise ratio (PSNR) is the natural fidelity criterion for automating parameter selection, yet it requires a ground-truth reference that is typically unavailable. To our knowledge, no learning-based method addresses no-reference PSNR prediction for low-light image enhancement; the natural surrogate, no-reference image quality assessment (NR-IQA), targets perceptual quality rather than signal fidelity, and all seven baselines we test achieve 0% top-1 selection accuracy on our benchmark. With paired training data, the ground-truth PSNR is analytically computable, providing exact supervision without a separate teacher network. Building on this, we propose BlindPSNR, a lightweight no-reference network that fuses the enhanced image with the degraded low-light input via windowed cross-attention and estimates PSNR through heteroscedastic regression. While a scalar-regression baseline achieves top-1 accuracy of 54.4%, BlindPSNR raises this to 89.5% with regret dropping from 1.62 dB to 0.026 dB, and generalizes to unseen datasets (SRCC = 0.61-0.67).

cs.CV

Enhancing Glass Surface Reconstruction via Depth Prior for Robot Navigation

Indoor robot navigation is often compromised by glass surfaces, which severely corrupt depth sensor measurements. While foundation models like Depth Anything 3 provide excellent geometric priors, they lack an absolute metric scale. We propose a training-free framework that leverages depth foundation models as a structural prior, employing a robust local RANSAC-based alignment to fuse it with raw sensor depth. This naturally avoids contamination from erroneous glass measurements and recovers an accurate metric scale. Furthermore, we introduce \ti{GlassRecon}, a novel RGB-D dataset with geometrically derived ground truth for glass regions. Extensive experiments demonstrate that our approach consistently outperforms state-of-the-art baselines, especially under severe sensor depth corruption. The dataset and related code will be released at https://github.com/jarvisyjw/GlassRecon.

cs.RO