Search arXivSearch

arXiv subjects

Geng Li

Publications and source records attributed to Geng Li.

At least 19 recordsLinked to original sources

Electromagnetic Selection Rules and Coherent Manipulation of Quantum Skyrmion via Surface Acoustic Wave Phonons

Skyrmions are competitive candidates for information-storage units and have great prospect in quantum information processing. We here study coherent control of a quantum skyrmion via surface acoustic wave (SAW) phonons in a piezoelectric SAW cavity. Exact diagonalization of a finite Dzyaloshinskii-Moriya cluster reveals the eigenstates of quantum skyrmion with robust scalar chirality. The underlying lattice-spin symmetry imposes polarization-dependent selection rules for the electromagnetic transitions between eigenstates states. Considering the small mode volume of the SAW phonon easy to reach strong coupling, we derive the interaction Hamiltonian between the quantum skyrmion as a qubit and a single-mode quantized SAW via the electric field induced by piezoelectric effect, and find that the coupling strength at the single-phonon level increases linearly with skyrmion radius. This enables enhanced spin-acoustic coupling for larger skyrmionic textures and for low magnetic dissipation. Our results establish few-spin quantum skyrmions as compact building blocks for hybrid quantum information processing and suggest a route toward more densely integrated on-chip quantum devices.

quant-ph

Energy Transport Structure and Fluctuation Theorem in Nonreciprocal Harmonic Chains

Nonreciprocal interactions fundamentally alter energy transport by breaking the symmetry between forward and backward responses. Here, we uncover an exact transport structure for a one-dimensional harmonic chain with asymmetric nearest-neighbor couplings. Using a Green-function approach, we demonstrate that the direction of heat transport is determined not solely by the temperature bias, but by its competition with an effective directional asymmetry. This competition can reverse the direction of heat flow, enabling cold-to-hot transport. We further show that nonreciprocity introduces an additional power channel associated with the antisymmetric sector of the interaction, leading to a generalized steady-state energy balance involving two reservoir heat currents and a nonreciprocal power current. At the fluctuation level, the heat exchanges with the two reservoirs constitute correlated yet distinct stochastic currents, whose joint scaled cumulant generating function obeys an exact Gallavotti-Cohen symmetry. Together, these results establish a unified energetic and fluctuation framework for nonreciprocal heat transport and demonstrate that directional interactions can serve as an active resource for controlling nonequilibrium energy flows.

cond-mat.stat-mech

Scalar glueball-$s\bar{s}$ mixing in one flavor lattice QCD

We investigate the mixing between the lowest-lying scalar glueball and the $s\overline{s}$ meson in $N_f=1$ lattice quantum chromodynamics (QCD) utilizing an anisotropic $16^3 \times 128$ lattice ensemble at a lattice spacing $a_s\approx 0.148\,\rm{fm}$. By solving a generalized eigenvalue problem (GEVP) for the optimized glueball and $s\overline{s}$ scalar operators in the $J^\text{PC} = 0^{++}$ channel, the masses of the two lowest-lying eigenstates are determined to be $m_1 = 1.290(20)\,\mathrm{GeV}$ and $m_2 = 1.777(49)\,\mathrm{GeV}$. By extracting the couplings of these mass eigenstates to the glueball and $s\bar{s}$ operators, we determine a substantial mixing angle $|θ| \approx 40.7(2.7)^\circ$ and a large mixing energy $x_s = 239(24)$ MeV. These results indicate a strong glueball-$s\bar{s}$ mixing in the scalar sector, providing important non-perturbative inputs for understanding the nature of the experimental isoscalar scalar mesons. The continuum limit of the mixing energy and its quark mass dependence need to be investigated in the future.

hep-lat

GraphMemix: Query-Aware Evidence Forests for Long-Term Multimodal Agent Memory

Organizing long-term memory for multimodal agents remains challenging because existing methods either suffer from expensive question-agnostic offline summaries or naive embedding similarity matching that introduces incomplete and redundant context. To address these issues, we propose GraphMemix, a combinatorial-optimization graph memory framework that models memory organization as query-aware evidence-forest construction. Specifically, our method consists of three key components:(1) candidate graph construction, which expands multi-view seed memories through schema and semantic relations to acquire query-aware original context; (2) evidence utility and activation costs, which decouples direct memory support from anchor-conditioned relation verification to suppress redundant or conflicting information; and (3) forest optimization, which jointly selects a forest-format memory context under a maximum evidence budget and its reliable relational structure. By organizing memory into a query-relevant subgraph, the method avoids substantial lifecycle cost and recovers low-similarity complementary evidence. Experimental results across four long-term multimodal memory benchmarks demonstrate significant improvements with different foundation models and establish a new Pareto frontier between accuracy and lifecycle cost.

cs.AI

R2M-Bench: Evaluating Revisit Memory via Relative Consistency in Interactive Video World Models

High similarity between first-visit and return frames does not necessarily show that a video world model remembered the scene; the intervening rollout may simply have changed very little. This ambiguity makes absolute revisit scores sensitive to rendering stability, repetitive content, and failed motion. We introduce \emph{R2M-Bench} (\textbf{R}elative \textbf{R}evisit \textbf{M}emory Benchmark), a benchmark of observable revisit-selective consistency. For every detected return, R2M-Bench compares the revisit pair with two controls from the same rollout: a gap-matched non-revisit pair that measures generic temporal stability and a short-range pair that estimates short-horizon consistency. These comparisons produce \emph{MemoryGain} (MG), the revisit advantage over the temporal baseline, and the \emph{Normalized Memory Ratio} (NMR), which normalizes this advantage by the short-to-baseline dynamic range. R2M-Bench combines 100 reference scenes with three leave-and-return trajectories to form 300 instances and evaluates appearance fidelity, scene and object identity, local geometry, and persistent state. Across seven action-conditioned video world models, Overall NMR correlates with human consistency judgments at Spearman's $ρ=0.547$ (95\% CI $[0.45,0.63]$). Its within-model correlation magnitude with generated motion is $0.072$, compared with $0.207$ for raw revisit similarity, indicating that relative calibration substantially reduces the slow-motion shortcut. DreamX-World-Memo achieves the highest Overall NMR among the evaluated video models. Together, these results support same-rollout relative calibration as a practical way to distinguish revisit-specific consistency from generic temporal stability.

cs.CV

Scalar and tensor structures in $J/ψJ/ψ$ scattering from lattice QCD

The $^1S_0$ and $^5S_2$ amplitudes of $J/ψ~J/ψ$ scattering are determined up to 6.6~GeV from $N_f=2$ lattice QCD simulations at $m_π\approx 420$~MeV and $250$~MeV. We observe an attractive interaction in the ${}^1S_0$ channel, which permits the existence of a near-threshold scalar structure. This structure may correspond to $X(6200)$, but the possible left-hand-cut effect prevents us from a precise pole determination. In the ${}^5S_2$ channel, while a near-threshold repulsive interaction is observed, a resonance appears at higher energies associated with a Castillejo-Dalitz-Dyson zero $\sqrt{s}_{\rm CDD}\approx 6.45~\rm{GeV}$ of the scattering amplitude. The tensor resonance has the parameters $(m_R,Γ_R)=(6.543(10), 0.548(34))~\rm{GeV}$ for $m_π\approx 420~\rm{MeV}$ and $(6.538(13),0.537(56))~\rm{GeV}$ for $m_π\approx 250~\rm{MeV}$, which are compatible with those of $X(6600)$ (or $X(6400)$) reported by ATLAS and CMS, and also support the $2^{++}$ assignment of the recent CMS study. The difference of the $^1S_0$ and $^5S_2$ $J/ψJ/ψ$ interaction is attributed to the dominance of the quark rearrangement effects. The systematic uncertainties owing to the unphysical pion mass, the finite lattice spacing, and the finite volume should be investigated by future lattice QCD studies with more sophisticated lattice setups.

hep-lat

$η_cη_c$ and $J/ψJ/ψ$ scatterings from lattice QCD

We investigate the $S$-wave $η_cη_c$ and $J/ψJ/ψ$ scattering in the $J^{PC}=(0,2)^{++}$ channels up to a center-of-mass energy of 6.6~GeV. The calculations are carried out at two unphysical pion masses, $m_π\approx 420$~MeV and 250~MeV, in $N_f=2$ lattice QCD. For each $m_π$, we extract the finite-volume energy levels on two lattices with an identical lattice spacing ($a\simeq 0.136~\mathrm{fm}$) but different spatial volumes. Since the coupled-channel effects between the $η_cη_c$ and $J/ψJ/ψ$ channels are found to be negligible, we analyze the corresponding scattering properties using the single-channel L"{u}scher method. We find that the interactions in these dicharmonium systems are dominated by the quark rearrangement effect. In the $0^{++}$ channel, the near-threshold attraction in $J/ψJ/ψ$ and repulsion in $η_cη_c$ can be explained through the Fierz rearrangement. The attractive interaction in the ${}^1S_0$ $J/ψJ/ψ$ channel allows for the existence of a near-threshold scalar structure, which may correspond to the $X(6200)$. However, possible left-hand-cut effects due to light-hadron exchange prevent us from precisely determining the pole positions. In the $2^{++}$ channel, while the ${}^5S_2$ $J/ψJ/ψ$ system exhibits a repulsive interaction near threshold, the scattering amplitude has a Castillejo-Dalitz-Dyson zero at $\sqrt{s}=6.45~\mathrm{GeV}$ and a resonance pole at $\sqrt{s}=\big[6.543(10)-i,0.548(34)/2\big]~\mathrm{GeV}$ for $m_π\approx 420~\mathrm{MeV}$ and $\big[6.538(13)-i,0.537(56)/2\big]~\mathrm{GeV}$ for $m_π\approx 250~\mathrm{MeV}$, where the uncertainties are statistical. This resonance may correspond to the $X(6600)$ (or $X(6400)$) reported by the ATLAS and CMS Collaborations. Our result supports its $2^{++}$ assignment, in agreement with the latest spin-parity determination by CMS.

hep-lat

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

We present \textbf{DreamX-Phi 1.0}, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee faithfulness: a convincing rollout can still move the wrong arm or lose the manipulated object. To ensure the prediction respects each arm's commanded path, we inject per-arm $\mathrm{SE}(3)$ transformations into attention via \textbf{PRoPE-style geometric encoding}, preserving arm identity and rigid-motion structure. Action control alone does not fully constrain scene geometry or the evolution of small manipulated objects. We therefore add a lightweight \textbf{depth branch} for scene-level geometry and use \textbf{SAM3 masks} with a frozen \textbf{V-JEPA teacher} to maintain object consistency throughout grasping. We further distill the multi-step generator into a few-step student via distribution-matching distillation for efficient deployment. At the time of writing, \model{} achieves first place on Track~1 and second place on Track~2 of the WorldArena~2.0 Challenge. Our model and code will be publicly available.

cs.CV

Shortcuts to Parameter Sweeps

Efficient evaluation of stationary parametric sensitivities over broad parameter ranges is important for identifying influential training data, fitting force fields, and predicting material responses, but standard pointwise approaches require repeated relaxation and sampling. Here we introduce Shortcuts to Parameter Sweeps (STPS), an engineered control strategy that uses an auxiliary control to transport the probability density along a prescribed family of instantaneous stationary states during a finite-time parameter sweep. This enables the continuous response curve over the full parameter interval to be estimated from a single controlled sweep using covariance-based response relations. STPS applies to both equilibrium and nonequilibrium steady-state systems, including those with unknown stationary distributions, and can be implemented directly using stationary samples in high-dimensional settings. Numerical tests on single-particle and interacting many-body systems show that STPS yields response curves in close agreement with reference results. These findings establish STPS as an efficient, sample-based framework for continuous sensitivity analysis in stochastic simulations.

cond-mat.stat-mech

Learning Physical Interaction: A Survey of Tactile- and Force-aware Robot Learning

Physically grounded robot intelligence requires robots to perceive, reason about, and regulate their interactions with the physical world. This capability is particularly critical in contact-sensitive manipulation, where successful task execution depends not only on visual perception and motion generation, but also on force regulation and adaptive control. In this context, recent robot learning methods have made substantial progress by integrating force, tactile, vision, language, and proprioceptive sensing into learned manipulation policies. In parallel, many systems adopt multi-phase architectures that combine high-level policies, action-refinement modules, and low-level controllers to bridge semantic task understanding with reactive physical execution. Despite these advances, existing surveys have not explicitly reviewed force- and tactile-aware robot learning from a unified perspective that jointly captures multimodal sensing and multi-phase system design. This survey addresses this gap by proposing TF-ART, a Tactile/Force-Aware Robot learning Taxonomy for multimodal and multi-phase frameworks, which maps individual methods into a unified hierarchical structure. The framework characterizes how recent works organize observation modalities, encode and fuse heterogeneous sensory inputs, generate and refine actions across multiple phases, and connect learned policies to reactive robot-end control. Building on this methodological view, we further examine the task settings and infrastructure requirements of physical interaction, thereby integrating both algorithmic and practical perspectives on force- and tactile-aware robot learning.

cs.RO

Peak-End-Net: A Peak-End Rule Inspired Framework for Generalizable Video Aesthetic Assessment

Video aesthetic assessment (VAA) aims to predict how aesthetically pleasing a video is, yet remains far less explored than other visual assessment tasks. Its progress is hindered not only by the scarcity of large-scale benchmarks, but also by the intrinsic subjectivity of aesthetic judgment, which is shaped by human perception. In this paper, we revisit VAA from a psychological perspective and propose \textit{Peak-End-Net}, a lightweight and interpretable framework inspired by the \textit{peak-end rule}, which suggests that people tend to judge a temporal experience mainly according to its salient moments and the ending. Building on this intuition, we first transfer knowledge from image aesthetic assessment (IAA) to VAA by introducing a pretrained IAA head to produce frame-wise aesthetic priors, which serve as surrogate signals for identifying aesthetically salient moments and guiding \textit{peak-end rule}-based temporal aggregation. To further capture how a video evolves aesthetically over time, we design an aesthetic rhythm encoder that models temporal progression beyond isolated moments. Additionally, we refine the overall assessment through a dynamic gated fusion mechanism to improve robustness under distribution shift. Our method is built on a frozen vision transformer (ViT) and requires only a small number of trainable parameters, making it scalable and parameter-efficient. Extensive experiments on two existing VAA benchmarks, including in-domain evaluation on VADB and cross-domain testing on DIVIDE-3K, demonstrate that our approach achieves state-of-the-art performance, affirming the value of psychologically grounded modeling for VAA. Our code and models are available at https://github.com/AMAP-ML/Peak-End-Net.

cs.CV

Reproducing human biases in route choice using large language models: Toward scalable behavioral modeling

Human choice behavior, including route choice, exhibits systematic behavioral biases that deviate from the assumptions of full rationality. Cumulative prospect theory (CPT) has been widely recognized as an effective framework for characterizing such behavioral patterns. However, its large-scale application, particularly in simulation and agent-based modeling, critically depends on specifying individual-level CPT parameters, which remain a major bottleneck. Conventional approaches typically rely on surveys and controlled experiments to calibrate CPT parameters, yet these methods are difficult to generalize and often fail to capture the full diversity of human decision-making. To address this challenge, this paper investigates whether large language models (LLMs) can reproduce human behavioral biases in choice-making without explicit specification of prospect-theoretic parameters. Using route choice as a representative scenario, we design a behavioral evaluation framework and systematically compare LLM-generated decisions with established human behavioral patterns predicted by CPT. Experimental results demonstrate that LLMs are capable of reproducing non-rational human choice biases and can exhibit decision behaviors consistent with prospect-theoretic effects under uncertainty. These findings suggest that generative AI models may provide a scalable alternative for modeling human decision processes and offer a promising foundation for next-generation large-scale agent-based simulation and AI-driven behavioral research.

cs.AI

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception

While Multimodal Large Language Models (MLLMs) demonstrate impressive general capabilities, they struggle with fine-grained perception in ultra-high-resolution (UHR) images, particularly for tiny objects in cluttered scenes. Existing methods face a dilemma: they either rely on inefficient prior-free scanning, or depend on static prior-driven heuristics that lack posterior correction to rectify initial model biases. To address this, we propose BVS (Bayesian Visual Search), a framework that formulates perception as a global optimization problem over a continuous spatial-scale manifold. Specifically, BVS bridges prior guidance with posterior correction: it utilizes an early-stop attention rollout of MLLM to construct reasoning-aware priors, while employing a scale-aware non-stationary kernel and GP-UCB to dynamically rectify noise and recover missing information in the prior through iterative local observations. We provide theoretical guarantees via sub-linear regret bounds, and extensive experiments demonstrate that BVS significantly outperforms state-of-the-art baselines with a superior trade-off between accuracy and efficiency.

cs.CV

DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive fine-grained perception capabilities. However, existing benchmarks predominantly rely on explicit textual cues or low-resolution inputs, failing to evaluate a model's ability to autonomously perceive implicit visual cues in high-resolution. To bridge this gap, we introduce DiCoBench, a comprehensive, multi-image high-resolution benchmark designed for cross-image fine-grained perception. DiCoBench consists of 765 meticulously curated samples categorized into two progressive tracks: Differential Visual Cues and Commonality Visual Cues, covering 8 distinct perception tasks. By formulating the benchmark as a multiple-choice question task and utilizing high-resolution imagery (approaching 2K), we eliminate evaluation metric bias and pose a substantial challenge to current state-of-the-art MLLMs. Our extensive evaluation of 18 diverse MLLMs reveals a striking performance gap compared to human accuracy (98.3\%), with top-performing models struggling significantly with micro-scale detail capture. We believe DiCoBench will serve as a challenging testbed to drive future research in autonomous, high-resolution multi-image perception.

cs.CV

ADAPT: Analytical Disturbance-Aware Policy Training for Humanoid Locomotion

Humanoids deployed in human-centered environments must handle force-interactive tasks, where external contacts introduce unexpected disturbances that disrupt locomotion accuracy and stability. Existing learning-based approaches rely on broad domain randomization, task-specific force objectives, or learning-based force estimators from motion history, each of which compromises accuracy, task transferability, or out-of-distribution (OOD) robustness. We present Analytical Disturbance-Aware Policy Training (ADAPT), a framework that equips humanoid policies with a physically grounded disturbance observer. The core of ADAPT is an analytical whole-body disturbance observer that estimates residual force/torque online with the accessible robot dynamics, without requiring force/torque sensors. Fed directly into the policy, the estimated disturbances give the humanoid an explicit, physics-derived sense of external force/torque that can generalize across diverse unseen scenes. Experiments on a Unitree G1 humanoid show that ADAPT achieves accurate disturbance prediction and stronger robustness than a proprioception-only baseline under torso perturbations, standing pushes, and asymmetric hand payloads, with improved velocity tracking even on OOD disturbances. Moreover, ADAPT enables penalizing inferred disturbances at lower-body joints to encourage lighter locomotion.

cs.RO

DreamX-World 1.0: A General-Purpose Interactive World Model

DreamX-World 1.0 is a general-purpose interactive text/image-to-video world model for controllable long-horizon generation. It supports camera navigation, revisits to previously observed regions, and promptable events across photorealistic, game-style, and stylized domains. Our data engine combines camera-accurate Unreal Engine rendering, action-rich gameplay recordings, and real-world videos with recovered camera geometry. For camera control, we introduce E-PRoPE, a lightweight variant of projective positional encoding that retains PRoPE's projective camera geometry while applying camera-aware attention to spatially reduced tokens. We convert a bidirectional video generator into a few-step autoregressive world model using causal forcing, DMD-style distillation, and long-rollout training. Training on self-generated long-horizon contexts exposes the model to its own generated history and reduces the style and color drift that accumulates across autoregressive chunks. Memory-Conditioned Scene Persistence retrieves earlier views through camera-geometry-based retrieval, while residual recycling makes the conditioning path less sensitive to imperfect memory latents. Event Instruction Tuning adds composable event control, and reinforcement learning alignment recovers camera control and visual quality after distillation. With mixed-precision DiT execution, residual reuse, 75\%-pruned VAE decoding, and asynchronous pipeline parallelism, DreamX-World 1.0 reaches up to 16\,FPS on eight RTX\,5090 GPUs. On our 5-second basic evaluation, DreamX-World 1.0 achieves a camera-control score of 73.75 and an overall score of 84.76, outperforming HY-WorldPlay 1.5 and LingBot-World in overall score, which achieve 80.79 and 80.45, respectively.

cs.CV

Geometric Bounds on the Finite-Time Performance of Active Machines

Optimizing energy conversion in active matter remains a central challenge in nonequilibrium physics. Here, we develop a unified thermodynamic framework that characterizes the finite-time performance of interacting active machines. We show that cyclic work admits a geometric decomposition into an antisymmetric thermodynamic curvature, governing work extraction, and a symmetric metric, controlling dissipation. Minimal-dissipation protocols follow geodesics in parameter space, while optimal work extraction deviates from them due to a curvature-induced, Lorentz-like effect. This geometric structure directly determines the finite-time scaling of work and dissipation, enabling a mapping onto Onsager-type quasi-linear current--force relations. We show that both the maximal efficiency and the efficiency at maximum power are governed by an asymmetry parameter and a figure of merit, establishing a formal correspondence between active machines and thermoelectric devices with broken time-reversal symmetry. Our results reveal a fundamental geometric origin of energy-conversion performance and provide a general framework for optimizing active machines.

cond-mat.stat-mech

YARD: Y-Architecture Register Decoding for Efficient Hallucination Mitigation in Large Vision-Language Models

Contrastive decoding (CD) seeks to mitigate hallucinations in Large Vision-Language Models (LVLMs) by contrasting the output distributions of a standard model and a visually degraded model. However, existing training-free CD methods suffer from sub-optimal degraded branches: completely dropping visual tokens is too extreme and induces language hallucinations, while corrupting input images offers coarse control over visual evidence and suffers from high inference latency due to requiring two full forward passes. To address these dilemmas, we propose YARD, a training-free Y-Architecture Register Decoding framework. Motivated by the observation that reliable text-to-vision grounding predominantly emerges in the middle decoder layers, YARD constructs the degraded branch internally by sharing shallow-layer computations and branching exactly at this critical stage. For the degraded branch, YARD replaces patch-level visual tokens with register tokens, which preserve global image semantics but lack fine-grained local evidence. This image-aware yet locally under-grounded design provides a faithful contrastive signal without extreme modality mismatch, while the Y-architecture strictly avoids a costly second forward pass. Extensive experiments on generative and discriminative hallucination benchmarks demonstrate that YARD consistently achieves state-of-the-art hallucination mitigation across multiple LVLMs, alongside a significant reduction in inference latency.

cs.CV