Search arXivSearch

arXiv subjects

Jie Gu

Publications and source records attributed to Jie Gu.

At least 19 recordsLinked to original sources

Collective advantage from a minimal record in a quantum information engine

For a single particle, a quantum information engine can turn measurement fluctuations into transport by raising a barrier behind the particle each time its position is measured. When many particles are present, the barrier still depends on just one number, the position of the leftmost particle, so we let the demon measure that order statistic and nothing else. A demon that resolves every position instead pays a record entropy that grows with particle number, even though the extra distinctions never change where the barrier is placed. We show that the coarser measurement supplies an energy that is bounded independently of particle number, and that this bound yields a ceiling on the record entropy that falls as the filling increases. At the same tilt, the collective engine then delivers $44\%$ more work per recorded nat than independent single-particle engines, each with its own optimized cycle time and each running at equal or greater power per particle. Pauli blocking ends this advantage once the accessible region just above the wall fills. We also evaluate the coherence that the coarse measurement leaves behind, and the cost of placing the barrier imprecisely.

quant-ph

Modularity of resurgent topological string

Topological string free energy has a rich collection of non-perturbative contributions which are labeled by D-brane charge vectors, and the associated Stokes constants are conjectured to coincide with BPS or DT invariants, i.e. D-brane multiplicities. In this paper, we provide additional evidence to this conjecture by studying modular properties of non-perturbative contributions. We argue using resurgence theory that non-perturbative contributions form orbits of local monodromy group induced by singular points inside a stability chamber, and that the associated Stokes constants must be the same across the orbits. In some examples, this allows generation of infinitely many Stokes constants, which reproduce the entire BPS spectrum. In addition, following [DK26], we also show that generators of Stokes transformations of non-holomorphic partition function satisfy Lie brackets of the Kontsevich-Soibelman Lie algebra, making it possible to identify the global Stokes transformation with the Kontsevich-Soibelman wall-crossing invariant.

hep-th

DeformSmith: Physics Harness-Guided Hierarchical Generation of Deformable Assets for Robot Manipulation

Creating deformable assets for robot manipulation requires jointly specifying their geometry, appearance, and physical properties. This is especially challenging for deformable objects, since text and images provide limited evidence about how they deform and respond to contact, yet these responses directly affect their suitability for interaction. Automated generation therefore needs to resolve coupled physical requirements and use interaction evidence to guide construction and refinement. We present DeformSmith, a framework that enables automated generation of interactive, physically credible deformable assets from text or a single image. Through hierarchical agentic construction and a shared physics-grounded harness, it progressively builds, tests, and refines geometry, physical models, material behavior, and robot interaction until the resulting asset is ready for simulation and manipulation. Robot interaction closes the generation loop through manipulation feedback and replayable interaction data. Results show that DeformSmith generates assets with better visual quality and physical plausibility than state-of-the-art baselines, including PhysGen3D, PhysGM, and PhysX-Omni, while supporting the synthesis of data for robotic manipulation of deformable objects. Project page: https://can-lee.github.io/deformsmith-web/

cs.RO

Rare-History Transitions in Temporally Random Integrable Quantum Circuits

We study current fluctuations in a temporally random integrable quantum circuit. Commutativity reduces every drive history exactly to its layer composition, turning annealed fluctuations into a competition between current gain and the large-deviation cost of rare compositions. For each fixed composition, homogeneous thermodynamic Bethe ansatz dressing supplemented by ballistic fluctuation theory yields the conditional current statistics. Their annealed large-deviation contraction predicts a first-order switch between two dominant history classes, terminating on a line of regular cusp endpoints. Finite-time analysis shows how the switch is rounded. Thus temporal randomness can act as an emergent order-parameter-like coordinate in trajectory space.

cond-mat.stat-mech

SANTS: A State-Adaptive Scheduler for World Action Models

World Action Models (WAMs) improve robot manipulation by using video-based future representations to condition action generation. In pixel-space WAMs, however, the best action condition is not necessarily the fully denoised video. Controlled denoising-depth scans show that video refinement can reduce action error up to a state-dependent point, after which the gain may saturate or even reverse when late predictions become less action-relevant or physically unreliable. This suggests that action generation should use a state-dependent point along the video noise trajectory rather than a fixed terminal denoising depth. We introduce State-Adaptive Noise Trajectory Scheduler (SANTS), a lightweight scheduler for video-to-action diffusion policies. At each video decision point, SANTS reads the current video-state representation and noise level, then jointly predicts a cumulative stopping hazard and a relative noise-progression ratio. SANTS is post-trained with a path-level reward computed after the frozen action branch generates the final action chunk, so the scheduler is optimized for downstream action quality rather than intermediate video fidelity, while redundant video-state updates are explicitly penalized. Experiments show that SANTS reaches \(94.4\%\) overall success on RoboTwin 2.0 and \(73.1\%\) average success across seven real-robot tasks, while reducing latency by \(81.7\%\) and \(79.0\%\) relative to full video denoising, respectively. These results indicate that adaptive selection along the video noise trajectory can preserve the control benefits of WAM-style future reasoning while removing much of its redundant inference cost.

cs.RO

Thou shalt not tunnel: Complex instantons and tunneling suppression in deformed quantum mechanics

The quantization of the Seiberg-Witten curve of ${\cal N}=2$ super Yang-Mills theory leads to a deformation of one-dimensional quantum mechanics with unconventional behavior. Most notably, quantum tunneling is suppressed at special points in parameter space. In this paper we examine these deformed models in the case of double-well and cubic potentials, and we find that they have a rich phase structure. In what we call the strong coupling phase, the theory behaves like conventional quantum mechanics, instantons are real, and tunneling is not suppressed. In the weak coupling phase, the instantons responsible for tunneling become complex, and tunneling suppression takes place at the so-called Toda lattice points. At the critical point between the two phases, which corresponds to a monopole point in super Yang-Mills theory, the non-perturbative amplitudes display an anomalous scaling as a function of $\hbar$. This phase structure reflects the physics of the underlying super Yang-Mills theory and can be regarded as a physical manifestation of wall-crossing behavior of the BPS spectrum, which we determine in our problem by using resurgent techniques.

hep-th

Reconfiguration-Complete Motion Primitives with Constructive Planning for Deformable Planar Modular Robots

The continuously deformable geometry of modular robots makes it difficult to define a fixed representation for reconfiguration planning and analysis. This letter introduces a square-cell abstraction that maps deformable rhombus modules to fixed-size grid cells while retaining physically interpretable local motions through two primitives, pivoting and shearing. Under this abstraction, we prove that every non-straight edge-connected configuration with $N \geq 7$ can be transformed to a fixed canonical staircase using only admissible primitive motions. Since these motions are reversible, any two configurations in this class are mutually reconfigurable. The proof is constructive and directly yields a staircase-canonicalization planner that transports removable boundary modules while preserving connectivity. As a practical enhancement, we further introduce a boundary-to-delivery lookahead selector that ranks admissible high level choices without affecting the completeness guarantee. Experiments demonstrate the constructive reconfiguration process and show that the selector substantially reduces planning time, while reference comparisons indicate lower planning times than the prior framework over the shared module counts.

cs.RO

Hamiltonian Thresholds for Objective Records

Classical objectivity requires an environment to witness a pointer value, not merely to decohere it. We derive microscopic Hamiltonian laws for this distinction. Exact QND product monitoring yields a certified two-edge strong-Darwinism window: sufficient witnesses cross the optimal record-discrimination edge while retaining an unobserved complement that crosses the residual-coherence edge. Under additional regularity assumptions, the supplement gives converse bounds at the same logarithmic scale. For controlled-phase monitors, environmental asymmetry gives a single-fragment information bound, a redundancy converse, and, in efficient i.i.d. streams, an achievable block-capacity law for redundant witnesses. A finite-temperature spin-phase bath shows the physical separation: hot probes can decohere efficiently while losing the local asymmetry needed for objective witnesses.

quant-ph

Move-Then-Operate: Behavioral Phasing for Human-Like Robotic Manipulation

We present Move-Then-Operate, a Vision language action framework that explicitly decouples robotic manipulation into two distinct behavioral phases: coarse relocation (move) and contact-critical interaction (operate). Unlike monolithic policies that conflate these heterogeneous regimes, our architecture employs a dual-expert policy routed by a learnable phase selector, introducing a structural inductive bias that isolates phase-specific dynamics. Phase labels are automatically generated via an MLLM-based pipeline conditioned on lightweight contextual cues such as end-effector velocity and subtask decomposition to ensure alignment with human motor patterns. Evaluated on the RoboTwin2 benchmark, our method achieves an average success rate of $68.9\%$, outperforming the monolithic $π_0$ baseline by $24\%$. It matches or exceeds models trained on $10\times$ more data and reaches peak performance in $40\%$ fewer training steps, demonstrating that architectural disentanglement of move and operate phases is a highly effective and efficient strategy for mastering high-precision manipulation.

cs.RO

ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams

Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agents equipped with long-term memory over video streams have attracted increasing interest. Unfortunately, existing systems either keep their memory inside the model context or in a flat feature store, and organize it around frames rather than around the persistent entities a stream is really about, which confines them to bounded videos and weakens their ability to track who and what reappears over time. In this paper, we propose ReflectWorld-MM, an entity-oriented multimodal memory system for open-ended video streams. It consists of three parts. The first is a perception front-end that turns an audiovisual stream into entity-resolved observations under a bounded short-term memory. The second is a hierarchical long-term memory, grounded in human memory theory, that couples a multi-scale episodic memory, an evolving entity-centric semantic memory, and a procedural memory. The third is a complete realization, built for real-world operation, that ingests arbitrary streams and plugs into off-the-shelf assistants. Across six long-video and lifelong-memory benchmarks, ReflectWorld-MM achieves the best accuracy on all six, outperforming strong memory agents and a frontier model.

cs.CV

Response kinetic uncertainty relation for Markovian open quantum systems

Response uncertainty relations in stochastic thermodynamics extend precision bounds to the sensitivity of observables under external perturbations. Here we derive a quantum response kinetic uncertainty relation for continuously monitored Markovian open quantum systems in the steady state of the Lindblad master equation. The response precision of a measured trajectory observable is bounded by two contributions: the conventional quantum dynamical activity and a perturbation-induced intersubspace transition term. The latter is absent in the classical limit and captures a genuinely quantum part of the response cost. We identify simple conditions under which either contribution vanishes, and we further clarify the structure of the intersubspace term through a symmetry-resolved decomposition and exact sector-selection rules. The bound and its structure are illustrated in a driven two-level atom.

quant-ph

Dissipation-coherence tradeoff for stochastic oscillations

Autonomous noisy oscillations in biochemical and mesoscopic systems require nonequilibrium driving and therefore dissipation. A striking conjecture by Oberreiter, Barato, and Seifert (OBS) proposes a universal lower bound on the entropy produced per oscillation period in terms of the coherence number of the slowest oscillatory mode. Here we derive a weaker but rigorous lower bound that preserves the OBS structure while introducing a mode-uniformity factor that quantifies how evenly the oscillatory eigenmode is distributed across states in the steady-state inner product. The result makes explicit that an eigenvalue-only prefactor can fail when the dominant oscillatory mode is localized. We also outline a proof-of-principle route for estimating this factor from low-dimensional data under single-mode dominance and sufficiently informative measurements, and derive an eigenvector-free corollary using only the smallest stationary probability. Translation-invariant Markov jump processes on a ring provide a symmetry-protected class with $η=1$, so the refinement reduces to the OBS form; the drift--diffusion limit on a circle saturates the bound.

cond-mat.stat-mech

Restoring Velocity Immunity via Dynamic Mirror Compensation in a Large-Area Dual-Atom-Interferometer Gyroscope

We propose and demonstrate a dynamical mirror compensation scheme to restore velocity immunity in a large-area dual-atom-interferometer gyroscope. In an ideal Mach-Zehnder configuration, the phase shift is inherently immune to atomic velocity, but this property is broken by the Earth's rotation via the Coriolis effect. We overcome this by actively rotating the Raman mirrors during the pulse sequence to cancel the time-dependent angular offset. The implementation relies on a decouplable calibration-compensation chain to remove rotation-induced time-dependent terms. The scheme is validated on a dual-atom-interferometer gyroscope with an interference area of 21.1 cm^2. After compensation, the phase's dependence on atomic velocity is reduced 40-fold, and the velocity contribution to scale-factor stability is evaluated to be 0.13 ppm. The sensor achieves a rotation sensitivity of 1.3\times10^{-8} rad/s/Hz^{1/2} and a stability of 1.9\times10^{-10} rad/s at 4500 s integration, together with a common-mode noise rejection ratio of up to 459, demonstrated in a seismic event. This work removes a key obstacle to scale-factor stabilization in atom-interferometer gyroscopes and paves the way for their applications in inertial navigation and geophysics.

physics.atom-ph

DeformMaster: An Interactive Physics-Neural World Model for Deformable Objects from Videos

World models for deformable objects should recover not only geometry and appearance, but also underlying physical dynamics, interaction grounding, and material behavior. Learning such a model from real videos is challenging because deformable linear, planar, and volumetric objects evolve under high-dimensional deformation, noisy interactions, and complex material response. The model must therefore infer a physical state from visual observations, roll it forward under new interactions, and render the resulting dynamics with high visual fidelity. We present DeformMaster, a video-derived interactive physics-neural world model that turns real interaction videos into an online interactive model of deformable objects within a unified dynamics-and-appearance framework. DeformMaster preserves structured physical rollout while using a neural residual to compensate for unmodeled effects, grounds sparse hand motion as distributed compliant actuator for hand-continuum interaction, represents material response with spatially varying constitutive experts, and drives high-fidelity 4D appearance from the predicted physical evolution. Experiments on real-world deformable-object sequences demonstrate DeformMaster's ability to roll out future dynamics and render dynamic appearance, outperforming state-of-the-art baselines while supporting novel action rollout, material-parameter variation, and dynamic novel-view synthesis. Project page: https://can-lee.github.io/deformmaster-web/

cs.CV

Geometric Floquet Condition for Quantum Adiabaticity

Quantum adiabaticity is the evolution of a quantum system that remains close to an instantaneous eigenstate of a time-dependent Hamiltonian. Using Floquet formalism, we derive a rigorous sufficient condition for adiabaticity in closed, finite-dimensional periodically driven systems that is valid for arbitrarily many driving periods. The condition is stroboscopic and geometric, depending only on single-cycle information: the Fubini--Study length of the instantaneous eigenray and a quasienergy-separation measure extracted from the Floquet operator. We also formulate a state-targeted refinement that reduces conservativeness when only one adiabatic branch is relevant. Rather than synthesizing control pulses, the result provides a certification criterion for a given periodic protocol. We illustrate the criterion and contrast it with conventional instantaneous-gap conditions in three representative examples.

quant-ph

Beyond Hagedorn: A Harmonic Approach to $T\bar{T}$-deformation

We apply harmonic analysis to study the $T\bar{T}$-deformed torus partition function. We first express the CFT partition functions in terms of Maass waveforms, including the Eisenstein series and cusp forms. These basis functions turn out to deform in a very simple way under the $T\bar{T}$-deformation. The spectral decomposition provides a numerically stable and efficient method to compute the partition function at finite values of the deformation parameter $λ$, allowing us to clearly resolve the analytic structure of the partition function as a function of $λ$. The resulting deformed partition function exhibits a Hagedorn singularity. Building on harmonic analysis approach, we propose a natural analytic continuation beyond the Hagedorn singularity, which enables us to compute the full partition function for any value of $λ$.

hep-th

Finite-frequency fluctuation-response bounds for open quantum systems

We derive a finite-frequency fluctuation-response inequality for Markovian open quantum systems in an input-output setting. For any downstream measurement of the emitted field, the measured lock-in response-to-noise matrix is bounded by the output-field quantum Fisher information rate. For dissipative amplitude modulation with vacuum inputs, this information rate is further bounded by a frequency-independent signal-channel activity, which reduces for kinetic modulation to the stationary channel fluxes. The result is detector-facing but unraveling-independent: it applies after choosing a measurement record, while the information ceiling is set by the quantum field before any detection scheme or trajectory representation is selected. We formulate the bound for multiple signal channels and real finite-frequency quadratures, and illustrate it with a single-sided cavity, resonance fluorescence, and a truncated Kerr-parametric cat resonator.

quant-ph

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective

Existing Multimodal Large Language Models (MLLMs) process a large number of visual tokens, leading to significant computational costs and inefficiency. Instruction-related visual token compression demonstrates strong task relevance, which aligns well with MLLMs ultimate goal of instruction following. Previous works generally assume that visual tokens achieve better vision-language alignment in the shallow layers of LLMs, which have led to task-related token compression being primarily applied in intermediate LLM layers. In contrast, our study reveals that with proper selection, task-related token compression is feasible at the input stage of LLM with negligible performance loss. This new paradigm significantly reduces task-irrelevant visual tokens and its model-agnostic design enables application without modifying the LLM architecture. Specifically, we suggest that explainability methods for transformer-based architechtures can evaluate the global importance of each visual token with respect to the given instruction, which can effectively guide the task-related token compression for MLLMs. Furthermore, we propose to learn a mapping from the attention map of the first LLM layer to the explanation results, thereby avoiding the need for a full inference pass. Interestingly, this mapping can be learned using a simple and lightweight convolutional network, whose training is efficient and independent of MLLMs. Extensive experiments on 13 image and video benchmarks across three leading MLLMs (Qwen2-VL, LLaVA-OneVision, and VILA1.5) demonstrate the remarkable effectiveness and strong generalization of our approach. Additionally, our new compression paradigm achieves faster inference with reductions in both prefilling time and KV cache memory.

cs.CV