Search arXiv⌕ Search

arXiv subjects

Zhuorui Zhang

Publications and source records attributed to Zhuorui Zhang.

6 recordsLinked to original sources

Match4Annotate: Cross-Video Annotation Transfer in Ultrasound via Implicit Feature Flow-Guided Matching

Acquiring per-frame annotations for ultrasound videos is costly and requires clinical expertise, limiting learning-based analysis. We study cross-video annotation transfer: propagating user-specified annotations from a labeled ultrasound video to an independently acquired target video with no target-side labels or manual initialization. Video trackers and segmentation propagators rely on temporal continuity and require a prompt in every new sequence, whereas cross-image feature matching and one-shot segmentation estimate correspondences independently, without enforcing coherent deformations or supporting both point and mask annotations. We present Match4Annotate, a test-time framework with three stages. A spatiotemporal implicit feature representation lifts frozen vision foundation-model features into a continuous field over space and time, enabling queries beyond the backbone resolution. A continuous implicit feature flow then aligns the source and target fields under a smooth-deformation prior, estimating correspondence in feature space rather than relying on intensity consistency, which is often violated in ultrasound by speckle and acquisition-dependent appearance. Finally, flow-guided annotation transfer uses the estimated flow as a spatial prior over feature similarity. This formulation unifies sparse point and dense mask transfer and includes unconstrained feature matching and direct flow warping as limiting cases. On four clinical ultrasound datasets spanning echocardiography and musculoskeletal imaging, Match4Annotate achieves state-of-the-art annotation transfer, outperforming dense feature-matching baselines across PCK thresholds and one-shot segmentation methods in Dice score. It also demonstrates bidirectional transfer of left-ventricular annotations across datasets. It requires no task-specific training and adapts to each video in minutes on a single consumer GPU.

cs.CV↗

Cohort-Scale Neural Atlases of Ultrasound Video

Ultrasound is the most widely used real-time imaging modality in clinical practice, yet per-frame video annotation remains a major bottleneck: expert labels are scarce and costly, and image appearance varies with speckle, shadowing, attenuation, and operator-dependent probe pose. This is especially limiting because clinically relevant information is often dynamic, from left-ventricular motion in echocardiography to muscle and bone kinematics in musculoskeletal imaging. Population atlases can amortize annotation cost by registering observations to a shared canonical coordinate system, but existing neural atlas methods mainly target single videos, small test-time image sets, or object-centric image collections. We introduce a cohort-scale neural atlas for ultrasound video: a single canonical chart with per-video Generative Latent Optimization embeddings, trained jointly over thousands of frames in DINOv3 feature space. Across five cardiac and musculoskeletal datasets with point landmarks and segmentation masks, our method learns coherent canonical templates and enables accurate atlas-space annotation transfer. On EchoNet-Dynamic and MSK-Bone, it supports single- and few-shot transfer with accuracy competitive with strong dense-correspondence baselines, while training in minutes on a single consumer GPU. The learned embeddings are interpretable: linear projections reveal structured cohort variation, image-decoder interpolation produces anatomically plausible intermediate frames, and test-time latent inversion reconstructs held-out frames through the atlas. These results suggest that cohort-scale neural atlases offer a practical, interpretable representation for reducing expert annotation burden in ultrasound video analysis.

cs.CV↗

Emergent Decoherence Dynamics in Doubly Disordered Spin Networks

Elucidating the emergence of irreversible macroscopic laws from reversible quantum many-body dynamics is a question of broad importance across all quantum science. Many-body decoherence plays a key role in this transition, yet connecting microscopic dynamics to emergent macroscopic behavior remains challenging. Here, in a doubly disordered electron-nuclear spin network, we uncover an emergent decoherence law for nuclear polarization, $e^{-\sqrt{R_{p}t}}e^{-R_{d}t}$, that is robust across broad parameter regimes. We trace its microscopic origins to two interdependent decoherence channels: long-range interactions mediated by the electron network and spin transport within the nuclear network exhibiting anomalous, sub-diffusive dynamics. We demonstrate the capacity to control--and even eliminate--either channel individually through a combination of Floquet engineering and (optical) environment modulation. We find that disorder, typically viewed as detrimental, here proves protective, generating isolated electron-free clusters that localize polarization and prolong coherence lifetimes. These findings establish a microscopic framework for manipulating decoherence pathways and suggests engineered disorder as a new design principle for realizing long-lived quantum memories and sensors.

quant-ph↗

Dynamic Reconstruction of Hand-Object Interaction with Distributed Force-aware Contact Representation

We present ViTaM-D, a novel visual-tactile framework for reconstructing dynamic hand-object interaction with distributed tactile sensing to enhance contact modeling. Existing methods, relying solely on visual inputs, often fail to capture occluded interactions and object deformation. To address this, we introduce DF-Field, a distributed force-aware contact representation leveraging kinetic and potential energy in hand-object interactions. ViTaM-D first reconstructs interactions using a visual network with contact constraint, then refines contact details through force-aware optimization, improving object deformation modeling. To evaluate deformable object reconstruction, we introduce the HOT dataset, featuring 600 hand-object interaction sequences in a high-precision simulation environment. Experiments on DexYCB and HOT datasets show that ViTaM-D outperforms state-of-the-art methods in reconstruction accuracy for both rigid and deformable objects. DF-Field also proves more effective in refining hand poses and enhancing contact modeling than previous refinement methods. The code, models, and datasets are available at https://sites.google.com/view/vitam-d/.

cs.CV↗

Cryogenic field-cycling instrument for optical NMR hyperpolarization studies

Optical dynamic nuclear polarization (DNP) offers an attractive approach to enhancing the sensitivity of nuclear magnetic resonance (NMR) spectroscopy. Efficient, optically-generated electron polarization can be leveraged to operate across a broad range of temperatures and magnetic fields, making it particularly appealing for applications requiring high DNP efficiency or spatial resolution. While a large class of systems hold promise for optical DNP, many candidates display both variable electron polarizability and electron and nuclear T1 relaxation times as functions of magnetic field and temperature. This necessitates tools capable of studying DNP under diverse experimental conditions. To address this, we introduce a cryogenic field cycling instrument that facilitates optical DNP studies across a wide range of magnetic fields (10mT to 9.4T) and temperatures (10K to 300K). Continuous cryogen replenishment enables sustained, long-term operation. Additionally, the system supports the ability to manipulate and probe hyperpolarized nuclear spins via pulse sequences involving millions of RF pulses. We describe innovations in the device design and demonstrate its operation on a model system of 13C nuclear spins in diamond polarized through optically pumped nitrogen vacancy (NV) centers. We anticipate the use of the instrument for a broad range of optical DNP systems and studies.

quant-ph↗

Anything in Any Scene: Photorealistic Video Object Insertion

Realistic video simulation has shown significant potential across diverse applications, from virtual reality to film production. This is particularly true for scenarios where capturing videos in real-world settings is either impractical or expensive. Existing approaches in video simulation often fail to accurately model the lighting environment, represent the object geometry, or achieve high levels of photorealism. In this paper, we propose Anything in Any Scene, a novel and generic framework for realistic video simulation that seamlessly inserts any object into an existing dynamic video with a strong emphasis on physical realism. Our proposed general framework encompasses three key processes: 1) integrating a realistic object into a given scene video with proper placement to ensure geometric realism; 2) estimating the sky and environmental lighting distribution and simulating realistic shadows to enhance the light realism; 3) employing a style transfer network that refines the final video output to maximize photorealism. We experimentally demonstrate that Anything in Any Scene framework produces simulated videos of great geometric realism, lighting realism, and photorealism. By significantly mitigating the challenges associated with video data generation, our framework offers an efficient and cost-effective solution for acquiring high-quality videos. Furthermore, its applications extend well beyond video data augmentation, showing promising potential in virtual reality, video editing, and various other video-centric applications. Please check our project website https://anythinginanyscene.github.io for access to our project code and more high-resolution video results.

cs.CV↗