Search arXiv⌕ Search

arXiv · 2610.04258

FloVMos: Optical Flow-based Medical Video Mosaicking

Abstract

Biomedical imaging modalities often require a trade-off among resolution, field of view (FOV), and acquisition speed. Video mosaicking offers a strategy to overcome this limitation by computationally stitching sequential high-resolution frames into a wide-FOV composite. However, existing methods struggle with non-rigid deformations, and modality-specific artifacts arising in clinical and research imaging. Here, we present FloVMos, a generalizable, optical-flow-based deep learning framework for real-time video mosaicking across diverse biomedical imaging modalities. FloVMos achieves robust, pixel-level registration by fine-tuning an optical flow model on synthetic training data with ground-truth deformation fields. We introduce a pipeline for generating this training data, simulating realistic tissue motion and imaging distortions from existing mosaics or raw videos. Our automated synthetic data generation and optical flow model training based on this data allow users to adapt FloVMos to different imaging modalities. To demonstrate this, we applied FloVMos to seven diverse imaging modalities: reflection confocal microscopy, open-top light-sheet microscopy, fetoscopy, laparoscopy, dermoscopy, sparse spectral microscopy, and endoscopy. FloVMos outperforms conventional baselines in accuracy, robustness, and speed for all the tested modalities. This adaptable and training-efficient framework enables large-area visualization with real-time performance and may support broader use of video-based biomedical imaging in research and clinical workflows.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jinyang Liu, Sandesh Ghimire, Chaman Singh, Jennifer Dy, Milind Rajadhyaksha, Dana H. Brooks, Octavia Camps, Kivanc Kose. 2026-10-03. FloVMos: Optical Flow-based Medical Video Mosaicking. https://arxiv.org/abs/2610.04258

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Underwater imaging without color distortions requires RAW capture

1. Consumer cameras are not designed to be scientific instruments, yet they are increasingly used in science not just to see, but to measure. This is possible because camera sensors respond approximately linearly to the light they receive, a property that can, in principle, support quantitative measurements of color. 2. Whether that quantitative measurement potential is preserved depends on what happens to the image after capture. Minimally processed sensor images, commonly known as RAW, retain the approximately linear response needed. Instead of RAW, however, the default output of most cameras is a heavily processed image optimized for visual appearance, compact storage, and immediate use, typically saved as a JPEG. 3. The convenience offered by an in-camera processed image comes at a scientific cost. RAW capture requires more storage and processing effort but leaves calibration choices open. Processed images reduce those demands at acquisition and in post-processing but transfer irreversible decisions about color manipulation from the researcher to the camera. 4. The preservation of RAW sensor data is necessary but not sufficient for quantitative color measurement. Pixel values must also be calibrated to account for the light conditions under which the image was acquired. Underwater, this calibration is more challenging than in clear air, because light is absorbed and scattered as it travels through water, resulting in effects that depend on both wavelength and distance. Consequently, color distortions vary across the scene and generally cannot be corrected with a single global color transformation like white balancing; instead, calibration requires information about the imaging distance. We explain the consequences for quantitative imaging and introduce the 3P protocol for distortion-free color imaging underwater.

eess.IV↗

OASIS: Optimized Lightweight Autoencoder System for Distributed In-Sensor computing

In-sensor computing, which integrates computation directly within the sensor, has emerged as a promising paradigm for machine vision applications such as AR/VR and smart home systems. By processing data on-chip before transmission, it alleviates the bandwidth bottleneck caused by high-resolution, high-frame-rate image transmission. We propose a system architecture that integrates a CMOS image sensor (CIS) with a logic chip via advanced packaging, where the logic chip executes a lightweight encoder that aggressively compresses activations before off-chip transmission. However, conventional deep neural network (DNN) partitioning strategies require multiple layers before achieving meaningful dimensionality reduction, limiting bandwidth savings. To address this, we propose a dual-branch autoencoder-based vision architecture trained with a triple loss function combining task-specific, entropy, and reconstruction objectives to produce compact, Huffman-codable representations. The encoder is prototyped on a FPGA to provide hardware-validated energy measurements, complementing circuit-simulated CIS front-end characterization. Our approach achieves a four-order-of-magnitude reduction in transmitted data volume compared to raw input images, resulting in $2{\times}{-}4.5{\times}$ energy savings. We evaluate on CNN and ViT-based models across diverse tasks, achieving state-of-the-art accuracy with energy efficiency of up to 22.7 TOPS/W.

eess.IV↗

One Photon, Many Worlds: Posteriors and Predictions with Single-Photon Cameras

Single-photon avalanche diode (SPAD) cameras operate fundamentally differently from conventional cameras due to their photon-counting nature. Each frame produces a binary image: pixels report zero if no photons arrived during exposure, and one if one or more photons arrived. Reconstructing a scene or inferring its properties from a single binary frame is difficult because many different images could produce the same measurement; thus, the inverse problem is fundamentally one-to-many. As we gather more binary measurements, the inherent uncertainty associated with the inverse problem and any associated inference diminishes. With sufficient photon counts, photon noise becomes negligible relative to the signal mean, enabling near-deterministic scene recovery and inference. This work characterizes the transition from stochastic to near-deterministic scene understanding as photon budget increases, analyzing how the stochasticity in photon arrival affects downstream inference tasks. Technically, we develop a conditional generative framework based on a Hypergeometric frame-thinning process for accumulated binary SPAD measurements. Generative models capture the one-to-many nature of photon-starved inverse problems, enabling empirical characterization of how this ambiguity diminishes with increasing measurements and its impact on downstream tasks like character recognition, QR code decoding, and facial analysis.

eess.IV↗