Search arXiv⌕ Search

arXiv · 2609.32399

Endo-TSR: Temporal Spectral Modeling of Appearance and Motion for Endoscopic Reconstruction

Abstract

Endoscopic scene reconstruction requires modeling tissue motion and temporal appearance while recovering fine surface detail. Deformable Gaussian models provide explicit trajectories, but their fixed colour coefficients lack a dedicated temporal representation for photometric changes. We propose Endo-TSR, which augments deformable Gaussian splatting with bounded Fourier colour residuals and independent translation residuals on shared temporal frequencies. The colour residuals capture local appearance changes, while a Matérn spectral prior regularises motion corrections. Multi-scale Laplacian supervision guides tissue-detail recovery during joint image fitting. Extensive experiments on the EndoNeRF and StereoMIS datasets demonstrate state-of-the-art rendering quality, with the highest PSNR across all evaluated sequences. Ablation studies show that temporal appearance yields the largest PSNR gain among the tested component additions, while appearance and detail supervision jointly improve rendering with fixed Gaussian counts.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Taoyu Wu, Yiyi Miao, Qi Shao, Zhuoxiao Li, Zhe Tang, Limin Yu, Baoru Huang. 2026-09-26. Endo-TSR: Temporal Spectral Modeling of Appearance and Motion for Endoscopic Reconstruction. https://arxiv.org/abs/2609.32399

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Self-eXplainable AI for Medical Image Analysis: A Survey and New Outlooks

The increasing demand for transparent and reliable models, particularly in high-stakes decision-making areas such as medical image analysis, has led to the emergence of eXplainable Artificial Intelligence (XAI). Post-hoc XAI techniques, which aim to explain black-box models after training, have raised concerns about their fidelity to model predictions. In contrast, Self-eXplainable AI (S-XAI) offers a compelling alternative by incorporating explainability directly into the training process of deep learning models. This approach allows models to generate inherent explanations that are closely aligned with their internal decision-making processes, enhancing transparency and supporting the trustworthiness, robustness, and accountability of AI systems in real-world medical applications. To facilitate the development of S-XAI methods for medical image analysis, this survey presents a comprehensive review across various image modalities and clinical applications. It covers more than 200 papers from three novel perspectives: 1) input explainability through the integration of explainable feature engineering and knowledge graph, 2) model explainability via attention-based learning, concept-based learning, and prototype-based learning, and 3) output explainability by providing textual explanations and visual counterfactual explanations. This paper also outlines desirable characteristics of S-XAI models for assessing explanation quality, while discussing major challenges and future research directions in developing S-XAI for medical image analysis.

cs.CV↗

Imitating Radiological Scrolling: A Global-Local Attention Model for 3D Chest CT Volumes Multi-Label Anomaly Classification

The rapid increase in the number of Computed Tomography (CT) scan examinations has created an urgent need for automated tools, such as organ segmentation, anomaly classification, and report generation, to assist radiologists with their growing workload. Multi-label classification of Three-Dimensional (3D) CT scans is a challenging task due to the volumetric nature of the data and the variety of anomalies to be detected. Existing deep learning methods based on Convolutional Neural Networks (CNNs) struggle to capture long-range dependencies effectively, while Vision Transformers require extensive pre-training, posing challenges for practical use. Additionally, these existing methods do not explicitly model the radiologist's navigational behavior while scrolling through CT scan slices, which requires both global context understanding and local detail awareness. In this study, we present CT-Scroll, a novel global-local attention model specifically designed to emulate the scrolling behavior of radiologists during the analysis of 3D CT scans. Our approach is evaluated on two public datasets, demonstrating its efficacy through comprehensive experiments and an ablation study that highlights the contribution of each model component.

cs.CV↗

Noise-adapted Neural Operator for Robust Non-Line-of-Sight Imaging

This work has been submitted for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible. Computational imaging, especially non-line-of-sight (NLOS) imaging, the extraction of information from obscured or hidden scenes is achieved through the utilization of indirect light signals resulting from multiple reflections or scattering. The inherently weak nature of these signals, coupled with their susceptibility to noise, necessitates the integration of physical processes to ensure accurate reconstruction. This paper presents a parameterized inverse problem framework tailored for large-scale linear problems in 3D imaging reconstruction. Initially, a noise estimation module is employed to adaptively assess the noise levels present in transient data. Subsequently, a parameterized neural operator is developed to approximate the inverse mapping, facilitating end-to-end rapid image reconstruction. Our 3D image reconstruction framework, grounded in operator learning, is constructed through deep algorithm unfolding, which not only provides commendable model interpretability but also enables dynamic adaptation to varying noise levels in the acquired data, thereby ensuring consistently robust and accurate reconstruction outcomes. Furthermore, we introduce a novel method for the fusion of global and local spatiotemporal data features. By integrating structural and detailed information, this method significantly enhances both accuracy and robustness. Comprehensive numerical experiments conducted on both simulated and real datasets substantiate the efficacy of the proposed method. It demonstrates remarkable performance with fast scanning data and sparse illumination point data, offering a viable solution for NLOS imaging in complex scenarios.

cs.CV↗