Search arXivSearch

arXiv subjects

Dimitris Visvikis

Publications and source records attributed to Dimitris Visvikis.

At least 19 recordsLinked to original sources

Continuous 3-D Latent Diffusion for Medical Image Generation and Reconstruction

High-resolution three-dimensional (3-D) medical diffusion models remain constrained by the cost of processing full volumes, even when denoising is performed in a compact latent space. We introduce a continuous 3-D latent diffusion model (LDM) framework for computed tomography (CT) and magnetic resonance imaging (MRI) generation and measurement-guided reconstruction. Its central component is a compact autoencoder (AE) with a coordinate-conditioned local implicit image function (LIIF) decoder that represents a volume as a continuous function of spatial coordinates. By evaluating the convolutional decoder once on the latent grid and restricting repeated computation to a lightweight implicit head, the proposed design avoids overlapping sub-volume decoding while remaining differentiable for inverse-problem optimization. We evaluate the framework on CT volumes of 512^3 voxels and MRI volumes of 256^3 voxels. On high-resolution CT, the proposed AE is approximately x12-32 faster than the evaluated reference autoencoders, achieves the lowest peak graphics processing unit (GPU) memory use, and retains comparable structural fidelity despite a moderate reduction in voxel-level accuracy. The resulting frozen 3-D latent prior generates coherent full volumes without visible patch seams and can be applied, without task-specific retraining, to sparse-view CT and accelerated MRI reconstruction through hard data consistency. Although direct pixel-domain reconstruction remains more accurate, the results demonstrate that a single volumetric latent prior can support both unconditional generation and measurement-conditioned reconstruction on one GPU. Overall, the framework provides a practical trade-off between continuous volumetric decoding, computational efficiency, and fine-detail preservation. Our code will be made available at https://github.com/mellak/.

physics.med-ph

Multilevel Stochastic Plug-and-Play for Sparse-View CT Reconstruction

Sparse-view computed tomography (SVCT) reduces radiation exposure and acquisition time, but the limited number of projection views makes the reconstruction problem severely ill-posed and leads to streak artifacts when analytical methods are used. Plug-and-Play (PnP) methods provide an effective way to combine data fidelity with learned image priors, while stochastic PnP methods further improve robustness by matching the denoiser input distribution through re-noising. However, these methods often require many iterations to converge, which limits their practical efficiency. In this work, we propose a multilevel (ML) stochastic PnP method for SVCT that accelerates stochastic PnP reconstruction. We highlight that, in the stochastic setting, directly enforcing prior coherence across levels would require accurately estimating fine-level prior gradients through multiple denoiser function evaluations, which substantially increases the computational cost. Motivated by this observation, we perform the multilevel steps in multiresolution analysis (MRA) approximation spaces. This choice is supported by the structure of the wavelet decomposition, which causes the prior-coherence correction to vanish in expectation, thereby avoiding costly estimation of fine-level stochastic prior gradients for the coarse-level corrections. Experiments on SVCT reconstruction show that our method, called Multilevel Stochastic Plug-and-Play (ML-SPnP), achieves reconstruction quality comparable to state-of-the-art methods while substantially reducing runtime.

cs.CV

A Fast and Generic Energy-Shifting Transformer for Hybrid Monte Carlo Radiotherapy Calculation

We introduce a novel learning framework for accelerated Monte Carlo (MC) dose calculation termed Energy-Shifting. This approach leverages deep learning to synthesize highly complex polyenergetic dose distributions directly from simple monoenergetic inputs under identical beam configurations. Unlike conventional denoising techniques, which rely on noisy low-count dose maps that compromise beam profile integrity, our method achieves superior cross-domain generalization on unseen datasets by integrating high-fidelity anatomical textures and source-specific beam similarity into the model's input space. Furthermore, we propose a novel 3D architecture termed TransUNetSE3D, featuring Transformer blocks for global context and Residual Squeeze-and-Excitation (SE) modules for adaptive channel-wise feature recalibration. Hierarchical representations of these blocks are fused into the network's latent space alongside the primary dose-map parameters, allowing physics-aware reconstruction. This hybrid design outperforms existing UNet and Transformer-based benchmarks in both spatial precision and structural preservation, while maintaining the execution speed necessary for real-time use. Our proposed pipeline achieves a Gamma Passing Rate exceeding 98% (3%/3mm) compared to the MC reference, evaluated within the framework of a treatment planning system (TPS) using 6MV TrueBeam Lineac Accelerator (LINAC) for prostate radiotherapy. These results offer a robust solution for fast volumetric dosimetry in adaptive radiotherapy.

physics.med-ph

Joint Reconstruction of Activity and Attenuation in PET by Diffusion Posterior Sampling in Wavelet Coefficient Space

Attenuation correction (AC) is necessary for accurate activity quantification in positron emission tomography (PET). Conventional reconstruction methods typically rely on attenuation maps derived from a co-registered computed tomography (CT) or magnetic resonance (MR) scan. However, this additional scan may complicate the imaging workflow, introduce misalignment artifacts and increase radiation exposure. In this paper, we propose a joint reconstruction of activity and attenuation (JRAA) approach that eliminates the need for auxiliary anatomical imaging by relying solely on emission data. This framework combines wavelet diffusion model (WDM) and diffusion posterior sampling (DPS) to reconstruct fully three-dimensional (3-D) data. Experimental results on simulated data show our method outperforms maximum likelihood activity and attenuation (MLAA) and MLAA-UNet with U-Net-based postprocessing, and yields high-quality noise-free reconstructions across various count settings with time-of-flight (TOF). It is also able to reconstruct non-TOF data, although the reconstruction quality significantly degrades in low-count (LC) conditions, limiting its practical effectiveness in such settings. Nonetheless, a non-TOF Biograph mMR real data reconstruction with joint scatter estimation highlights the potential of the method for clinical applications. This approach represents a step towards stand-alone PET imaging by reducing the dependence on anatomical modalities while maintaining quantification accuracy, even in LC scenarios when TOF information is available. Our code is available on GitHub at https://github.com/clemphg/jraa-dps.

physics.med-ph

Adaptive Diffusion Models for Sparse-View Motion-Corrected Head Cone-beam CT

Cone-beam computed tomography (CBCT) is an imaging modality widely used in head and neck diagnostics due to its accessibility and lower radiation dose. However, its relatively long acquisition times make it susceptible to patient motion, especially under sparse-view settings used to reduce dose, which can result in severe image artifacts. In this work, we propose a novel framework combining joint reconstruction and motion estimation (JRM) with an adaptive diffusion model (ADM) that simultaneously addresses motion compensation and sparse-view reconstruction in head CBCT. Leveraging recent advances in diffusion-based generative models, our method integrates a wavelet-domain diffusion prior into an iterative reconstruction pipeline to guide the solution toward anatomically plausible volumes while estimating rigid motion parameters in a blind fashion. We evaluate our method on simulated motion-affected CBCT data derived from real clinical computed tomography (CT) volumes. Experimental results demonstrate that JRM- ADM achieves consistent quantitative improvements over both traditional and learning-based baselines. In highly undersampled cases, JRM-ADM improves peak signal-to-noise ratio (PSNR) by more than 4 dB and structural similarity index measure (SSIM) by 0.10 compared to the baseline motion-corrected (MC) reconstruction method. These results highlight the potential of our approach to enable motion-robust, low-dose CBCT imaging, paving the way for improved clinical viability. The project page is available at https://antoinedepaepe.github.io/jrm-adm-io/.

physics.med-ph

Material Decomposition in Photon-Counting Computed Tomography with Diffusion Models: Comparative Study and Hybridization with Variational Regularizers

Photon-counting computed tomography (PCCT) has emerged as a promising imaging technique, enabling spectral imaging and material decomposition (MD). However, images typically suffer from a low signal-to-noise ratio (SNR) due to constraints such as low photon counts and sparse-view settings which provoke artifacts. To prevent this, variational methods minimize a data-fit function coupled with handcrafted regularizers that mimic a prior by enforcing image properties such as gradient sparsity. In the last few years, diffusion models (DMs) have become predominant in the field of generative models and have been used as a learned prior for image reconstruction. This work investigates the use of DMs as regularizers for MD tasks in PCCT, specifically using diffusion posterior sampling (DPS) guidance. Three DPS-based approaches -- image-domain two-step DPS (im-TDPS), projection-domain two-step DPS (proj-TDPS), and one-step DPS (ODPS) -- are evaluated. The first two methods achieve MD in two steps by performing reconstruction and MD separately. The last method, ODPS, samples the material images directly from the measurement data. The results indicate that ODPS achieves superior performance compared to im-TDPS and proj-TDPS, providing sharper, noise-free and crosstalk-free images. Furthermore, we introduce a novel hybrid method for scenarios involving materials absent from the training dataset which combines DM priors with standard variational handcrafted regularizers for the materials unknown to the DM. This hybrid method demonstrates improved MD quality compared to a standard variational method and does not require additional training of the DM neural network (NN).

physics.med-ph

Dual-Input Dynamic Convolution for Positron Range Correction in PET Image Reconstruction

Positron range (PR) blurring degrades positron emission tomography (PET) image resolution, particularly for high-energy emitters like gallium-68 (68 Ga). We introduce Dual-Input Dynamic Convolution (DDConv), a novel computationally efficient approach trained with voxel-specific PR point spread functions (PSFs) from Monte Carlo (MC) simulations and designed to be utilized within an iterative reconstruction algorithm to perform PR correction (PRC). By dynamically inferring local blurring kernels through a trained convolutional neural network (CNN), DDConv captures complex tissue interfaces more accurately than prior methods. Additionally, it also computes the transpose operator, ensuring consistency within iterative PET reconstruction. Comparisons with a state-of-the-art, tissue-dependent correction confirm the advantages of DDConv in recovering higher-resolution details in heterogeneous regions, including bone-soft tissue and lung-soft tissue boundaries. Experiments across digital phantoms and MC-simulated data show that DDConv offers near-MC accuracy and outperforms the state-of-the-art technique, namely spatially-variant and tissue-dependent (SVTD), especially in areas with complex material interfaces. Results from real phantom experiments further confirm DD-Conv's robustness and practical applicability: while both DD-Conv and SVTD performed similarly in homogeneous soft-tissue regions, DDConv provided more accurate activity recovery and sharper delineation at heterogeneous lung-soft tissue interfaces. Our code available at https://github.com/mellak/ddconv-prc.

physics.med-ph

Solving Blind Inverse Problems: Adaptive Diffusion Models for Motion-corrected Sparse-view 4DCT

Four-dimensional computed tomography (4DCT) is essential for medical imaging applications like radiotherapy, which demand precise respiratory motion representation. Traditional methods for reconstructing 4DCT data suffer from artifacts and noise, especially in sparse-view, low-dose contexts. Motion-corrected (MC) reconstruction is a blind inverse problem that we propose to solve with a novel diffusion model (DM) framework that calibrates an adaptive unknown forward model for motion correction. Furthermore, we used a wavelet diffusion model (WDM) to address computational cost and memory usage. By leveraging the prior probability distribution function (PDF) from the DMs, we enhance the joint reconstruction and motion estimation (JRM) process, improving image quality and preserving resolution. Experiments on extended cardiac-torso (XCAT) phantom data demonstrate that our method outperforms existing techniques, yielding artifact-free, high-resolution reconstructions even under irregular breathing conditions. These results showcase the potential of combining DMs with motion correction to advance sparse-view 4DCT imaging.

physics.med-ph

Evaluation of Deep Learning-based Scatter Correction on a Long-axial Field-of-view PET scanner

Objective: Long-axial field-of-view (LAFOV) positron emission tomography (PET) systems allow higher sensitivity, with an increased number of detected lines of response induced by a larger angle of acceptance. However, this extended angle increases the number of multiple scatters and the scatter contribution within oblique planes. As scattering affects both quality and quantification of the reconstructed image, it is crucial to correct this effect with more accurate methods than the state-of-the-art single scatter simulation (SSS) that can reach its limits with such an extended field-of-view (FOV). In this work, which is an extension of our previous assessment of deep learning-based scatter estimation (DLSE) carried out on a conventional PET system, we aim to evaluate the DLSE method performance on LAFOV total-body PET. Approach: The proposed DLSE method based on a convolutional neural network (CNN) U-Net architecture uses emission and attenuation sinograms to estimate scatter sinogram. The network was trained from Monte-Carlo (MC) simulations of XCAT phantoms [18F]-FDG PET acquisitions using a Siemens Biograph Vision Quadra scanner model, with multiple morphologies and dose distributions. We firstly evaluated the method performance on simulated data in both sinogram and image domain by comparing it to the MC ground truth and SSS scatter sinograms. We then tested the method on seven [18F]-FDG and seven [18F]-PSMA clinical datasets, and compare it to SSS estimations. Results: DLSE showed superior accuracy on phantom data, greater robustness to patient size and dose variations compared to SSS, and better lesion contrast recovery. It also yielded promising clinical results, improving lesion contrasts in [18F]-FDG datasets and performing consistently with [18F]-PSMA datasets despite no training with [18F]-PSMA.

physics.med-ph

Joint Reconstruction of the Activity and the Attenuation in PET by Diffusion Posterior Sampling: a Feasibility Study

This study introduces a novel framework for joint reconstruction of the activity and the attenuation (JRAA) in positron emission tomography (PET) using diffusion posterior sampling (DPS). By leveraging diffusion models (DMs), this approach directly addresses activity-attenuation dependencies, mitigating crosstalk issues prevalent in non-time-of-flight (TOF) settings. Experimental evaluations, conducted using 2-dimensional (2-D) XCAT phantom data, demonstrate that DPS significantly outperforms traditional maximum likelihood activity and attenuation (MLAA) methods, producing consistent and high-quality reconstructions even in the absence of TOF information. Ongoing work aims to extend our method to real 3-dimensional (3-D) data with encouraging preliminary findings.

physics.med-ph

Synergistic PET/CT Reconstruction Using a Joint Generative Model

We propose in this work a framework for synergistic positron emission tomography (PET)/computed tomography (CT) reconstruction using a joint generative model as a penalty. We use a synergistic penalty function that promotes PET/CT pairs that are likely to occur together. The synergistic penalty function is based on a generative model, namely $\beta$-variational autoencoder ($\beta$-VAE). The model generates a PET/CT image pair from the same latent variable which contains the information that is shared between the two modalities. This sharing of inter-modal information can help reduce noise during reconstruction. Our result shows that our method was able to utilize the information between two modalities. The proposed method was able to outperform individually reconstructed images of PET (i.e., by maximum likelihood expectation maximization (MLEM)) and CT (i.e., by weighted least squares (WLS)) in terms of peak signal-to-noise ratio (PSNR). Future work will focus on optimizing the parameters of the $\beta$-VAE network and further exploration of other generative network models.

physics.med-ph

Semi-overcomplete convolutional auto-encoder embedding as shape priors for deep vessel segmentation

The extraction of blood vessels has recently experienced a widespread interest in medical image analysis. Automatic vessel segmentation is highly desirable to guide clinicians in computer-assisted diagnosis, therapy or surgical planning. Despite a good ability to extract large anatomical structures, the capacity of U-Net inspired architectures to automatically delineate vascular systems remains a major issue, especially given the scarcity of existing datasets. In this paper, we present a novel approach that integrates into deep segmentation shape priors from a Semi-Overcomplete Convolutional Auto-Encoder (S-OCAE) embedding. Compared to standard Convolutional Auto-Encoders (CAE), it exploits an over-complete branch that projects data onto higher dimensions to better characterize tiny structures. Experiments on retinal and liver vessel extraction, respectively performed on publicly-available DRIVE and 3D-IRCADb datasets, highlight the effectiveness of our method compared to U-Net trained without and with shape priors from a traditional CAE.

eess.IV

Scale-specific auxiliary multi-task contrastive learning for deep liver vessel segmentation

Extracting hepatic vessels from abdominal images is of high interest for clinicians since it allows to divide the liver into functionally-independent Couinaud segments. In this respect, an automated liver blood vessel extraction is widely summoned. Despite the significant growth in performance of semantic segmentation methodologies, preserving the complex multi-scale geometry of main vessels and ramifications remains a major challenge. This paper provides a new deep supervised approach for vessel segmentation, with a strong focus on representations arising from the different scales inherent to the vascular tree geometry. In particular, we propose a new clustering technique to decompose the tree into various scale levels, from tiny to large vessels. Then, we extend standard 3D UNet to multi-task learning by incorporating scale-specific auxiliary tasks and contrastive learning to encourage the discrimination between scales in the shared representation. Promising results, depicted in several evaluation metrics, are revealed on the public 3D-IRCADb dataset.

eess.IV

Deep vessel segmentation with joint multi-prior encoding

The precise delineation of blood vessels in medical images is critical for many clinical applications, including pathology detection and surgical planning. However, fully-automated vascular segmentation is challenging because of the variability in shape, size, and topology. Manual segmentation remains the gold standard but is time-consuming, subjective, and impractical for large-scale studies. Hence, there is a need for automatic and reliable segmentation methods that can accurately detect blood vessels from medical images. The integration of shape and topological priors into vessel segmentation models has been shown to improve segmentation accuracy by offering contextual information about the shape of the blood vessels and their spatial relationships within the vascular tree. To further improve anatomical consistency, we propose a new joint prior encoding mechanism which incorporates both shape and topology in a single latent space. The effectiveness of our method is demonstrated on the publicly available 3D-IRCADb dataset. More globally, the proposed approach holds promise in overcoming the challenges associated with automatic vessel delineation and has the potential to advance the field of deep priors encoding.

eess.IV

Direct3{\gamma}: A Pipeline for Direct Three-gamma PET Image Reconstruction

This paper presents a novel image reconstruction pipeline for three-gamma (3-{\gamma}) positron emission tomography (PET) aimed at improving spatial resolution and reducing noise in nuclear medicine; the proposed Direct3{\gamma} pipeline addresses the inherent challenges in 3-{\gamma} PET systems, such as detector imperfections and uncertainty in photon interaction points, with a key feature being its ability to determine the order of interactions through a model trained on Monte Carlo (MC) simulations using the Geant4 Application for Tomography Emission (GATE) toolkit, thus providing the necessary information to construct Compton cones which intersect with the line of response (LOR) to estimate the emission point; the pipeline processes 3-{\gamma} PET raw data, reconstructs histoimages by propagating energy and spatial uncertainties along the LOR, and applies a 3-D convolutional neural network (CNN) to refine these intermediate images into high-quality reconstructions, further enhancing image quality through supervised learning and adversarial losses that preserve fine structural details; experimental results show that Direct3{\gamma} consistently outperforms conventional 200-ps time-of-flight (TOF) PET in terms of structural similarity index measure (SSIM) and peak signal-to-noise ratio (PSNR).

physics.med-ph

Multibranch Generative Models for Multichannel Imaging with an Application to PET/CT Synergistic Reconstruction

This paper presents a novel approach for learned synergistic reconstruction of medical images using multibranch generative models. Leveraging variational autoencoders (VAEs), our model learns from pairs of images simultaneously, enabling effective denoising and reconstruction. Synergistic image reconstruction is achieved by incorporating the trained models in a regularizer that evaluates the distance between the images and the model. We demonstrate the efficacy of our approach on both Modified National Institute of Standards and Technology (MNIST) and positron emission tomography (PET)/computed tomography (CT) datasets, showcasing improved image quality for low-dose imaging. Despite challenges such as patch decomposition and model limitations, our results underscore the potential of generative models for enhancing medical imaging reconstruction.

eess.IV

CT respiratory motion synthesis using joint supervised and adversarial learning

Objective: Four-dimensional computed tomography (4DCT) imaging consists in reconstructing a CT acquisition into multiple phases to track internal organ and tumor motion. It is commonly used in radiotherapy treatment planning to establish planning target volumes. However, 4DCT increases protocol complexity, may not align with patient breathing during treatment, and lead to higher radiation delivery. Approach: In this study, we propose a deep synthesis method to generate pseudo respiratory CT phases from static images for motion-aware treatment planning. The model produces patient-specific deformation vector fields (DVFs) by conditioning synthesis on external patient surface-based estimation, mimicking respiratory monitoring devices. A key methodological contribution is to encourage DVF realism through supervised DVF training while using an adversarial term jointly not only on the warped image but also on the magnitude of the DVF itself. This way, we avoid excessive smoothness typically obtained through deep unsupervised learning, and encourage correlations with the respiratory amplitude. Main results: Performance is evaluated using real 4DCT acquisitions with smaller tumor volumes than previously reported. Results demonstrate for the first time that the generated pseudo-respiratory CT phases can capture organ and tumor motion with similar accuracy to repeated 4DCT scans of the same patient. Mean inter-scans tumor center-of-mass distances and Dice similarity coefficients were $1.97$mm and $0.63$, respectively, for real 4DCT phases and $2.35$mm and $0.71$ for synthetic phases, and compares favorably to a state-of-the-art technique (RMSim).

cs.CV

Spectral CT Two-step and One-step Material Decomposition using Diffusion Posterior Sampling

This paper proposes a novel approach to spectral computed tomography (CT) material decomposition that uses the recent advances in generative diffusion models (DMs) for inverse problems. Spectral CT and more particularly photon-counting CT (PCCT) can perform transmission measurements at different energy levels which can be used for material decomposition. It is an ill-posed inverse problem and therefore requires regularization. DMs are a class of generative model that can be used to solve inverse problems via diffusion posterior sampling (DPS). In this paper we adapt DPS for material decomposition in a PCCT setting. We propose two approaches, namely Two-step Diffusion Posterior Sampling (TDPS) and One-step Diffusion Posterior Sampling (ODPS). Early results from an experiment with simulated low-dose PCCT suggest that DPSs have the potential to outperform state-of-the-art model-based iterative reconstruction (MBIR). Moreover, our results indicate that TDPS produces material images with better peak signal-to-noise ratio (PSNR) than images produced with ODPS with similar structural similarity (SSIM).

physics.med-ph