Cross Modality Image Translation In Medical Imaging Using Generative Frameworks
Magnetic Resonance Imaging (MRI), Computed Tomography (CT), and Positron Emission Tomography (PET) provide complementary information about tissues. Medical image-to-image (I2I) translation enables virtual scanning by synthesizing a target modality from a source one without requiring an additional acquisition. Despite growing interest, many methods operate on 2D slices, are evaluated on isolated tasks under different experimental settings, and lack clinically oriented assessment. This work presents a reproducible benchmark for 3D I2I translation in oncological imaging that compares seven generative models: three Generative Adversarial Networks (Pix2Pix, CycleGAN, and SRGAN) and four latent models (Latent Diffusion Model, Latent Diffusion Model+ControlNet, Brownian Bridge, and Flow Matching). The benchmark comprises 77 experiments across eleven configurations drawn from five datasets, covering three anatomical regions (head/neck, lung, and pelvis) and four translation directions (cone-beam CT to CT, MRI to CT, CT to PET, and T2-weighted MRI to T2-FLAIR). Under the evaluated configurations, SRGAN achieves the highest quantitative image fidelity across all tasks, while latent models perform less well, due to information loss introduced by the variational autoencoder. A tumor-level analysis reveals that all models struggle with small lesions and that, in CT to PET synthesis, models reproduce tumor shape more reliably than tracer uptake values. A Visual Turing test involving 17 physicians, including 15 radiologists, shows near-chance classification accuracy (56.7\%), suggesting that experts struggle to distinguish real from synthetic volumes under the viewing conditions of the study. Expert preferences do not follow quantitative rankings, exposing a dissociation between quantitative metrics and clinical preference.