Search arXivSearch

arXiv · 2505.18058

A Pre-trained Foundation Model Framework for Multiplanar MRI Classification of Extramural Vascular Invasion and Mesorectal Fascia Invasion in Rectal Cancer

Abstract

Objectives Accurate MRI-based identification of extramural vascular invasion (EVI) and mesorectal fascia invasion (MFI) is crucial for risk-stratified rectal cancer treatment. However, subjective visual assessment and inter-institutional variability limit diagnostic consistency. This study developed and externally evaluated a multi-centre, foundation model-driven framework that automatically classifies EVI and MFI on axial and sagittal MRI. Methods A total of 331 pre-treatment rectal cancer T2-weighted MRI scans from three European hospitals were retrospectively recruited. A self-supervised frequency domain harmonization strategy was applied to reduce scanner variability. Three classifiers, SeResNet, the universal biomedical pretrained model (UMedPT) with a multilayer perceptron head, and a logistic-regression variant using frozen UMedPT features (UMedPT_LR), were trained (n=265) and tested (n=66). Gradient-weighted class activation mapping (Grad-CAM) visualized model predictions. Results UMedPT_LR achieved the best EVI performance with multiplanar fusion (AUC=0.82, test set). For MFI, UMedPT trained on axial harmonized images yielded the highest performance (AUC = 0.77). Both tasks outperformed the CHAIMELEON 2024 benchmark (EVI: 0.82 vs 0.74; MFI: 0.77 vs 0.75). Harmonization enhanced MFI classification, and multiplanar fusion further boosted EVI performance. Grad-CAM confirmed biologically plausible attention on peritumoral regions (EVI) and mesorectal fascia margins (MFI). Conclusion The proposed foundation model-driven framework, leveraging frequency domain harmonization and multiplanar fusion, achieves state-of-the-art performance for automated EVI and MFI classification on MRI, demonstrating strong generalizability across multiple centers.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yumeng Zhang, Shruti Atul Mali, Danial Khan, Sina Amirrajab, Eduardo Ibor-Crespo, Ana Jimenez-Pastor, Gloria Ribas, Silvia Flor-Arnal, Marta Zerunian, Christophe Aube, Luis Marti-Bonmati, Zohaib Salahuddin, Philippe Lambin. 2026-01-12. A Pre-trained Foundation Model Framework for Multiplanar MRI Classification of Extramural Vascular Invasion and Mesorectal Fascia Invasion in Rectal Cancer. https://arxiv.org/abs/2505.18058

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

KELP: K-space-conditioned Estimation of Learned Sampling Patterns for Scan-Adaptive Multi-Coil MRI

Deep learning techniques have gained considerable attention for accelerating MRI acquisition while maintaining image quality. In this work, we present a convolutional neural network (CNN)-based framework for predicting scan-adaptive undersampling patterns directly from low-frequency multi-coil $k$-space data. Unlike approaches that optimize sampling patterns during training or rely on nearest-neighbor search at inference, our method is trained using precomputed scan-adaptive optimized masks as supervised labels and predicts a scan-specific sampling pattern in a single forward pass. The training procedure alternates between optimizing the sampling network and a reconstruction network, allowing the learned masks to adapt to reconstruction performance. Experiments on the fastMRI multi-coil knee dataset demonstrate competitive reconstruction quality compared with existing population- and scan-adaptive sampling approaches at $4\times$ and $8\times$ acceleration factors. In addition, the proposed method enables efficient scan-specific mask prediction with inference cost independent of the training dictionary size.

eess.IV

LaminoDiff: Generative Computed Laminography via Near-Isotropic Spectral Supervision and Anisotropic Geometry

Computed Laminography (CL) is widely used for nondestructive inspection of extended planar objects, but its restricted angular coverage produces an anisotropic point spread function, missing-cone spectral incompleteness, and severe aliasing with interlayer leakage. This paper presents LaminoDiff, a physics-constrained diffusion framework for CL reconstruction. During training, a CT-derived near-isotropic supervision target is reconstructed from full-angle Computed Tomography (CT) projections physically degraded to match CL detector noise and focal-spot blur while retaining full angular coverage; it is withheld from inference. At inference, the reverse process uses only the CL observation. An Anisotropic Representation (AR) constructs depth-aware channels from three adjacent slices: a neighbor average, a center-neighbor residual, and the retained current slice, followed by directional in-plane feature extraction. Experiments on simulated and real multilayer printed circuit board data, including ball grid array and high-frequency stub samples, show that LaminoDiff improves artifact suppression, edge preservation, and depth stratification over analytic Feldkamp--Davis--Kress and representative learning-based baselines.

eess.IV

IViT: A Novel Interpretable Visual Transformer for Skin Disease Detection

The clinical diagnosis of skin diseases is susceptible to interference from inter-class similarity of skin lesions, and over-reliance on clinicians'experience easily leads to subjective bias. Although existing deep learning aided diagnosis methods achieve competitive accuracy, they suffer from the black-box opacity of Vision Transformer (ViT) and poor adaptability to medical few-shot scenarios. Moreover, mainstream explainable algorithms generally face the bottleneck of significant accuracy degradation when improving interpretability. This paper proposes an interpretable ViT (IViT) constrained by Quadratic Programming (QP). The introduced pre-trained transfer learning adapts to few-shot feature extraction. A discrete QP feature selection framework is constructed to screen generic and discriminative features consistent with clinical diagnostic logic. A multi-objective loss function is designed to reduce feature redundancy and optimize activation distribution while preserving classification performance. Experimental results on six standard skin disease datasets show that IViT achieves an accuracy of 93.80%, only 0.21% lower than the baseline, with feature redundancy reduced by 29.5%. Its core activation regions are consistent with clinically concerned lesion areas. The proposed model balances accuracy and interpretability, providing a reliable solution for the clinical deployment of few-shot intelligent skin disease diagnosis.

eess.IV