Search arXivSearch

arXiv · 2504.10916

Embedding Radiomics into Vision Transformers for Multimodal Medical Image Classification

Abstract

Background: Deep learning has significantly advanced medical image analysis, with Vision Transformers (ViTs) offering a powerful alternative to convolutional models by modeling long-range dependencies through self-attention. However, ViTs are inherently data-intensive and lack domain-specific inductive biases, limiting their applicability in medical imaging. In contrast, radiomics provides interpretable, handcrafted descriptors of tissue heterogeneity but suffers from limited scalability and integration into end-to-end learning frameworks. In this work, we propose the Radiomics-Embedded Vision Transformer (RE-ViT) that combines radiomic features with data-driven visual embeddings within a ViT backbone. Purpose: To develop a hybrid RE-ViT framework that integrates radiomics and patch-wise ViT embeddings through early fusion, enhancing robustness and performance in medical image classification. Methods: Following the standard ViT pipeline, images were divided into patches. For each patch, handcrafted radiomic features were extracted and fused with linearly projected pixel embeddings. The fused representations were normalized, positionally encoded, and passed to the ViT encoder. A learnable [CLS] token aggregated patch-level information for classification. We evaluated RE-ViT on three public datasets (including BUSI, ChestXray2017, and Retinal OCT) using accuracy, macro AUC, sensitivity, and specificity. RE-ViT was benchmarked against CNN-based (VGG-16, ResNet) and hybrid (TransMed) models. Results: RE-ViT achieved state-of-the-art results: on BUSI, AUC=0.950+/-0.011; on ChestXray2017, AUC=0.989+/-0.004; on Retinal OCT, AUC=0.986+/-0.001, which outperforms other comparison models. Conclusions: The RE-ViT framework effectively integrates radiomics with ViT architectures, demonstrating improved performance and generalizability across multimodal medical image classification tasks.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhenyu Yang, Haiming Zhu, Rihui Zhang, Haipeng Zhang, Jianliang Wang, Chunhao Wang, Minbin Chen, Fang-Fang Yin. 2025-04-23. Embedding Radiomics into Vision Transformers for Multimodal Medical Image Classification. https://arxiv.org/abs/2504.10916

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Prediction of biological radiation effects based on ionization clusters (nanodosimetry)

This article reviews approaches that link the formation of ionization clusters in nanometric volumes to radiobiological effectiveness. The corresponding models were developed as the field of nanodosimetry developed. Some address early biological radiation effects, such as DNA damage, while most aim to predict cell survival or inactivation. The models also differ in the nanodosimetric quantities considered, with many based on the probability distribution of ionization cluster formation in a single target. Some models account for the synergistic effects of pairs of ionization clusters formed in different targets. Several models feature macroscopic aggregation frameworks based on particle fluence, which are proposed for use in radiotherapy treatment planning, particularly in ion-beam radiotherapy. The models are presented here using harmonized terminology and notation for nanodosimetric quantities. An extension of the conceptual framework of nanodosimetry is also discussed. This extension transitions from a target-centered description to a track-centered description. It also introduces nanodosimetry-based analogs of dosimetric concepts, such as dose and linear energy transfer. This paper traces and summarizes the historical development of nanodosimetry-based biological effect models and discusses conceptual aspects of the models to reveal their underlying assumptions and the extent to which they are mechanistic or merely elucidate correlations. Eventually, an attempt is made to identify the key open questions in this field that still need to be addressed.

physics.med-ph

Contextual Cellular Growth (ConCeG) of neural cells for realistic grey matter tissue generation for diffusion MRI simulations

Accurate interpretation of diffusion magnetic resonance imaging (dMRI) signals in grey matter (GM) remains challenging due to the complex, heterogeneous, and densely packed cellular environment. Numerical phantoms provide a controlled framework for investigating the relationship between microstructure and diffusion signals, yet existing approaches often lack the morphological realism and multi-cellular organisation required to faithfully represent GM tissue. In this work, we introduce Contextual Cellular Growth (ConCeG), a generative framework for creating individual cells or constructing dense, three-dimensional, multi-cellular GM substrates informed by real neuronal and glial morphologies. The method combines topological neuron synthesis with a spatially constrained growth network, allowing for the controlled generation of heterogeneous cellular environments with realistic intra- and extracellular compartments. Synthetic cells are generated using morphological and topological characteristics derived from biological reconstructions. We validate the framework through comparisons of structural features with real cellular data, demonstrating strong agreement in branch order, length, angle, and tortuosity distributions. Power spectrum analysis further shows that both intracellular compartments reproduce the spatial correlations observed in biological tissue. Together, these results show ConCeG provides a biologically grounded framework for generating grey matter substrates suitable for large scale diffusion MRI simulation.

physics.med-ph

Magnetic Field of Firing Neuron in Humans: Measurable by Quantum Sensing MRI?

Firing neurons generate action potentials that propagate along axons to transmit signals supporting cognitive functions. These electrical currents generate magnetic fields, yet direct detection of these neuronal magnetic fields by MRI remains elusive. This Mini Review investigates why this goal has been proven difficult to achieve and whether an emerging approach, quantum sensing MRI, can overcome the challenge.

physics.med-ph