arXiv · 2602.23994
MINT: Multimodal Imaging-to-Speech Knowledge Transfer for Early Alzheimer's Screening
Abstract
Alzheimer's disease is a progressive neurodegenerative disorder in which mild cognitive impairment (MCI) precedes dementia. Structural MRI provides biomarkers but requires costly infrastructure, limiting population-scale deployment. Speech offers a non-invasive alternative, yet speech-only classifiers are developed independently of neuroimaging and lack biological grounding for CN-versus-MCI classification. We propose MINT (Multimodal Imaging-to-Speech Knowledge Transfer), a three-stage framework that transfers MRI-derived biomarker structure to speech during training. An MRI teacher defines a compact embedding space for CN-versus-MCI classification, while a residual projection head aligns speech representations to this space using a combined geometric loss. The frozen MRI classifier enables imaging-free inference. On ADNI-4, aligned speech achieves performance comparable to speech baselines, while multimodal fusion improves over MRI alone. Ablations identify dropout regularization and self-supervised pretraining as important design choices. To our knowledge, MINT is the first demonstration of MRI-to-speech knowledge transfer for early Alzheimer's screening without imaging at inference.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Vrushank Ahire, Yogesh Kumar, Anouck Girard, M. A. Ganaie. 2026-09-16. MINT: Multimodal Imaging-to-Speech Knowledge Transfer for Early Alzheimer's Screening. https://arxiv.org/abs/2602.23994
Cite the original work for its findings. Save a collection to share your selection of sources.