Search arXiv⌕ Search

arXiv · 2609.30855

MDSkin-Net: Multi-Task Skin Lesion Analysis Driven by Pattern Analysis Priors and Spatial Alignment Regularization

Abstract

Reliable skin lesion segmentation and classification are central to dermoscopic computer-aided diagnosis. Existing multi-task frameworks couple the two tasks architecturally without clinical knowledge, while knowledge-injecting approaches rely on the macroscopic ABCD rule, which was not designed for dermoscopy. Dermoscopic diagnosis is grounded in Pattern Analysis, a microscopic framework structured around dermoscopic features. We propose MDSkin-Net, which incorporates cue-level Pattern Analysis priors into a hybrid CNN-Transformer architecture. At its core is a Pattern Analysis-Guided Attention Module (PAGAM) comprising three priors motivated by distinct dermoscopic cues: an improved Efficient Channel Attention (iECA), a Multi-Scale Spatial Attention (MSSA), and a Biased Asymmetry Attention (BAA). We further introduce a multi-scale spatial alignment regularization (MSAR) that uses the segmentation ground-truth mask as hierarchical soft supervision, confining the classification head to lesion-localized evidence and coupling both task pathways through a shared spatial prior. Trained exclusively on the ISIC 2017 training split without external dermoscopy data, the MDSkin-Net ensemble transfers robustly under zero-shot evaluation, reaching a Dice Similarity Coefficient (DSC) of 92.38% and a melanoma AUC of 97.84%on PH2, and a DSC of 88.92% on the ISIC 2018 Task 1 test set. On the in-domain ISIC 2017 benchmark, the ensemble attains a mean Area Under the Curve (AUC) of 91.60% across the two classification tasks (melanoma and seborrheic keratosis vs. rest), and a DSC of 84.72% for segmentation. Classification remains competitive with baselines; in-domain segmentation trails single-task specialists, yet the proposed priors and alignment regularization yield representations that generalize consistently across cohorts of different scales.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yijian Li, Saad Bedros, Paul Bigliardi, Mei Bigliardi Qi, Vassilios Morellas, Nikolaos Papanikolopoulos. 2026-09-25. MDSkin-Net: Multi-Task Skin Lesion Analysis Driven by Pattern Analysis Priors and Spatial Alignment Regularization. https://arxiv.org/abs/2609.30855

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Sonicmesh: Enhancing 3D Human Mesh Reconstruction in Vision-Impaired Environments With Acoustic Signals

3D human mesh reconstruction (HMR) from RGB images often degrades under poor illumination, occlusion, and non-line-of-sight conditions. Acoustic sensing provides complementary spatial cues but suffers from low spatial resolution. We propose SonicMesh, which, to the best of our knowledge, is the first acoustic--visual framework for robust 3D human mesh reconstruction. SonicMesh first converts ultrasonic echoes into range--azimuth acoustic images through an Inverse Synthetic Aperture Radar (ISAR)-based imaging process. It then introduces a cross-dimensional anatomical registration module that maps modality-specific 2D joint features into a common canonical 3D human space. The registered anatomical representations are further integrated with acoustic and visual features through a two-stage fusion network for final mesh reconstruction. Experiments demonstrate that SonicMesh achieves accurate and robust 3D human reconstruction across normal, poor-light, occluded, and non-line-of-sight environments, consistently outperforming existing RGB-, radio-frequency (RF)-, and mmWave-based approaches under challenging sensing conditions.

cs.CV↗

Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward

Recent years have witnessed significant progress in Unified Multimodal Models, yet a fundamental question remains: Does understanding truly inform generation in Unified Multimodal Models? To investigate this, we introduce UniSandbox, a decoupled evaluation framework paired with controlled, synthetic datasets to avoid data leakage and enable detailed analysis. Our findings reveal a significant understanding-generation gap, which is mainly reflected in two key dimensions: reasoning generation and knowledge transfer. Specifically, for reasoning generation tasks, we observe that explicit Chain-of-Thought (CoT) in the understanding module effectively bridges the gap, and further demonstrate that a self-training approach can successfully internalize this ability, enabling implicit reasoning during generation. Additionally, for knowledge transfer tasks, we find that CoT assists the generative process by helping retrieve newly learned knowledge, and also discover that query-based architectures inherently exhibit latent CoT-like properties that affect this transfer. UniSandbox provides preliminary insights for designing future unified architectures and training strategies that truly bridge the understanding-generation gap.

cs.CV↗

Geometric-Photometric Event-based 3D Gaussian Ray Tracing

Event cameras offer a high temporal resolution over traditional frame-based cameras, which makes them suitable for motion and structure estimation. However, it has been unclear how event-based 3D Gaussian Splatting (3DGS) approaches could leverage fine-grained temporal information of sparse events. This work proposes GPERT, a framework to address the trade-off between accuracy and temporal resolution in event-based 3DGS. Our key idea is to decouple the rendering into two branches: event-by-event geometry (depth) rendering and snapshot-based radiance (intensity) rendering, by using ray-tracing and the image of warped events. The extensive evaluation shows that our method achieves state-of-the-art performance on the real-world datasets and competitive performance on the synthetic dataset. Also, the proposed method works without prior information (e.g., pretrained image reconstruction models) or COLMAP-based initialization, is more flexible in the event selection number, and achieves sharp reconstruction on scene edges with fast training time. We hope that this work deepens our understanding of the sparse nature of events for 3D reconstruction. https://github.com/e3ai/gpert

cs.CV↗