Search arXiv⌕ Search

arXiv · 2511.20956

BUSTR: Descriptor-Aware Vision-Language Learning for Breast Ultrasound Report Generation

Abstract

Breast ultrasound (BUS) reporting relies on clinically meaningful lesion descriptors, including BI-RADS category, lesion shape, margin, echogenicity, posterior features, pathology, and histology. However, many public BUS datasets provide structured annotations and lesion masks without paired radiologist-written reports, limiting the development of vision--language models for BUS report generation. We propose BUSTR, a descriptor-aware vision--language framework that uses structured lesion information to enable report generation under limited report supervision. BUSTR first constructs descriptor-derived reports from available annotations and radiomics features extracted from lesion masks. It then trains a multi-head Swin Transformer encoder with multitask supervision to learn descriptor-aware visual representations across datasets with partially overlapping annotation sets. The projected visual tokens condition a frozen LLaMA-based language model, and training is guided by a dual-level objective combining token-level cross-entropy with representation-level cosine alignment. At inference, BUSTR generates reports from BUS images without access to structured descriptors, lesion masks, or radiomics features. We evaluate BUSTR on the public BrEaST and BUS-BRA datasets using natural language generation and clinical efficacy metrics. BUSTR improves report similarity and descriptor recovery compared with representative report-generation baselines, with notable gains for lesion shape, margin, posterior features, and pathology, as well as improved BI-RADS sensitivity and F1-score on BrEaST. These results suggest that structured BUS descriptors, lesion masks, and radiomics features can provide useful supervision for descriptor-aware BUS report generation when paired radiologist-written reports are unavailable.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Rawa Mohammed, Mina Attin, Laxmi Gewali, Bryar Shareef. 2026-07-20. BUSTR: Descriptor-Aware Vision-Language Learning for Breast Ultrasound Report Generation. https://arxiv.org/abs/2511.20956

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Band-Attention Modulation Network for Robust Face Forgery Detection

Face forgery detection faces critical challenges in generalizing to unseen manipulation techniques and remaining robust under image compression, which often obscures subtle artifacts. Existing methods typically rely on fixed filters or coarse band separation, lacking the adaptability to learn task-specific spectral cues. To address this, we propose the Band-Attention Modulation Network (BAM-Net), a novel framework that pioneers learnable, fine-grained modulation of frequency components for forgery detection. At its core is the Band-Attention Modulation (BAM) mechanism, which transforms an image into its Discrete Cosine Transform (DCT) spectrogram and learns to dynamically reweight frequency bands along anti-diagonals. This process effectively enhances forgery-related spectral signatures while suppressing less informative ones, simulating an adaptive "inverse compression" that counters information loss. The modulated frequency information is then fused with the spatial domain to guide a lightweight yet effective spatial backbone equipped with distance-decayed attention for comprehensive feature extraction. Extensive experiments on FaceForensics++, Celeb-DF, and DFDC datasets demonstrate that BAM-Net achieves state-of-the-art performance. More importantly, it exhibits exceptional generalization in cross-dataset, cross-compression, and cross-manipulation scenarios, underscoring the vital role of adaptive frequency band modulation in building robust forgery detectors.

cs.CV↗

Cross-Task Generalization in Handwriting-Based Alzheimer's Screening via Vision Language Adaptation

Alzheimer's disease (AD) is a prevalent neurodegenerative disorder for which early detection is critical. Handwriting, which can be disrupted by subtle motor and cognitive decline, provides a non-invasive and cost-effective window for AD screening. Existing handwriting-based AD studies mostly rely on online trajectories and hand-crafted features, while the influence of handwriting task type on diagnostic performance and cross-task generalization remains underexplored. Meanwhile, large-scale vision--language models have demonstrated strong transfer and adaptation ability in natural-image anomaly detection and several medical modalities, such as chest X-ray and brain MRI. However, handwriting-based disease detection remains unexplored within this paradigm. To address this gap, we introduce a lightweight Cross-Layer Fusion Adapter (CLFA) framework that repurposes Contrastive Language--Image Pre-training (CLIP) for handwriting-based AD screening. CLFA inserts multi-level adapters into a frozen visual encoder, combining cross-layer feature fusion with depthwise 2D convolution on patch grids to capture both local stroke irregularities and higher-level handwriting structure. This design progressively aligns pretrained vision--language representations with AD-related handwriting cues and supports transfer from supervised source tasks to task-disjoint unseen target tasks. On the Darwin dataset, under the subject-disjoint cross-task protocol, averaged over all 600 task-disjoint source-target pairs, CLFA achieves 74.63\% AUC, 74.85\% accuracy, and 73.72\% F1 score, outperforming the best competing model by 2.15, 1.79, and 1.87 percentage points, respectively.

cs.CV↗

LeafTrackNet: A Deep Learning Framework for Robust Leaf Tracking in Top-Down Plant Phenotyping

High-resolution phenotyping at the level of individual leaves offers fine-grained insights into plant development and stress responses. However, the full potential of accurate leaf tracking over time remains largely unexplored due to the absence of robust tracking methods, particularly for structurally complex crops such as canola. Existing plant-specific tracking methods are typically limited to small-scale species or rely on constrained imaging conditions. In contrast, generic multi-object tracking (MOT) methods are not designed for dynamic biological scenes. Progress in the development of accurate leaf tracking models has also been hindered by a lack of large-scale datasets captured under realistic conditions. In this work, we introduce CanolaTrack, a new benchmark dataset comprising 5704 RGB images with 31,840 annotated leaf instances collected from 184 canola plants during their early growth stages. To enable accurate leaf tracking over time, we introduce LeafTrackNet, an efficient framework that combines a YOLOv10-based leaf detector with a MobileNetV3-based embedding network. During inference, leaf identities are maintained over time through an embedding-based memory association strategy. When trained directly on each target dataset without prior CanolaTrack fine-tuning, LeafTrackNet achieves HOTA scores of 88.03, 87.33, and 74.20 on CanolaTrack, KOMATSUNA, and MSU-PID, respectively, outperforming the corresponding second-best methods by 8.35, 4.94, and 1.62 HOTA points. This work provides a new benchmark for leaf-level tracking under realistic conditions and introduces CanolaTrack, which, to the best of our knowledge, is the largest leaf-tracking dataset for agricultural crops. Our code and dataset are publicly available at GitHub.

cs.CV↗