arXiv · 2609.22734
Clinical Domain Classification from Medical Transcriptions
Abstract
Clinical domain classification plays an important role in organizing and analyzing large volumes of unstructured medical text. However, medical transcription datasets are often highly imbalanced, which can substantially degrade classification performance, particularly for underrepresented clinical specialties. In this work, we present a comparative study of machine learning and transformer-based approaches for clinical domain classification from medical transcriptions. We evaluate six traditional machine learning classifiers---Naive Bayes, Support Vector Machine (SVM), Decision Tree, Random Forest, K-Nearest Neighbors (KNN), and XGBoost---along with two pretrained transformer models, BERT and XLNet, and a few-shot large language model prompting approach. Experiments are conducted on medical transcription data collected from MTSamples, comprising 5,013 samples across 40 clinical specialties. To address severe class imbalance, we investigate two balancing strategies: text augmentation using NLP-based synonym replacement and Synthetic Minority Over-sampling Technique (SMOTE). Experimental results demonstrate that data balancing substantially improves classification performance across the evaluated models. In particular, BERT achieves the highest F1-score of 0.996 on the SMOTE-balanced dataset while requiring lower training time than XLNet. The results highlight the effectiveness of transformer-based representations combined with appropriate data balancing strategies for clinical domain classification and provide a systematic comparison of classical machine learning, transformer models, and few-shot prompting for medical transcription analysis.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sravani Pottipati, Lakshmikar R. Polamreddy. 2026-09-19. Clinical Domain Classification from Medical Transcriptions. https://arxiv.org/abs/2609.22734
Cite the original work for its findings. Save a collection to share your selection of sources.