arXiv · 2311.00768
Language Model Training Paradigms for Clinical Feature Embeddings
Abstract
In research areas with scarce data, representation learning plays a significant role. This work aims to enhance representation learning for clinical time series by deriving universal embeddings for clinical features, such as heart rate and blood pressure. We use self-supervised training paradigms for language models to learn high-quality clinical feature embeddings, achieving a finer granularity than existing time-step and patient-level representation learning. We visualize the learnt embeddings via unsupervised dimension reduction techniques and observe a high degree of consistency with prior clinical knowledge. We also evaluate the model performance on the MIMIC-III benchmark and demonstrate the effectiveness of using clinical feature embeddings. We publish our code online for replication.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yurong Hu, Manuel Burger, Gunnar Rätsch, Rita Kuznetsova. 2024-02-06. Language Model Training Paradigms for Clinical Feature Embeddings. https://arxiv.org/abs/2311.00768
Cite the original work for its findings. Save a collection to share your selection of sources.