arXiv · 2311.00768
Language Model Training Paradigms for Clinical Feature Embeddings
Abstract
In research areas with scarce data, representation learning plays a significant role. This work aims to enhance representation learning for clinical time series by deriving universal embeddings for clinical features, such as heart rate and blood pressure. We use self-supervised training paradigms for language models to learn high-quality clinical feature embeddings, achieving a finer granularity than existing time-step and patient-level representation learning. We visualize the learnt embeddings via unsupervised dimension reduction techniques and observe a high degree of consistency with prior clinical knowledge. We also evaluate the model performance on the MIMIC-III benchmark and demonstrate the effectiveness of using clinical feature embeddings. We publish our code online for replication.
Explore related subjects
Keep this discovery
Yurong Hu, Manuel Burger, Gunnar Rätsch, Rita Kuznetsova. 2023-11-01. Language Model Training Paradigms for Clinical Feature Embeddings. https://arxiv.org/abs/2311.00768
Cite the original work for its findings. Save a collection to share your selection of sources.