Developing Healthcare Language Model Embedding Spaces
Pre-trained Large Language Models (LLMs) often struggle on out-of-domain datasets like healthcare focused text. We explore specialized pre-training to adapt smaller LLMs to different healthcare datasets. Three methods are assessed: traditional masked language modeling, Deep Contrastive Learning for Unsupervised Textual Representations (DeCLUTR), and a novel pre-training objective utilizing metadata categories from the healthcare settings. These schemes are evaluated on downstream document classification tasks for each dataset, with additional analysis of the resultant embedding spaces. Contrastively trained models outperform other approaches on the classification tasks, delivering strong performance from limited labeled data and with fewer model parameter updates required. While metadata-based pre-training does not further improve classifications across the datasets, it yields interesting embedding cluster separability. All domain adapted LLMs outperform their publicly available general base LLM, validating the importance of domain-specialization. This research illustrates efficient approaches to instill healthcare competency in compact LLMs even under tight computational budgets, an essential capability for responsible and sustainable deployment in local healthcare settings. We provide pre-training guidelines for specialized healthcare LLMs, motivate continued inquiry into contrastive objectives, and demonstrates adaptation techniques to align small LLMs with privacy-sensitive medical tasks.
Code (0)
등록된 구현이 없습니다.
Tasks
Contrastive LearningDocument ClassificationLanguage ModelingLanguage ModellingMasked Language ModelingmodelMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Unsupervised Cross-Modal Alignment of Speech and Text Embedding Spaces
Recent research has shown that word embedding spaces learned from text corpora of different languages can be aligned without any parallel data supervision. Inspired by the success in unsupervised cross-lingual word embed…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Cross-Lingual Word Embeddingscross-modal alignment+6Snomed2Vec: Random Walk and Poincaré Embeddings of a Clinical Knowledge Base for Healthcare Analytics
Representation learning methods that transform encoded data (e.g., diagnosis and drug codes) into continuous vector spaces (i.e., vector embeddings) are critical for the application of deep learning in healthcare. Initia…
Clinical KnowledgeLink PredictionNode ClassificationRepresentation LearningA Comparative Study on Structural and Semantic Properties of Sentence Embeddings
Sentence embeddings encode natural language sentences as low-dimensional dense vectors. A great deal of effort has been put into using sentence embeddings to improve several important natural language processing tasks. R…
Knowledge GraphsRelationRelation ExtractionSentence+3AURORA: Contextual Orthogonalization for Geometric Representation Learning in Healthcare Foundation Models
Recent healthcare foundation models have achieved strong predictive performance through large scale self supervised learning, yet their latent representations frequently entangle physiologic severity, intervention intens…
Representation LearningLearning in Hilbert vs. Banach Spaces: A Measure Embedding Viewpoint
The goal of this paper is to investigate the advantages and disadvantages of learning in Banach spaces over Hilbert spaces. While many works have been carried out in generalizing Hilbert methods to Banach spaces, in this…