A Comparison of Word-based and Context-based Representations for Classification Problems in Health Informatics
Distributed representations of text can be used as features when training a statistical classifier. These representations may be created as a composition of word vectors or as context-based sentence vectors. We compare the two kinds of representations (word versus context) for three classification problems: influenza infection classification, drug usage classification and personal health mention classification. For statistical classifiers trained for each of these problems, context-based representations based on ELMo, Universal Sentence Encoder, Neural-Net Language Model and FLAIR are better than Word2Vec, GloVe and the two adapted using the MESH ontology. There is an improvement of 2-4% in the accuracy when these context-based representations are used instead of word-based representations.
Code (0)
등록된 구현이 없습니다.
Tasks
ClassificationGeneral ClassificationLanguage ModelingLanguage ModellingSentenceMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
A Comparison of Context-sensitive Models for Lexical Substitution
Word embedding representations provide good estimates of word meaning and give state-of-the art performance in semantic tasks. Embedding approaches differ as to whether and how they account for the context surrounding a …
Word EmbeddingsComparison of Turkish Word Representations Trained on Different Morphological Forms
Increased popularity of different text representations has also brought many improvements in Natural Language Processing (NLP) tasks. Without need of supervised data, embeddings trained on large corpora provide us meanin…
Language ModelingLanguage ModellingLEMMAtext-classification+1Comparing Text Representations: A Theory-Driven Approach
Much of the progress in contemporary NLP has come from learning representations, such as masked language model (MLM) contextual embeddings, that turn challenging problems into simple classification tasks. But how do we q…
Language ModelingLanguage ModellingLearning TheoryNatural Language InferenceComparative Analysis of Text Classification Approaches in Electronic Health Records
Text classification tasks which aim at harvesting and/or organizing information from electronic health records are pivotal to support clinical and translational research. However these present specific challenges compare…
ClassificationGeneral Classificationtext-classificationText ClassificationWord Usage Similarity Estimation with Sentence Representations and Automatic Substitutes
Usage similarity estimation addresses the semantic proximity of word instances in different contexts. We apply contextualized (ELMo and BERT) word and sentence embeddings to this task, and propose supervised models that …
SentenceSentence Embeddings