paper-with-me

홈 › Papers

Assessing the Corpus Size vs. Similarity Trade-off for Word Embeddings in Clinical NLP

2016-12-01 · WS 2016 12 · Kirk Roberts

The proliferation of deep learning methods in natural language processing (NLP) and the large amounts of data they often require stands in stark contrast to the relatively data-poor clinical NLP domain. In particular, large text corpora are necessary to build high-quality word embeddings, yet often large corpora that are suitably representative of the target clinical data are unavailable. This forces a choice between building embeddings from small clinical corpora and less representative, larger corpora. This paper explores this trade-off, as well as intermediate compromise solutions. Two standard clinical NLP tasks (the i2b2 2010 concept and assertion tasks) are evaluated with commonly used deep learning models (recurrent neural networks and convolutional neural networks) using a set of six corpora ranging from the target i2b2 data to large open-domain datasets. While combinations of corpora are generally found to work best, the single-best corpus is generally task-dependent.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningWord Embeddings

Similar Papers 제목 키워드 기반

Statistical Uncertainty in Word Embeddings: GloVe-V

2024-06-18 · Andrea Vallebueno, Cassandra Handan-Nader, Christopher D. Manning, Daniel E. Ho

Static word embeddings are ubiquitous in computational social science applications and contribute to practical decision-making in a variety of fields including law and healthcare. However, assessing the statistical uncer…

Decision MakingModel SelectionWord Embeddings

ParCourE: A Parallel Corpus Explorer for a Massively Multilingual Corpus

2021-07-14 · ACL 2021 5 · Ayyoob Imani, Masoud Jalili Sabet, Philipp Dufter, Michael Cysouw 외

With more than 7000 languages worldwide, multilingual natural language processing (NLP) is essential both from an academic and commercial perspective. Researching typological properties of languages is fundamental for pr…

Multilingual NLPTransfer Learning

Size vs. Structure in Training Corpora for Word Embedding Models: Araneum Russicum Maximum and Russian National Corpus

2018-01-19 · Andrey Kutuzov, Maria Kunilovskaya

In this paper, we present a distributional word embedding model trained on one of the largest available Russian corpora: Araneum Russicum Maximum (over 10 billion words crawled from the web). We compare this model to the…

Semantic SimilaritySemantic Textual Similarity

Using Word Familiarities and Word Associations to Measure Corpus Representativeness

2014-05-01 · LREC 2014 5 · Reinhard Rapp

The definition of corpus representativeness used here assumes that a representative corpus should reflect as well as possible the average language use a native speaker encounters in everyday life over a longer period of …

Language Acquisition

GNEG: Graph-Based Negative Sampling for word2vec

2018-07-01 · ACL 2018 7 · Zheng Zhang, Pierre Zweigenbaum

Negative sampling is an important component in word2vec for distributed word representation learning. We hypothesize that taking into account global, corpus-level information and generating a different noise distribution…

Language ModelingLanguage ModellingRepresentation LearningWord Similarity