Resource-Size matters: Improving Neural Named Entity Recognition with Optimized Large Corpora
This study improves the performance of neural named entity recognition by a margin of up to 11% in F-score on the example of a low-resource language like German, thereby outperforming existing baselines and establishing a new state-of-the-art on each single open-source dataset. Rather than designing deeper and wider hybrid neural architectures, we gather all available resources and perform a detailed optimization and grammar-dependent morphological processing consisting of lemmatization and part-of-speech tagging prior to exposing the raw data to any training process. We test our approach in a threefold monolingual experimental setup of a) single, b) joint, and c) optimized training and shed light on the dependency of downstream-tasks on the size of corpora used to compute word embeddings.
Code (1)
Tasks
Lemmatizationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Part-Of-Speech TaggingWord EmbeddingsSimilar Papers 제목 키워드 기반
What Matters for Neural Cross-Lingual Named Entity Recognition: An Empirical Analysis
Building named entity recognition (NER) models for languages that do not have much training data is a challenging task. While recent work has shown promising results on cross-lingual transfer from high-resource languages…
Cross-Lingual NERCross-Lingual Transfernamed-entity-recognitionNamed Entity Recognition+3Distant Supervision and Noisy Label Learning for Low Resource Named Entity Recognition: A Study on Hausa and Yorùbá
The lack of labeled training data has limited the development of natural language processing tools, such as named entity recognition, for many languages spoken in developing countries. Techniques such as distant and weak…
Low Resource Named Entity Recognitionnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1Soft Gazetteers for Low-Resource Named Entity Recognition
Traditional named entity recognition models use gazetteers (lists of entities) as features to improve performance. Although modern neural network models do not require such hand-crafted features for strong performance, r…
Cross-Lingual Entity LinkingEntity LinkingLow Resource Named Entity Recognitionnamed-entity-recognition+2Broad Twitter Corpus: A Diverse Named Entity Recognition Resource
One of the main obstacles, hampering method development and comparative evaluation of named entity recognition in social media, is the lack of a sizeable, diverse, high quality annotated corpus, analogous to the CoNLL{'}…
Diversitynamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Using Domain Knowledge for Low Resource Named Entity Recognition
In recent years, named entity recognition has always been a popular research in the field of natural language processing, while traditional deep learning methods require a large amount of labeled data for model training,…
Chinese Named Entity RecognitionLow Resource Named Entity Recognitionnamed-entity-recognitionNamed Entity Recognition+2