Judicious Selection of Training Data in Assisting Language for Multilingual Neural NER
Multilingual learning for Neural Named Entity Recognition (NNER) involves jointly training a neural network for multiple languages. Typically, the goal is improving the NER performance of one of the languages (the primary language) using the other assisting languages. We show that the divergence in the tag distributions of the common named entities between the primary and assisting languages can reduce the effectiveness of multilingual learning. To alleviate this problem, we propose a metric based on symmetric KL divergence to filter out the highly divergent training instances in the assisting language. We empirically show that our data selection strategy improves NER performance in many languages, including those with very limited training data.
Code (1)
Tasks
Domain AdaptationMachine Translationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NERTAGSimilar Papers 제목 키워드 기반
The impact of responding to patient messages with large language model assistance
Documentation burden is a major contributor to clinician burnout, which is rising nationally and is an urgent threat to our ability to care for patients. Artificial intelligence (AI) chatbots, such as ChatGPT, could redu…
ChatbotDecision MakingLanguage ModelingLanguage Modelling+1Pre-training via Leveraging Assisting Languages and Data Selection for Neural Machine Translation
Sequence-to-sequence (S2S) pre-training using large monolingual data is known to improve performance for various S2S NLP tasks in low-resource settings. However, large monolingual corpora might not always be available fo…
Machine TranslationNMTTranslationNARMADA: Need and Available Resource Managing Assistant for Disasters and Adversities
Although a lot of research has been done on utilising Online Social Media during disasters, there exists no system for a specific task that is critical in a post-disaster scenario -- identifying resource-needs and resour…
Information RetrievalManagementRetrievalMultilingual training set selection for ASR in under-resourced Malian languages
We present first speech recognition systems for the two severely under-resourced Malian languages Bambara and Maasina Fulfulde. These systems will be used by the United Nations as part of a monitoring system to inform an…
Humanitarianspeech-recognitionSpeech RecognitionAction Shapley: A Training Data Selection Metric for World Model in Reinforcement Learning
Numerous offline and model-based reinforcement learning systems incorporate world models to emulate the inherent environments. A world model is particularly important in scenarios where direct interactions with the real …
Computational EfficiencyReinforcement Learning