paper-with-me

홈 › Papers

A Multi-way Parallel Named Entity Annotated Corpus for English, Tamil and Sinhala

2024-12-03 · Surangika Ranathunga, Asanka Ranasinghea, Janaka Shamala, Ayodya Dandeniyaa, Rashmi Galappaththia, Malithi Samaraweeraa

This paper presents a multi-way parallel English-Tamil-Sinhala corpus annotated with Named Entities (NEs), where Sinhala and Tamil are low-resource languages. Using pre-trained multilingual Language Models (mLMs), we establish new benchmark Named Entity Recognition (NER) results on this dataset for Sinhala and Tamil. We also carry out a detailed investigation on the NER capabilities of different types of mLMs. Finally, we demonstrate the utility of our NER system on a low-resource Neural Machine Translation (NMT) task. Our dataset is publicly released: https://github.com/suralk/multiNER.

📄 PDF Abstract BibTeX arXiv:2412.02056

Code (1)

suralk/multiner 공식 구현

Tasks

Low Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine Translationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NERNMTTranslation

Similar Papers 제목 키워드 기반

Building a Bilingual Vietnamese-French Named Entity Annotated Corpus through Cross-Linguistic Projection

2015-06-01 · JEPTALNRECITAL 2015 6 · Ngoc Tan Le, Fatiha Sadat

The creation of high-quality named entity annotated resources is time-consuming and an expensive process. Most of the gold standard corpora are available for English but not for less-resourced languages such as Vietnames…

UNER: Universal Named-Entity RecognitionFramework

2020-10-23 · Diego Alves, Tin Kuculo, Gabriel Amaral, Gaurish Thakkar 외

We introduce the Universal Named-Entity Recognition (UNER)framework, a 4-level classification hierarchy, and the methodology that isbeing adopted to create the first multilingual UNER corpus: the SETimesparallel corpus a…

Knowledge Graphsnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

QTLeap WSD/NED Corpora: Semantic Annotation of Parallel Corpora in Six Languages

2016-05-01 · LREC 2016 5 · Arantxa Otegi, Nora Aranberri, Antonio Branco, Jan Haji{\v{c}} 외

This work presents parallel corpora automatically annotated with several NLP tools, including lemma and part-of-speech tagging, named-entity recognition and classification, named-entity disambiguation, word-sense disambi…

Cross-Lingual TransferEntity DisambiguationGeneral ClassificationLEMMA+7

Constructing Uyghur Name Entity Recognition System using Neural Machine Translation Tag Projection

2020-10-01 · CCL 2020 10 · Anwar Azmat, Li Xiao, Yang Yating, Dong Rui 외

Although named entity recognition achieved great success by introducing the neural networks, it is challenging to apply these models to low resource languages including Uyghur while it depends on a large amount of annota…

Machine Translationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+4

The SETimes.HR Linguistically Annotated Corpus of Croatian

2014-05-01 · LREC 2014 5 · {\v{Z}}eljko Agi{\'c}, Nikola Ljube{\v{s}}i{\'c}

We present SETimes.HR ― the first linguistically annotated corpus of Croatian that is freely available for all purposes. The corpus is built on top of the SETimes parallel corpus of nine Southeast European languages an…

AllBoundary DetectionDependency ParsingLemmatization+4