paper-with-me

홈 › Papers

Annotated Guidelines and Building Reference Corpus for Myanmar-English Word Alignment

2019-09-25 · Nway Nway Han, Aye Thida

Reference corpus for word alignment is an important resource for developing and evaluating word alignment methods. For Myanmar-English language pairs, there is no reference corpus to evaluate the word alignment tasks. Therefore, we created the guidelines for Myanmar-English word alignment annotation between two languages over contrastive learning and built the Myanmar-English reference corpus consisting of verified alignments from Myanmar ALT of the Asian Language Treebank (ALT). This reference corpus contains confident labels sure (S) and possible (P) for word alignments which are used to test for the purpose of evaluation of the word alignments tasks. We discuss the most linking ambiguities to define consistent and systematic instructions to align manual words. We evaluated the results of annotators agreement using our reference corpus in terms of alignment error rate (AER) in word alignment tasks and discuss the words relationships in terms of BLEU scores.

📄 PDF Abstract BibTeX arXiv:1909.11288

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningWord Alignment

Similar Papers 제목 키워드 기반

Syllable-based Neural Named Entity Recognition for Myanmar Language

2019-03-12 · Hsu Myat Mo, Khin Mar Soe

Named Entity Recognition (NER) for Myanmar Language is essential to Myanmar natural language processing research work. In this work, NER for Myanmar language is treated as a sequence tagging problem and the effectiveness…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

Building an Arabic Machine Translation Post-Edited Corpus: Guidelines and Annotation

2016-05-01 · LREC 2016 5 · Wajdi Zaghouani, Nizar Habash, Ossama Obeid, Behrang Mohit 외

We present our guidelines and annotation procedure to create a human corrected machine translated post-edited corpus for the Modern Standard Arabic. Our overarching goal is to use the annotated corpus to develop automati…

ArticlesMachine TranslationTranslation

Building The Sense-Tagged Multilingual Parallel Corpus

2014-05-01 · LREC 2014 5 · Shan Wang, Francis Bond

Sense-annotated parallel corpora play a crucial role in natural language processing. This paper introduces our progress in creating such a corpus for Asian languages using English as a pivot, which is the first such corp…

Automatic Myanmar Image Captioning using CNN and LSTM-Based Language Model

2020-05-01 · LREC 2020 5 · San Pa Pa Aung, Win Pa Pa, Tin Lay Nwe

An image captioning system involves modules on computer vision as well as natural language processing. Computer vision module is for detecting salient objects or extracting features of images and Natural Language Process…

Image CaptioningLanguage ModelingLanguage Modelling

Building a Manually Annotated Hungarian Coreference Corpus: Workflow and Tools

2022-10-01 · COLING (CRAC) 2022 10 · Noémi Vadász

This paper presents the complete workflow of building a manually annotated Hungarian corpus, KorKor, with particular reference to anaphora and coreference annotation. All linguistic annotation layers were corrected manua…