paper-with-me

Papers

Bootstrapping Method for Developing Part-of-Speech Tagged Corpus in Low Resource Languages Tagset - A Focus on an African Igbo

2019-03-12 · Onyenwe Ikechukwu E, Onyedinma Ebele G, Aniegwu Godwin E, Ezeani Ignatius M

Most languages, especially in Africa, have fewer or no established part-of-speech (POS) tagged corpus. However, POS tagged corpus is essential for natural language processing (NLP) to support advanced researches such as machine translation, speech recognition, etc. Even in cases where there is no POS tagged corpus, there are some languages for which parallel texts are available online. The task of POS tagging a new language corpus with a new tagset usually face a bootstrapping problem at the initial stages of the annotation process. The unavailability of automatic taggers to help the human annotator makes the annotation process to appear infeasible to quickly produce adequate amounts of POS tagged corpus for advanced NLP research and training the taggers. In this paper, we demonstrate the efficacy of a POS annotation method that employed the services of two automatic approaches to assist POS tagged corpus creation for a novel language in NLP. The two approaches are cross-lingual and monolingual POS tags projection. We used cross-lingual to automatically create an initial 'errorful' tagged corpus for a target language via word-alignment. The resources for creating this are derived from a source language rich in NLP resources. A monolingual method is applied to clean the induce noise via an alignment process and to transform the source language tags to the target language tags. We used English and Igbo as our case study. This is possible because there are parallel texts that exist between English and Igbo, and the source language English has available NLP resources. The results of the experiment show a steady improvement in accuracy and rate of tags transformation with score ranges of 6.13% to 83.79% and 8.67% to 98.37% respectively. The rate of tags transformation evaluates the rate at which source language tags are translated to target language tags.

📄 PDF Abstract BibTeX arXiv:1903.05225

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationPOSPOS Taggingspeech-recognitionSpeech RecognitionWord Alignment

Similar Papers 제목 키워드 기반

Part of Speech Tagging (POST) of a Low-resource Language using another Language (Developing a POS-Tagged Lexicon for Kurdish (Sorani) using a Tagged Persian (Farsi) Corpus)

2022-01-30 · Hossein Hassani

Tagged corpora play a crucial role in a wide range of Natural Language Processing. The Part of Speech Tagging (POST) is essential in developing tagged corpora. It is time-and-effort-consuming and costly, and therefore, i…

Part-Of-Speech TaggingPOS

Creating a morphological and syntactic tagged corpus for the Uzbek language

2022-10-27 · Maksud Sharipov, Jamolbek Mattiev, Jasur Sobirov, Rustam Baltayev

Nowadays, creation of the tagged corpora is becoming one of the most important tasks of Natural Language Processing (NLP). There are not enough tagged corpora to build machine learning models for the low-resource Uzbek l…

POS

Open-Source Boundary-Annotated Corpus for Arabic Speech and Language Processing

2012-05-01 · LREC 2012 5 · Claire Brierley, Majdi Sawalha, Eric Atwell

A boundary-annotated and part-of-speech tagged corpus is a prerequisite for developing phrase break classifiers. Boundary annotations in English speech corpora are descriptive, delimiting intonation units perceived by th…

ChunkingDescriptiveSpeech SynthesisText-To-Speech Synthesis

A New Annotation Scheme for the Sejong Part-of-speech Tagged Corpus

2019-08-01 · WS 2019 8 · Jungyeul Park, Francis Tyers

In this paper we present a new annotation scheme for the Sejong part-of-speech tagged corpus based on Universal Dependencies style annotation. By using a new annotation scheme, we can produce Sejong-style morphological a…

Morphological Analysisnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

A Tagged Corpus and a Tagger for Urdu

2014-05-01 · LREC 2014 5 · Bushra Jawaid, Amir Kamran, Ond{\v{r}}ej Bojar

In this paper, we describe a release of a sizeable monolingual Urdu corpus automatically tagged with part-of-speech tags. We extend the work of Jawaid and Bojar (2012) who use three different taggers and then apply a vot…