paper-with-me

Papers

Part-of-Speech Annotation of English-Assamese code-mixed texts: Two Approaches

2018-08-01 · COLING 2018 8 · Ritesh Kumar, Manas Jyoti Bora

In this paper, we discuss the development of a part-of-speech tagger for English-Assamese code-mixed texts. We provide a comparison of 2 approaches to annotating code-mixed data {--} a) annotation of the texts from the two languages using monolingual resources from each language and b) annotation of the text through a different resource created specifically for code-mixed data. We present a comparative study of the efforts required in each approach and the final performance of the system. Based on this, we argue that it might be a better approach to develop new technologies using code-mixed data instead of monolingual, {`}clean{'} data, especially for those languages where we do not have significant tools and technologies available till now.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AsNER - Annotated Dataset and Baseline for Assamese Named Entity recognition

2022-06-01 · LREC 2022 6 · Dhrubajyoti Pathak, Sukumar Nandi, Priyankoo Sarmah

We present the AsNER, a named entity annotation dataset for low resource Assamese language with a baseline Assamese NER model. The dataset contains about 99k tokens comprised of text from the speech of the Prime Minister…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1

AsNER -- Annotated Dataset and Baseline for Assamese Named Entity recognition

2022-07-07 · Dhrubajyoti Pathak, Sukumar Nandi, Priyankoo Sarmah

We present the AsNER, a named entity annotation dataset for low resource Assamese language with a baseline Assamese NER model. The dataset contains about 99k tokens comprised of text from the speech of the Prime Minister…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1

Analyzing long-term rhythm variations in Mising and Assamese using frequency domain correlates

2024-10-26 · Parismita Gogoi, Priyankoo Sarmah, S. R. M. Prasanna

The current work explores long-term speech rhythm variations to classify Mising and Assamese, two low-resourced languages from Assam, Northeast India. We study the temporal information of speech rhythm embedded in low-fr…

Rhythm

Development and Transcription of Assamese Speech Corpus

2013-09-27 · Himangshu Sarma, Navanath Saharia, Utpal Sharma, Smriti Kumar Sinha 외

A balanced speech corpus is the basic need for any speech processing task. In this report we describe our effort on development of Assamese speech corpus. We mainly focused on some issues and challenges faced during deve…

Assamese-English Bilingual Machine Translation

2014-07-08 · Kalyanee Kanchan Baruah, Pranjal Das, Abdul Hannan, Shikhar Kr. Sarma

Machine translation is the process of translating text from one language to another. In this paper, Statistical Machine Translation is done on Assamese and English language by taking their respective parallel corpus. A s…

Language ModelingLanguage ModellingMachine TranslationTranslation+1