Part-of-Speech Annotation of English-Assamese code-mixed texts: Two Approaches
In this paper, we discuss the development of a part-of-speech tagger for English-Assamese code-mixed texts. We provide a comparison of 2 approaches to annotating code-mixed data {--} a) annotation of the texts from the two languages using monolingual resources from each language and b) annotation of the text through a different resource created specifically for code-mixed data. We present a comparative study of the efforts required in each approach and the final performance of the system. Based on this, we argue that it might be a better approach to develop new technologies using code-mixed data instead of monolingual, {`}clean{'} data, especially for those languages where we do not have significant tools and technologies available till now.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
AsNER - Annotated Dataset and Baseline for Assamese Named Entity recognition
We present the AsNER, a named entity annotation dataset for low resource Assamese language with a baseline Assamese NER model. The dataset contains about 99k tokens comprised of text from the speech of the Prime Minister…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1AsNER -- Annotated Dataset and Baseline for Assamese Named Entity recognition
We present the AsNER, a named entity annotation dataset for low resource Assamese language with a baseline Assamese NER model. The dataset contains about 99k tokens comprised of text from the speech of the Prime Minister…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1Analyzing long-term rhythm variations in Mising and Assamese using frequency domain correlates
The current work explores long-term speech rhythm variations to classify Mising and Assamese, two low-resourced languages from Assam, Northeast India. We study the temporal information of speech rhythm embedded in low-fr…
RhythmDevelopment and Transcription of Assamese Speech Corpus
A balanced speech corpus is the basic need for any speech processing task. In this report we describe our effort on development of Assamese speech corpus. We mainly focused on some issues and challenges faced during deve…
Assamese-English Bilingual Machine Translation
Machine translation is the process of translating text from one language to another. In this paper, Statistical Machine Translation is done on Assamese and English language by taking their respective parallel corpus. A s…
Language ModelingLanguage ModellingMachine TranslationTranslation+1