Building Odia Shallow Parser
Shallow parsing is an essential task for many NLP applications like machine translation, summarization, sentiment analysis, aspect identification and many more. Quality annotated corpora is critical for building accurate shallow parsers. Many Indian languages are resource poor with respect to the availability of corpora in general. So, this paper is an attempt towards creating quality corpora for shallow parsers. The contribution of this paper is two folds: creation pos and chunk annotated corpora for Odia and development of baseline systems for pos tagging and chunking in Odia.
Code (2)
Tasks
ChunkingMachine TranslationPOSPOS TaggingSentiment AnalysisTranslationSimilar Papers 제목 키워드 기반
Universal Dependency Treebank for Odia Language
This paper presents the first publicly available treebank of Odia, a morphologically rich low resource Indian language. The treebank contains approx. 1082 tokens (100 sentences) in Odia selected from "Samantar", the larg…
BIG-bench Machine LearningMorphological AnalysisPost-OCR parsing: building simple and robust parser via BIO tagging
Parsing textual information embedded in images is important for various down- stream tasks. However, many previously developed parsers are limited to handling the information presented in one dimensional sequence format.…
Optical Character RecognitionOptical Character Recognition (OCR)Building a Public Domain Voice Database for Odia
Projects like Mozilla Common Voice were born to address the challenges of unavailability of voice data or the high cost of available data for use in speech technology such as Automatic Speech Recognition (ASR) research a…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionBuilding a Llama2-finetuned LLM for Odia Language Utilizing Domain Knowledge Instruction Set
Building LLMs for languages other than English is in great demand due to the unavailability and performance of multilingual LLMs, such as understanding the local context. The problem is critical for low-resource language…
OdiEnCorp 2.0: Odia-English Parallel Corpus for Machine Translation
The preparation of parallel corpora is a challenging task, particularly for languages that suffer from under-representation in the digital world. In a multi-lingual country like India, the need for such parallel corpora …
Machine TranslationNMTOptical Character RecognitionOptical Character Recognition (OCR)+1