paper-with-me

Papers

Building Odia Shallow Parser

2022-04-19 · Pruthwik Mishra, Dipti Misra Sharma

Shallow parsing is an essential task for many NLP applications like machine translation, summarization, sentiment analysis, aspect identification and many more. Quality annotated corpora is critical for building accurate shallow parsers. Many Indian languages are resource poor with respect to the availability of corpora in general. So, this paper is an attempt towards creating quality corpora for shallow parsers. The contribution of this paper is two folds: creation pos and chunk annotated corpora for Odia and development of baseline systems for pos tagging and chunking in Odia.

📄 PDF Abstract BibTeX arXiv:2204.08960

Code (2)

pruthwik/odia-chunker 공식 구현
Pruthwik/Odia-POS-Tagger

Tasks

ChunkingMachine TranslationPOSPOS TaggingSentiment AnalysisTranslation

Similar Papers 제목 키워드 기반

Universal Dependency Treebank for Odia Language

2022-05-24 · WILDRE (LREC) 2022 6 · Shantipriya Parida, Kalyanamalini Sahoo, Atul Kr. Ojha, Saraswati Sahoo 외

This paper presents the first publicly available treebank of Odia, a morphologically rich low resource Indian language. The treebank contains approx. 1082 tokens (100 sentences) in Odia selected from "Samantar", the larg…

BIG-bench Machine LearningMorphological Analysis

Post-OCR parsing: building simple and robust parser via BIO tagging

2019-09-14 · NeurIPS Workshop Document_Intelligen 2019 12 · Wonseok Hwang, Seonghyeon Kim, Minjoon Seo, Jinyeong Yim 외

Parsing textual information embedded in images is important for various down- stream tasks. However, many previously developed parsers are limited to handling the information presented in one dimensional sequence format.…

Optical Character RecognitionOptical Character Recognition (OCR)

Building a Public Domain Voice Database for Odia

2022-08-16 · WWW '22: Companion Proceedings of the Web Conference 2022 8 · Subhashish Panigrahi

Projects like Mozilla Common Voice were born to address the challenges of unavailability of voice data or the high cost of available data for use in speech technology such as Automatic Speech Recognition (ASR) research a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Building a Llama2-finetuned LLM for Odia Language Utilizing Domain Knowledge Instruction Set

2023-12-19 · Guneet Singh Kohli, Shantipriya Parida, Sambit Sekhar, Samirit Saha 외

Building LLMs for languages other than English is in great demand due to the unavailability and performance of multilingual LLMs, such as understanding the local context. The problem is critical for low-resource language…

OdiEnCorp 2.0: Odia-English Parallel Corpus for Machine Translation

2020-05-01 · LREC 2020 5 · Shantipriya Parida, Satya Ranjan Dash, Ond{\v{r}}ej Bojar, Petr Motlicek 외

The preparation of parallel corpora is a challenging task, particularly for languages that suffer from under-representation in the digital world. In a multi-lingual country like India, the need for such parallel corpora …

Machine TranslationNMTOptical Character RecognitionOptical Character Recognition (OCR)+1