paper-with-me

Papers

Challenges in Developing LRs for Non-Scheduled Languages: A Case of Magahi

2021-11-30 · Ritesh Kumar

Magahi is an Indo-Aryan Language, spoken mainly in the Eastern parts of India. Despite having a significant number of speakers, there has been virtually no language resource (LR) or language technology (LT) developed for the language, mainly because of its status as a non-scheduled language. The present paper describes an attempt to develop an annotated corpus of Magahi. The data is mainly taken from a couple of blogs in Magahi, some collection of stories in Magahi and the recordings of conversation in Magahi and it is annotated at the POS level using BIS tagset.

📄 PDF Abstract BibTeX arXiv:2111.15322

Code (0)

등록된 구현이 없습니다.

Tasks

POS

Similar Papers 제목 키워드 기반

Developing Universal Dependency Treebanks for Magahi and Braj

2022-04-26 · Mohit Raj, Shyam Ratan, Deepak Alok, Ritesh Kumar 외

In this paper, we discuss the development of treebanks for two low-resourced Indian languages - Magahi and Braj based on the Universal Dependencies framework. The Magahi treebank contains 945 sentences and Braj treebank …

Developing Universal Dependencies Treebanks for Magahi and Braj

2021-12-01 · PAIL (ICON) 2021 12 · Mohit Raj, Shyam Ratan, Deepak Alok, Ritesh Kumar 외

In this paper, we discuss the development of treebanks for two low-resourced Indian languages - Magahi and Braj - based on the Universal Dependencies framework. The Magahi treebank contains 945 sentences and Braj treeban…

Bengali and Magahi PUD Treebank and Parser

2022-06-01 · WILDRE (LREC) 2022 6 · Pritha Majumdar, Deepak Alok, Akanksha Bansal, Atul Kr. Ojha 외

This paper presents the development of the Parallel Universal Dependency (PUD) Treebank for two Indo-Aryan languages: Bengali and Magahi. A treebank of 1,000 sentences has been created using a parallel corpus of English …

Development of a Dataset and a Deep Learning Baseline Named Entity Recognizer for Three Low Resource Languages: Bhojpuri, Maithili and Magahi

2020-09-14 · Rajesh Kumar Mundotiya, Shantanu Kumar, Ajeet kumar, Umesh Chandra Chaudhary 외

In Natural Language Processing (NLP) pipelines, Named Entity Recognition (NER) is one of the preliminary problems, which marks proper nouns and other named entities such as Location, Person, Organization, Disease etc. Su…

Machine Translationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+2

Automatic Language Identification System for Hindi and Magahi

2018-04-13 · Priya Rani, Atul Kr. Ojha, Girish Nath Jha

Language identification has become a prerequisite for all kinds of automated text processing systems. In this paper, we present a rule-based language identifier tool for two closely related Indo-Aryan languages: Hindi an…

Language Identification