paper-with-me

홈 › Papers

Cross Script Hindi English NER Corpus from Wikipedia

2018-10-08 · Mohd Zeeshan Ansari, Tanvir Ahmad, Md Arshad Ali

The text generated on social media platforms is essentially a mixed lingual text. The mixing of language in any form produces considerable amount of difficulty in language processing systems. Moreover, the advancements in language processing research depends upon the availability of standard corpora. The development of mixed lingual Indian Named Entity Recognition (NER) systems are facing obstacles due to unavailability of the standard evaluation corpora. Such corpora may be of mixed lingual nature in which text is written using multiple languages predominantly using a single script only. The motivation of our work is to emphasize the automatic generation such kind of corpora in order to encourage mixed lingual Indian NER. The paper presents the preparation of a Cross Script Hindi-English Corpora from Wikipedia category pages. The corpora is successfully annotated using standard CoNLL-2003 categories of PER, LOC, ORG, and MISC. Its evaluation is carried out on a variety of machine learning algorithms and favorable results are achieved.

📄 PDF Abstract BibTeX arXiv:1810.03430

Code (0)

등록된 구현이 없습니다.

Tasks

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

Similar Papers 제목 키워드 기반

Translating Politeness Across Cultures: Case of Hindi and English

2021-12-03 · Ritesh Kumar, Girish Nath Jha

In this paper, we present a corpus based study of politeness across two languages-English and Hindi. It studies the politeness in a translated parallel corpus of Hindi and English and sees how politeness in a Hindi text …

Machine TranslationTranslation

On the effective transfer of knowledge from English to Hindi Wikipedia

2024-12-07 · Paramita Das, Amartya Roy, Ritabrata Chakraborty, Animesh Mukherjee

Although Wikipedia is the largest multilingual encyclopedia, it remains inherently incomplete. There is a significant disparity in the quality of content between high-resource languages (HRLs, e.g., English) and low-reso…

ArticlesIn-Context Learning

Low Resource Multimodal Neural Machine Translation of English-Hindi in News Domain

2021-09-01 · MMTLRL (RANLP) 2021 9 · Loitongbam Sanayai Meetei, Thoudam Doren Singh, Sivaji Bandyopadhyay

Incorporating multiple input modalities in a machine translation (MT) system is gaining popularity among MT researchers. Unlike the publicly available dataset for Multimodal Machine Translation (MMT) tasks, where the cap…

Machine TranslationMultimodal Machine TranslationNMTTranslation

A Topic-Aligned Multilingual Corpus of Wikipedia Articles for Studying Information Asymmetry in Low Resource Languages

2020-05-01 · LREC 2020 5 · Dwaipayan Roy, Sumit Bhatia, Prateek Jain

Wikipedia is the largest web-based open encyclopedia covering more than three hundred languages. However, different language editions of Wikipedia differ significantly in terms of their information coverage. We present a…

Articles

Generating Inflectional Errors for Grammatical Error Correction in Hindi

2020-12-01 · Asian Chapter of the Association for Computational Linguistics 2020 · Ankur Sonawane, Sujeet Kumar Vishwakarma, Bhavana Srivastava, Anil Kumar Singh

Automated grammatical error correction has been explored as an important research problem within NLP, with the majority of the work being done on English and similar resource-rich languages. Grammar correction using neur…

Grammatical Error Correction