paper-with-me

Papers

A Multilingual Parallel Corpora Collection Effort for Indian Languages

2020-07-15 · LREC 2020 5 · Shashank Siripragada, Jerin Philip, Vinay P. Namboodiri, C. V. Jawahar

We present sentence aligned parallel corpora across 10 Indian Languages - Hindi, Telugu, Tamil, Malayalam, Gujarati, Urdu, Bengali, Oriya, Marathi, Punjabi, and English - many of which are categorized as low resource. The corpora are compiled from online sources which have content shared across languages. The corpora presented significantly extends present resources that are either not large enough or are restricted to a specific domain (such as health). We also provide a separate test corpus compiled from an independent online source that can be independently used for validating the performance in 10 Indian languages. Alongside, we report on the methods of constructing such corpora using tools enabled by recent advances in machine translation and cross-lingual retrieval using deep neural network based methods.

📄 PDF Abstract BibTeX arXiv:2007.07691

Code (2)

jerinphilip/fairseq-ilmt pytorch
shashanksiripragada/pib-crawl

Tasks

Machine TranslationRetrievalSentenceTranslation

Similar Papers 제목 키워드 기반

Efficient Neural Machine Translation for Low-Resource Languages via Exploiting Related Languages

2020-07-01 · ACL 2020 6 · Vikrant Goyal, Sourav Kumar, Dipti Misra Sharma

A large percentage of the world{'}s population speaks a language of the Indian subcontinent, comprising languages from both Indo-Aryan (e.g. Hindi, Punjabi, Gujarati, etc.) and Dravidian (e.g. Tamil, Telugu, Malayalam, e…

Machine TranslationNMTTransfer LearningTranslation+1

A Large-scale Evaluation of Neural Machine Transliteration for Indic Languages

2021-04-01 · EACL 2021 2 · Anoop Kunchukuttan, Siddharth Jain, Rahul Kejriwal

We take up the task of large-scale evaluation of neural machine transliteration between English and Indic languages, with a focus on multilingual transliteration to utilize orthographic similarity between Indian language…

TranslationTransliteration

First Attempt at Building Parallel Corpora for Machine Translation of Northeast India's Very Low-Resource Languages

2023-12-08 · Atnafu Lambebo Tonja, Melkamu Mersha, Ananya Kalita, Olga Kolesnikova 외

This paper presents the creation of initial bilingual corpora for thirteen very low-resource languages of India, all from Northeast India. It also presents the results of initial translation efforts in these languages. I…

Machine TranslationTranslation

Statistical Analysis of Multilingual Text Corpus and Development of Language Models

2014-05-01 · LREC 2014 5 · Shyam Sundar Agrawal, {Abhimanue}, shweta bansal, Minakshi Mahajan

This paper presents two studies, first a statistical analysis for three languages i.e. Hindi, Punjabi and Nepali and the other, development of language models for three Indian languages i.e. Indian English, Punjabi and N…

Language IdentificationLanguage ModellingSpeech Language Identification

CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems

2025-09-24 · Soham Bhattacharjee, Mukund K Roy, Yathish Poojary, Bhargav Dave 외 arxiv

India's linguistic landscape is one of the most diverse in the world, comprising over 120 major languages and approximately 1,600 additional languages, with 22 officially recognized as scheduled languages in the Indian C…

Machine TranslationTransfer LearningDomain Adaptation