paper-with-me

Papers

Revisiting Low Resource Status of Indian Languages in Machine Translation

2020-08-11 · Jerin Philip, Shashank Siripragada, Vinay P. Namboodiri, C. V. Jawahar

Indian language machine translation performance is hampered due to the lack of large scale multi-lingual sentence aligned corpora and robust benchmarks. Through this paper, we provide and analyse an automated framework to obtain such a corpus for Indian language neural machine translation (NMT) systems. Our pipeline consists of a baseline NMT system, a retrieval module, and an alignment module that is used to work with publicly available websites such as press releases by the government. The main contribution towards this effort is to obtain an incremental method that uses the above pipeline to iteratively improve the size of the corpus as well as improve each of the components of our system. Through our work, we also evaluate the design choices such as the choice of pivoting language and the effect of iterative incremental increase in corpus size. Our work in addition to providing an automated framework also results in generating a relatively larger corpus as compared to existing corpora that are available for Indian languages. This corpus helps us obtain substantially improved results on the publicly available WAT evaluation benchmark and other standard evaluation benchmarks.

📄 PDF Abstract BibTeX arXiv:2008.04860

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationNMTRetrievalSentenceTranslation

Similar Papers 제목 키워드 기반

Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages

2025-10-08 · Amir Hossein Yari, Kalmit Kulkarni, Ahmad Raza Khan, Fajri Koto arxiv

While automatic metrics drive progress in Machine Translation (MT) and Text Summarization (TS), existing metrics have been developed and validated almost exclusively for English and other high-resource languages. This na…

Machine TranslationText Summarization

First Attempt at Building Parallel Corpora for Machine Translation of Northeast India's Very Low-Resource Languages

2023-12-08 · Atnafu Lambebo Tonja, Melkamu Mersha, Ananya Kalita, Olga Kolesnikova 외

This paper presents the creation of initial bilingual corpora for thirteen very low-resource languages of India, all from Northeast India. It also presents the results of initial translation efforts in these languages. I…

Machine TranslationTranslation

Machine Translation Advancements of Low-Resource Indian Languages by Transfer Learning

2024-09-24 · Bin Wei, Jiawei Zhen, Zongyao Li, Zhanglin Wu 외

This paper introduces the submission by Huawei Translation Center (HW-TSC) to the WMT24 Indian Languages Machine Translation (MT) Shared Task. To develop a reliable machine translation system for low-resource Indian lang…

Machine TranslationTransfer LearningTranslation

IndicIRSuite: Multilingual Dataset and Neural Information Models for Indian Languages

2023-12-15 · Saiful Haq, Ashutosh Sharma, Pushpak Bhattacharyya

In this paper, we introduce Neural Information Retrieval resources for 11 widely spoken Indian Languages (Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, and Telugu) from two major…

Information RetrievalMachine TranslationRetrieval

Exploring Pair-Wise NMT for Indian Languages

2020-12-10 · ICON 2020 12 · Kartheek Akella, Sai Himal Allu, Sridhar Suresh Ragupathi, Aman Singhal 외

In this paper, we address the task of improving pair-wise machine translation for specific low resource Indian languages. Multilingual NMT models have demonstrated a reasonable amount of effectiveness on resource-poor la…

Machine TranslationNMTTranslation