paper-with-me

Papers

IndicBART: A Pre-trained Model for Indic Natural Language Generation

2021-11-16 · ACL ARR November 2021 11 · Anonymous

We study pre-trained sequence-to-sequence model for a specific-language family with a focus on Indic languages. We present IndicBART, a multilingual, sequence-to-sequence pre-trained model focusing on 11 Indic languages and English. IndicBART utilizes the orthographic similarity between Indic scripts to improve transfer learning between similar Indic languages. We evaluate IndicBART on two NLG tasks: Neural Machine Translation (NMT) and extreme summarization. Our experiments on NMT and extreme summarization show that a language family-specific model like IndicBART is competitive with large pre-trained models like mBART50 despite being significantly smaller. It also performs well on very low-resource translation scenarios: languages not included in pre-training or fine-tuning. Script sharing, multilingual training and better utilization of limited model capacity contribute to the good performance of the compact IndicBART model.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Extreme SummarizationMachine TranslationNMTText GenerationTransfer LearningTranslation

Similar Papers 제목 키워드 기반

IndicBART: A Pre-trained Model for Indic Natural Language Generation

2021-09-07 · Findings (ACL) 2022 5 · Raj Dabre, Himani Shrotriya, Anoop Kunchukuttan, Ratish Puduppully 외

In this paper, we study pre-trained sequence-to-sequence models for a group of related languages, with a focus on Indic languages. We present IndicBART, a multilingual, sequence-to-sequence pre-trained model focusing on …

Extreme SummarizationMachine TranslationNMTText Generation+2

Summarizing Indian Languages using Multilingual Transformers based Models

2023-03-29 · Dhaval Taunk, Vasudeva Varma

With the advent of multilingual models like mBART, mT5, IndicBART etc., summarization in low resource Indian languages is getting a lot of attention now a days. But still the number of datasets is low in number. In this …

L3Cube-MahaSum: A Comprehensive Dataset and BART Models for Abstractive Text Summarization in Marathi

2024-10-11 · Pranita Deshmukh, Nikita Kulkarni, Sanhita Kulkarni, Kareena Manghani 외

We present the MahaSUM dataset, a large-scale collection of diverse news articles in Marathi, designed to facilitate the training and evaluation of models for abstractive summarization tasks in Indic languages. The datas…

Abstractive Text SummarizationArticlesText Summarization

Yes-MT's Submission to the Low-Resource Indic Language Translation Shared Task in WMT 2024

2025-12-17 · Yash Bhaskar, Parameswari Krishnamurthy arxiv

This paper presents the systems submitted by the Yes-MT team for the Low-Resource Indic Language Translation Shared Task at WMT 2024 (Pakray et al., 2024), focusing on translating between English and the Assamese, Mizo, …

Implementing Deep Learning-Based Approaches for Article Summarization in Indian Languages

2022-12-12 · Rahul Tangsali, Aabha Pingle, Aditya Vyawahare, Isha Joshi 외

The research on text summarization for low-resource Indian languages has been limited due to the availability of relevant datasets. This paper presents a summary of various deep-learning approaches used for the ILSUM 202…

ArticlesText Summarization