paper-with-me

홈 › Papers

Mukhyansh: A Headline Generation Dataset for Indic Languages

2023-11-29 · Lokesh Madasu, Gopichand Kanumolu, Nirmal Surange, Manish Shrivastava

The task of headline generation within the realm of Natural Language Processing (NLP) holds immense significance, as it strives to distill the true essence of textual content into concise and attention-grabbing summaries. While noteworthy progress has been made in headline generation for widely spoken languages like English, there persist numerous challenges when it comes to generating headlines in low-resource languages, such as the rich and diverse Indian languages. A prominent obstacle that specifically hinders headline generation in Indian languages is the scarcity of high-quality annotated data. To address this crucial gap, we proudly present Mukhyansh, an extensive multilingual dataset, tailored for Indian language headline generation. Comprising an impressive collection of over 3.39 million article-headline pairs, Mukhyansh spans across eight prominent Indian languages, namely Telugu, Tamil, Kannada, Malayalam, Hindi, Bengali, Marathi, and Gujarati. We present a comprehensive evaluation of several state-of-the-art baseline models. Additionally, through an empirical analysis of existing works, we demonstrate that Mukhyansh outperforms all other models, achieving an impressive average ROUGE-L score of 31.43 across all 8 languages.

📄 PDF Abstract BibTeX arXiv:2311.17743

Code (1)

ltrc/mukhyansh 공식 구현 jax

Tasks

Headline Generation

Similar Papers 제목 키워드 기반

L3Cube-IndicHeadline-ID: A Dataset for Headline Identification and Semantic Evaluation in Low-Resource Indian Languages

2025-09-02 · Nishant Tanksale, Tanmay Kokate, Darshan Gohad, Sarvadnyaa Barate 외 arxiv

Semantic evaluation in low-resource languages remains a major challenge in NLP. While sentence transformers have shown strong performance in high-resource settings, their effectiveness in Indic languages is underexplored…

Question Answering

Vārta: A Large-Scale Headline-Generation Dataset for Indic Languages

2023-05-10 · Rahul Aralikatte, Ziling Cheng, Sumanth Doddapaneni, Jackie Chi Kit Cheung

We present V\=arta, a large-scale multilingual dataset for headline generation in Indic languages. This dataset includes 41.8 million news articles in 14 different Indic languages (and English), which come from a variety…

ArticlesHeadline Generation

CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China

2025-09-12 · Guixian Xu, Zeli Su, Ziyin Zhang, Jianing Liu 외 arxiv

Minority languages in China, such as Tibetan, Uyghur, and Traditional Mongolian, face significant challenges due to their unique writing systems, which differ from international standards. This discrepancy has led to a s…

Headline Generation

IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages

2022-03-10 · Aman Kumar, Himani Shrotriya, Prachi Sahu, Raj Dabre 외

Natural Language Generation (NLG) for non-English languages is hampered by the scarcity of datasets in these languages. In this paper, we present the IndicNLG Benchmark, a collection of datasets for benchmarking NLG for …

ArticlesBenchmarkingHeadline GenerationMachine Translation+6

L3Cube-IndicNews: News-based Short Text and Long Document Classification Datasets in Indic Languages

2024-01-04 · Aishwarya Mirashi, Srushti Sonavane, Purva Lingayat, Tejas Padhiyar 외

In this work, we introduce L3Cube-IndicNews, a multilingual text classification corpus aimed at curating a high-quality dataset for Indian regional languages, with a specific focus on news headlines and articles. We have…

ArticlesClassificationDocument ClassificationMultilingual text classification+4