paper-with-me

홈 › Papers

Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages

2022-12-11 · Sumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal, Mitesh M. Khapra, Anoop Kunchukuttan, Pratyush Kumar

Building Natural Language Understanding (NLU) capabilities for Indic languages, which have a collective speaker base of more than one billion speakers is absolutely crucial. In this work, we aim to improve the NLU capabilities of Indic languages by making contributions along 3 important axes (i) monolingual corpora (ii) NLU testsets (iii) multilingual LLMs focusing on Indic languages. Specifically, we curate the largest monolingual corpora, IndicCorp, with 20.9B tokens covering 24 languages from 4 language families - a 2.3x increase over prior work, while supporting 12 additional languages. Next, we create a human-supervised benchmark, IndicXTREME, consisting of nine diverse NLU tasks covering 20 languages. Across languages and tasks, IndicXTREME contains a total of 105 evaluation sets, of which 52 are new contributions to the literature. To the best of our knowledge, this is the first effort towards creating a standard benchmark for Indic languages that aims to test the multilingual zero-shot capabilities of pretrained language models. Finally, we train IndicBERT v2, a state-of-the-art model supporting all the languages. Averaged across languages and tasks, the model achieves an absolute improvement of 2 points over a strong baseline. The data and models are available at https://github.com/AI4Bharat/IndicBERT.

📄 PDF Abstract BibTeX arXiv:2212.05409

Code (1)

ai4bharat/indicbert 공식 구현 tf

Tasks

Natural Language UnderstandingXLM-R

Methods 이 논문이 사용한 방법론

Test 설명 없음
BASE 설명 없음
XLM-R XLM-R

Similar Papers 제목 키워드 기반

Cross-Lingual Interleaving for Speech Language Models

2025-12-01 · Adel Moumen, Guangzhi Sun, Philip C. Woodland arxiv

Spoken Language Models (SLMs) aim to learn linguistic competence directly from speech using discrete units, widening access to Natural Language Processing (NLP) technologies for languages with limited written resources. …

IndicSQuAD: A Comprehensive Multilingual Question Answering Dataset for Indic Languages

2025-05-06 · Sharvi Endait, Ruturaj Ghatage, Aditya Kulkarni, Rajlaxmi Patil 외

The rapid progress in question-answering (QA) systems has predominantly benefited high-resource languages, leaving Indic languages largely underrepresented despite their vast native speaker base. In this paper, we presen…

Question Answering

PuoBERTa: Training and evaluation of a curated language model for Setswana

2023-10-13 · Vukosi Marivate, Moseli Mots'oehli, Valencia Wagner, Richard Lastrucci 외

Natural language processing (NLP) has made significant progress for well-resourced languages such as English but lagged behind for low-resource languages like Setswana. This paper addresses this gap by presenting PuoBERT…

Language ModelingLanguage Modellingnamed-entity-recognitionNamed Entity Recognition+5

Leveraging Multilingual Training for Authorship Representation: Enhancing Generalization across Languages and Domains

2025-09-20 · Junghwan Kim, Haotian Zhang, David Jurgens arxiv

Authorship representation (AR) learning, which models an author's unique writing style, has demonstrated strong performance in authorship attribution tasks. However, prior research has primarily focused on monolingual se…

Domain GeneralizationContrastive Learning

Unveiling Language-Specific Features in Large Language Models via Sparse Autoencoders

2025-05-08 · Boyi Deng, Yu Wan, Yidan Zhang, Baosong Yang 외

The mechanisms behind multilingual capabilities in Large Language Models (LLMs) have been examined using neuron-based or internal-activation-based methods. However, these methods often face challenges such as superpositi…