paper-with-me

홈 › Papers

IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages

2024-04-25 · Harman Singh, Nitish Gupta, Shikhar Bharadwaj, Dinesh Tewari, Partha Talukdar

As large language models (LLMs) see increasing adoption across the globe, it is imperative for LLMs to be representative of the linguistic diversity of the world. India is a linguistically diverse country of 1.4 Billion people. To facilitate research on multilingual LLM evaluation, we release IndicGenBench - the largest benchmark for evaluating LLMs on user-facing generation tasks across a diverse set 29 of Indic languages covering 13 scripts and 4 language families. IndicGenBench is composed of diverse generation tasks like cross-lingual summarization, machine translation, and cross-lingual question answering. IndicGenBench extends existing benchmarks to many Indic languages through human curation providing multi-way parallel evaluation data for many under-represented Indic languages for the first time. We evaluate a wide range of proprietary and open-source LLMs including GPT-3.5, GPT-4, PaLM-2, mT5, Gemma, BLOOM and LLaMA on IndicGenBench in a variety of settings. The largest PaLM-2 models performs the best on most tasks, however, there is a significant performance gap in all languages compared to English showing that further research is needed for the development of more inclusive multilingual language models. IndicGenBench is released at www.github.com/google-research-datasets/indic-gen-bench

📄 PDF Abstract BibTeX arXiv:2404.16816

Code (1)

google-research-datasets/indic-gen-bench 공식 구현

Tasks

Cross-Lingual Question AnsweringDiversityMachine TranslationQuestion Answering

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Attention 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Adafactor Adafactor is a stochastic optimization method based on Adam that reduces memory usage while retaining the empirical benefits of…
Position-Wise Feed-Forward Layer 설명 없음
SentencePiece 설명 없음
Inverse Square Root Schedule Inverse Square Root is a learning rate schedule 1 / $\sqrt{\max\left(n, k\right)}$ where $n$ is the current training iteration and $k$ is the number of warm-up steps. This…

Similar Papers 제목 키워드 기반

AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators

2025-08-12 · Jason Chou, Ao Liu, Yuchi Deng, Zhiying Zeng 외 arxiv

Large Language Models (LLMs) have demonstrated remarkable capabilities across various domains, with code generation emerging as a key area of focus. While numerous benchmarks have been proposed to evaluate their code gen…

Code Generation

Quantum-RAG and PunGPT2: Advancing Low-Resource Language Generation and Retrieval for the Punjabi Language

2025-08-03 · Jaskaranjeet Singh, Rakesh Thakur arxiv

Despite rapid advances in large language models (LLMs), low-resource languages remain excluded from NLP, limiting digital access for millions. We present PunGPT2, the first fully open-source Punjabi generative model suit…

Question Answering

MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language

2025-05-20 · Seyoung Song, Seogyeong Jeong, Eunsu Kim, Jiho Jin 외

Evaluating text generation capabilities of large language models (LLMs) is challenging, particularly for low-resource languages where methods for direct assessment are scarce. We propose MUG-Eval, a novel framework that …

Text Generation

Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss

2026-04-14 · Ronald Skorobogat, Ameya Prabhu, Matthias Bethge arxiv

Multilingual benchmarks guide the development of frontier models. Yet multilingual evaluations reported by frontier models are structured similar to popular reasoning and knowledge benchmarks, but across many languages. …

Mathematical Reasoning

MTG: A Benchmark Suite for Multilingual Text Generation

2021-08-13 · Findings (NAACL) 2022 7 · Yiran Chen, Zhenqiao Song, Xianze Wu, Danqing Wang 외

We introduce MTG, a new benchmark suite for training and evaluating multilingual text generation. It is the first-proposed multilingual multiway text generation dataset with the largest human-annotated data (400k). It in…

Question GenerationQuestion-GenerationStory GenerationText Generation+2