paper-with-me

Papers

MEGA: Multilingual Evaluation of Generative AI

2023-03-22 · Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Maxamed Axmed, Kalika Bali, Sunayana Sitaram

Generative AI models have shown impressive performance on many Natural Language Processing tasks such as language understanding, reasoning, and language generation. An important question being asked by the AI community today is about the capabilities and limits of these models, and it is clear that evaluating generative AI is very challenging. Most studies on generative LLMs have been restricted to English and it is unclear how capable these models are at understanding and generating text in other languages. We present the first comprehensive benchmarking of generative LLMs - MEGA, which evaluates models on standard NLP benchmarks, covering 16 NLP datasets across 70 typologically diverse languages. We compare the performance of generative LLMs including Chat-GPT and GPT-4 to State of the Art (SOTA) non-autoregressive models on these tasks to determine how well generative models perform compared to the previous generation of LLMs. We present a thorough analysis of the performance of models across languages and tasks and discuss challenges in improving the performance of generative LLMs on low-resource languages. We create a framework for evaluating generative LLMs in the multilingual setting and provide directions for future progress in the field.

📄 PDF Abstract BibTeX arXiv:2303.12528

Code (1)

microsoft/Multilingual-Evaluation-of-Generative-AI-MEGA

Tasks

Benchmarking

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Position-Wise Feed-Forward Layer 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Adam 설명 없음

Similar Papers 제목 키워드 기반

MegaWika 2: A More Comprehensive Multilingual Collection of Articles and their Sources

2025-08-05 · Samuel Barham, Chandler May, Benjamin Van Durme arxiv

We introduce MegaWika 2, a large, multilingual dataset of Wikipedia articles with their citations and scraped web sources; articles are represented in a rich data structure, and scraped source texts are stored inline wit…

Question AnsweringFact Checking

Synthetic Financial Data Generation for Enhanced Financial Modelling

2025-12-25 · Christophe D. Hounwanou, Yae Ulrich Gaba, Pierre Ntakirutimana arxiv

Data scarcity and confidentiality in finance often impede model development and robust testing. This paper presents a unified multi-criteria evaluation framework for synthetic financial data and applies it to three repre…

Portfolio Optimization

MegaChat: A Synthetic Persian Q&A Dataset for High-Quality Sales Chatbot Evaluation

2025-11-28 · Mahdi Rahmani, AmirHossein Saffari, Reyhane Rahmani arxiv

Small and medium-sized enterprises (SMEs) in Iran increasingly leverage Telegram for sales, where real-time engagement is essential for conversion. However, developing AI-driven chatbots for this purpose requires large, …

Question GenerationAnswer Generation

MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks

2023-11-13 · Sanchit Ahuja, Divyanshu Aggarwal, Varun Gumma, Ishaan Watts 외

There has been a surge in LLM evaluation research to understand LLM capabilities and limitations. However, much of this research has been confined to English, leaving LLM building and evaluation for non-English languages…

Benchmarking

Mega-COV: A Billion-Scale Dataset of 100+ Languages for COVID-19

2020-05-02 · EACL 2021 2 · Muhammad Abdul-Mageed, AbdelRahim Elmadany, El Moatez Billah Nagoudi, Dinesh Pabbi 외

We describe Mega-COV, a billion-scale dataset from Twitter for studying COVID-19. The dataset is diverse (covers 268 countries), longitudinal (goes as back as 2007), multilingual (comes in 100+ languages), and has a sign…

Misinformation