paper-with-me

홈 › Papers

LAG-MMLU: Benchmarking Frontier LLM Understanding in Latvian and Giriama

2025-03-14 · Naome A. Etori, Kevin Lu, Randu Karisa, Arturs Kanepajs

As large language models (LLMs) rapidly advance, evaluating their performance is critical. LLMs are trained on multilingual data, but their reasoning abilities are mainly evaluated using English datasets. Hence, robust evaluation frameworks are needed using high-quality non-English datasets, especially low-resource languages (LRLs). This study evaluates eight state-of-the-art (SOTA) LLMs on Latvian and Giriama using a Massive Multitask Language Understanding (MMLU) subset curated with native speakers for linguistic and cultural relevance. Giriama is benchmarked for the first time. Our evaluation shows that OpenAI's o1 model outperforms others across all languages, scoring 92.8% in English, 88.8% in Latvian, and 70.8% in Giriama on 0-shot tasks. Mistral-large (35.6%) and Llama-70B IT (41%) have weak performance, on both Latvian and Giriama. Our results underscore the need for localized benchmarks and human evaluations in advancing cultural AI contextualization.

📄 PDF Abstract BibTeX arXiv:2503.11911

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingMMLU

Similar Papers 제목 키워드 기반

Pretraining and Benchmarking Modern Encoders for Latvian

2026-03-16 · Arturs Znotins arxiv

Encoder-only transformers remain essential for practical NLP tasks. While recent advances in multilingual models have improved cross-lingual capabilities, low-resource languages such as Latvian remain underrepresented in…

IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding

2025-01-27 · Sankalp KJ, Ashutosh Kumar, Laxmaan Balaji, Nikunj Kotecha 외

Known by more than 1.5 billion people in the Indian subcontinent, Indic languages present unique challenges and opportunities for natural language processing (NLP) research due to their rich cultural heritage, linguistic…

BenchmarkingDiversityMMLUMulti-task Language Understanding

DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models

2025-10-31 · Malik H. Altakrori, Nizar Habash, Abed Alhakim Freihat, Younes Samih 외 arxiv

We present DialectalArabicMMLU, a new benchmark for evaluating the performance of large language models (LLMs) across Arabic dialects. While recently developed Arabic and multilingual benchmarks have advanced LLM evaluat…

Latvian National Corpora Collection – Korpuss.lv

2022-06-01 · LREC 2022 6 · Baiba Saulite, Roberts Darģis, Normunds Gruzitis, Ilze Auzina 외

LNCC is a diverse collection of Latvian language corpora representing both written and spoken language and is useful for both linguistic research and language modelling. The collection is intended to cover diverse Latvia…

Cultural Vocal Bursts Intensity PredictionLanguage Modelling

LaVA – Latvian Language Learner corpus

2022-06-01 · LREC 2022 6 · Roberts Darģis, Ilze Auziņa, Inga Kaija, Kristīne Levāne-Petrova 외

This paper presents the Latvian Language Learner Corpus (LaVA) developed at the Institute of Mathematics and Computer Science, University of Latvia. LaVA corpus contains 1015 essays (190k tokens and 790k characters exclu…