paper-with-me

홈 › Papers

LoraxBench: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages

2025-08-17 · Alham Fikri Aji, Trevor Cohn arxiv

As one of the world's most populous countries, with 700 languages spoken, Indonesia is behind in terms of NLP progress. We introduce LoraxBench, a benchmark that focuses on low-resource languages of Indonesia and covers 6 diverse tasks: reading comprehension, open-domain QA, language inference, causal reasoning, translation, and cultural QA. Our dataset covers 20 languages, with the addition of two formality registers for three languages. We evaluate a diverse set of multilingual and region-focused LLMs and found that this benchmark is challenging. We note a visible discrepancy between performance in Indonesian and other languages, especially the low-resource ones. There is no clear lead when using a region-specific model as opposed to the general multilingual model. Lastly, we show that a change in register affects model performance, especially with registers not commonly found in social media, such as high-level politeness `Krama' Javanese.

📄 PDF Abstract BibTeX arXiv:2508.12459

Code (0)

등록된 구현이 없습니다.

Tasks

Reading Comprehension

Similar Papers 제목 키워드 기반

SEA-HELM: Southeast Asian Holistic Evaluation of Language Models

2025-02-20 · Yosephine Susanto, Adithya Venkatadri Hulagadri, Jann Railey Montalan, Jian Gang Ngui 외

With the rapid emergence of novel capabilities in Large Language Models (LLMs), the need for rigorous multilingual and multicultural benchmarks that are integrated has become more pronounced. Though existing LLM benchmar…

Improving Indonesian Text Classification Using Multilingual Language Model

2020-09-12 · Ilham Firdausi Putra, Ayu Purwarianti

Compared to English, the amount of labeled data for Indonesian text classification tasks is very small. Recently developed multilingual language models have shown its ability to create multilingual representations effect…

ClassificationGeneral ClassificationLanguage ModelingLanguage Modelling+4

Does Language Shift Break Medical Vision-Language Models? Indonesian Radiology Visual Question Answering Case Study

2026-06-02 · Pieter Christy Yan Yudhistira, Dzaki Rafif Malik, Novanto Yudistira arxiv

Medical Vision-Language Models (VLMs) are typically evaluated on English radiology visual question answering benchmarks, leaving their robustness under non-English clinical language largely unexplored. We introduce IndoR…

Visual Question Answering

BHASA: A Holistic Southeast Asian Linguistic and Cultural Evaluation Suite for Large Language Models

2023-09-12 · Wei Qi Leong, Jian Gang Ngui, Yosephine Susanto, Hamsawardhini Rengarajan 외

The rapid development of Large Language Models (LLMs) and the emergence of novel abilities with scale have necessitated the construction of holistic, diverse and challenging benchmarks such as HELM and BIG-bench. However…

DiagnosticNatural Language UnderstandingSensitivity

COPAL-ID: Indonesian Language Reasoning with Local Culture and Nuances

2023-11-02 · Haryo Akbarianto Wibowo, Erland Hilman Fuadi, Made Nindyatama Nityasya, Radityo Eko Prasojo 외

We present COPAL-ID, a novel, public Indonesian language common sense reasoning dataset. Unlike the previous Indonesian COPA dataset (XCOPA-ID), COPAL-ID incorporates Indonesian local and cultural nuances, and therefore,…

Common Sense Reasoning