paper-with-me

홈 › Papers

AmharicStoryQA: A Multicultural Story Question Answering Benchmark in Amharic

2026-02-02 · Israel Abebe Azime, Abenezer Kebede Angamo, Hana Mekonen Tamiru, Dagnachew Mekonnen Marilign, Philipp Slusallek, Seid Muhie Yimam, Dietrich Klakow arxiv

With the growing emphasis on multilingual and cultural evaluation benchmarks for large language models, language and culture are often treated as synonymous, and performance is commonly used as a proxy for a models understanding of a given language. In this work, we argue that such evaluations overlook meaningful cultural variation that exists within a single language. We address this gap by focusing on narratives from different regions of Ethiopia and demonstrate that, despite shared linguistic characteristics, region-specific and domain-specific content substantially influences language evaluation outcomes. To this end, we introduce \textbf{\textit{AmharicStoryQA}}, a long-sequence story question answering benchmark grounded in culturally diverse narratives from Amharic-speaking regions. Using this benchmark, we reveal a significant narrative understanding gap in existing LLMs, highlight pronounced regional differences in evaluation results, and show that supervised fine-tuning yields uneven improvements across regions and evaluation settings. Our findings emphasize the need for culturally grounded benchmarks that go beyond language-level evaluation to more accurately assess and improve narrative understanding in low-resource languages.

📄 PDF Abstract BibTeX arXiv:2602.02774

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines

2024-10-16 · Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha 외

Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts. To evaluate their understanding of such knowledge, we …

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports

2025-09-24 · Punit Kumar Singh, Nishant Kumar, Akash Ghosh, Kunal Pasad 외 arxiv

Language Models (LMs) are primarily evaluated on globally popular sports, often overlooking regional and indigenous sporting traditions. To address this gap, we introduce \textbf{\textit{CultSportQA}}, a benchmark design…

VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding

2026-01-12 · Haorui Yu, Diji Yang, Hang He, Fengrui Zhang 외 arxiv

We introduce VULCA-Bench, a multicultural art-critique benchmark for evaluating Vision-Language Models' (VLMs) cultural understanding beyond surface-level visual perception. Existing VLM benchmarks predominantly measure …

Question AnsweringObject Recognition

StoryQA: Story Grounded Question Answering Dataset

2022-01-16 · ACL ARR January 2022 1 · Anonymous

The abundance of benchmark datasets supports the recent trend of increased attention given to Question Answering (QA) tasks. However, most of them lack a diverse selection of QA types and more challenging questions. In t…

Question Answering

Progressive Attention Memory Network for Movie Story Question Answering

2019-04-18 · CVPR 2019 6 · Junyeong Kim, Minuk Ma, Kyung-Su Kim, Sungjin Kim 외

This paper proposes the progressive attention memory network (PAMN) for movie story question answering (QA). Movie story QA is challenging compared to VQA in two aspects: (1) pinpointing the temporal parts relevant to an…

Question AnsweringVideo Story QAVisual Question Answering (VQA)