paper-with-me

홈 › Papers

CaLMQA: Exploring culturally specific long-form question answering across 23 languages

2024-06-25 · Shane Arora, Marzena Karpinska, Hung-Ting Chen, Ipsita Bhattacharjee, Mohit Iyyer, Eunsol Choi

Large language models (LLMs) are used for long-form question answering (LFQA), which requires them to generate paragraph-length answers to complex questions. While LFQA has been well-studied in English, this research has not been extended to other languages. To bridge this gap, we introduce CaLMQA, a collection of 1.5K complex culturally specific questions spanning 23 languages and 51 culturally agnostic questions translated from English into 22 other languages. We define culturally specific questions as those uniquely or more likely to be asked by people from cultures associated with the question's language. We collect naturally-occurring questions from community web forums and hire native speakers to write questions to cover under-resourced, rarely-studied languages such as Fijian and Kirundi. Our dataset contains diverse, complex questions that reflect cultural topics (e.g. traditions, laws, news) and the language usage of native speakers. We automatically evaluate a suite of open- and closed-source models on CaLMQA by detecting incorrect language and token repetitions in answers, and observe that the quality of LLM-generated answers degrades significantly for some low-resource languages. Lastly, we perform human evaluation on a subset of models and languages. Manual evaluation reveals that model performance is significantly worse for culturally specific questions than for culturally agnostic questions. Our findings highlight the need for further research in non-English LFQA and provide an evaluation framework.

📄 PDF Abstract BibTeX arXiv:2406.17761

Code (1)

2015aroras/calmqa 공식 구현

Tasks

FormLong Form Question AnsweringQuestion Answering

Similar Papers 제목 키워드 기반

CIVICS: Building a Dataset for Examining Culturally-Informed Values in Large Language Models

2024-05-22 · Giada Pistilli, Alina Leidinger, Yacine Jernite, Atoosa Kasirzadeh 외

This paper introduces the "CIVICS: Culturally-Informed & Values-Inclusive Corpus for Societal impacts" dataset, designed to evaluate the social and cultural variation of Large Language Models (LLMs) across multiple langu…

The Myth of Culturally Agnostic AI Models

2022-11-28 · Eva Cetinic

The paper discusses the potential of large vision-language models as objects of interest for empirical cultural studies. Focusing on the comparative analysis of outputs from two popular text-to-image synthesis models, DA…

Image GenerationMemorizationSpecificity

BengaliFig: A Low-Resource Challenge for Figurative and Culturally Grounded Reasoning in Bengali

2025-11-25 · Abdullah Al Sefat arxiv

Large language models excel on broad multilingual benchmarks but remain to be evaluated extensively in figurative and culturally grounded reasoning, especially in low-resource contexts. We present BengaliFig, a compact y…

Exploring Cultural Nuances in Emotion Perception Across 15 African Languages

2025-03-25 · Ibrahim Said Ahmad, Shiran Dudy, Tadesse Destaw Belay, Idris Abdulmumin 외

Understanding how emotions are expressed across languages is vital for building culturally-aware and inclusive NLP systems. However, emotion expression in African languages is understudied, limiting the development of ef…

Transfer Learning

CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation

2025-05-30 · Emilio Villa-Cueva, Sholpan Bolatzhanova, Diana Turmakhan, Kareem Elzeky 외

Cultural content poses challenges for machine translation systems due to the differences in conceptualizations between cultures, where language alone may fail to convey sufficient context to capture region-specific meani…

BenchmarkingMachine TranslationMultimodal Machine TranslationTranslation