paper-with-me

Papers

Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants

2025-10-28 · Hunzalah Hassan Bhatti, Firoj Alam arxiv

Large Language Models (LLMs) are increasingly used to answer everyday questions, yet their performance on culturally grounded and dialectal content remains uneven across languages. We propose a comprehensive method that (i) translates Modern Standard Arabic (MSA) multiple-choice questions (MCQs) into English and several Arabic dialects, (ii) converts them into open-ended questions (OEQs), (iii) benchmarks a range of zero-shot and fine-tuned LLMs under both MCQ and OEQ settings, and (iv) generates chain-of-thought (CoT) rationales to fine-tune models for step-by-step reasoning. Using this method, we extend an existing dataset in which QAs are parallelly aligned across multiple language varieties, making it, to our knowledge, the first of its kind. We conduct extensive experiments with both open and closed models. Our findings show that (i) models underperform on Arabic dialects, revealing persistent gaps in culturally grounded and dialect-specific knowledge; (ii) Arabic-centric models perform well on MCQs but struggle with OEQs; and (iii) CoT improves judged correctness while yielding mixed n-gram-based metrics. The developed dataset will be publicly released to support further research on culturally and linguistically inclusive evaluation.

📄 PDF Abstract BibTeX arXiv:2510.24328

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CamelEval: Advancing Culturally Aligned Arabic Language Models and Benchmarks

2024-09-19 · Zhaozhi Qian, Faroq Altam, Muhammad Alqurishi, Riad Souissi

Large Language Models (LLMs) are the cornerstones of modern artificial intelligence systems. This paper introduces Juhaina, a Arabic-English bilingual LLM specifically designed to align with the values and preferences of…

Instruction FollowingOpen-Ended Question AnsweringQuestion Answering

Beyond Understanding: Evaluating the Pragmatic Gap in LLMs' Cultural Processing of Figurative Language

2025-10-27 · Mena Attia, Aashiq Muhamed, Mai Alkhamissi, Thamar Solorio 외 arxiv

We present a comprehensive evaluation of the ability of large language models (LLMs) to process culturally grounded language, specifically to understand and pragmatically use figurative expressions that encode local know…

English Proverbs

OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA

2025-10-07 · Firoj Alam, Ali Ezzat Shahroor, Md. Arid Hasan, Zien Sheikh Ali 외 arxiv

Large-scale multimodal models achieve strong results on tasks like Visual Question Answering (VQA), but they are often limited when queries require cultural and visual information, everyday knowledge, particularly in low…

Visual Question AnsweringObject Recognition

AceGPT, Localizing Large Language Models in Arabic

2023-09-21 · Huang Huang, Fei Yu, Jianqing Zhu, Xuening Sun 외

This paper is devoted to the development of a localized Large Language Model (LLM) specifically for Arabic, a language imbued with unique cultural characteristics inadequately addressed by current mainstream models. Sign…

Instruction FollowingLanguage ModelingLanguage ModellingLarge Language Model+1

Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs

2025-05-23 · Wafa Alghallabi, Ritesh Thawkar, Sara Ghaboura, Ketan More 외

Arabic poetry is one of the richest and most culturally rooted forms of expression in the Arabic language, known for its layered meanings, stylistic diversity, and deep historical continuity. Although large language mode…