paper-with-me

홈 › Papers

EnDive: A Cross-Dialect Benchmark for Fairness and Performance in Large Language Models

2025-02-25 · Abhay Gupta, Jacob Cheung, Philip Meng, Shayan Sayyed, Austen Liao, Kevin Zhu, Sean O'Brien

The diversity of human language, shaped by social, cultural, and regional influences, presents significant challenges for natural language processing (NLP) systems. Existing benchmarks often overlook intra-language variations, leaving speakers of non-standard dialects underserved. To address this gap, we introduce EnDive (English Diversity), a benchmark that evaluates five widely-used large language models (LLMs) across tasks in language understanding, algorithmic reasoning, mathematics, and logic. Our framework translates Standard American English datasets into five underrepresented dialects using few-shot prompting with verified examples from native speakers, and compare these translations against rule-based methods via fluency assessments, preference tests, and semantic similarity metrics. Human evaluations confirm high translation quality, with average scores of at least 6.02/7 for faithfulness, fluency, and formality. By filtering out near-identical translations, we create a challenging dataset that reveals significant performance disparities - models consistently underperform on dialectal inputs compared to Standard American English. EnDive thus advances dialect-aware NLP by uncovering model biases and promoting more equitable language technologies.

📄 PDF Abstract BibTeX arXiv:2504.07100

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityFairnessSemantic SimilaritySemantic Textual Similarity

Methods 이 논문이 사용한 방법론

American 설명 없음

Similar Papers 제목 키워드 기반

One Language, Many Gaps: Evaluating Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks

2024-10-14 · Fangru Lin, Shaoguang Mao, Emanuele La Malfa, Valentin Hofmann 외

Language is not monolithic. While benchmarks, including those designed for multiple languages, are often used as proxies to evaluate the performance of Large Language Models (LLMs), they tend to overlook the nuances of w…

FairnessGSM8KHumanEvalMath

Disentangling Dialect from Social Bias via Multitask Learning to Improve Fairness

2024-06-14 · Maximilian Spliethöver, Sai Nikhil Menon, Henning Wachsmuth

Dialects introduce syntactic and lexical variations in language that occur in regional or social groups. Most NLP methods are not sensitive to such variations. This may lead to unfair behavior of the methods, conveying n…

Fairness

SD-QA: Spoken Dialectal Question Answering for the Real World

2021-09-24 · Findings (EMNLP) 2021 11 · Fahim Faisal, Sharlina Keshava, Md Mahfuz ibn Alam, Antonios Anastasopoulos

Question answering (QA) systems are now available through numerous commercial applications for a wide variety of domains, serving millions of users that interact with them via speech interfaces. However, current benchmar…

FairnessQuestion Answeringspeech-recognitionSpeech Recognition

UniQL: Towards Dialect-Universal Benchmarking for Text-to-SQL

2026-06-06 · Jianling Gao, Chongyang Tao, Jiayuan Bai, Liu Yang 외 arxiv

Existing text-to-SQL benchmarks are largely centered on SQLite, making it difficult to evaluate whether models can generalize across heterogeneous SQL dialects. However, real-world database systems differ substantially i…

DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models

2025-10-31 · Malik H. Altakrori, Nizar Habash, Abed Alhakim Freihat, Younes Samih 외 arxiv

We present DialectalArabicMMLU, a new benchmark for evaluating the performance of large language models (LLMs) across Arabic dialects. While recently developed Arabic and multilingual benchmarks have advanced LLM evaluat…