paper-with-me

홈 › Papers

AAVENUE: Detecting LLM Biases on NLU Tasks in AAVE via a Novel Benchmark

2024-08-27 · Abhay Gupta, Philip Meng, Ece Yurtseven, Sean O'Brien, Kevin Zhu

Detecting biases in natural language understanding (NLU) for African American Vernacular English (AAVE) is crucial to developing inclusive natural language processing (NLP) systems. To address dialect-induced performance discrepancies, we introduce AAVENUE ({AAVE} {N}atural Language {U}nderstanding {E}valuation), a benchmark for evaluating large language model (LLM) performance on NLU tasks in AAVE and Standard American English (SAE). AAVENUE builds upon and extends existing benchmarks like VALUE, replacing deterministic syntactic and morphological transformations with a more flexible methodology leveraging LLM-based translation with few-shot prompting, improving performance across our evaluation metrics when translating key tasks from the GLUE and SuperGLUE benchmarks. We compare AAVENUE and VALUE translations using five popular LLMs and a comprehensive set of metrics including fluency, BARTScore, quality, coherence, and understandability. Additionally, we recruit fluent AAVE speakers to validate our translations for authenticity. Our evaluations reveal that LLMs consistently perform better on SAE tasks than AAVE-translated versions, underscoring inherent biases and highlighting the need for more inclusive NLP models. We have open-sourced our source code on GitHub and created a website to showcase our work at https://aavenue.live.

📄 PDF Abstract BibTeX arXiv:2408.14845

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelNatural Language Understanding

Methods 이 논문이 사용한 방법론

American 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

One Language, Many Gaps: Evaluating Dialect Fairness and Robustness of Large Language Models in Reasoning Tasks

2024-10-14 · Fangru Lin, Shaoguang Mao, Emanuele La Malfa, Valentin Hofmann 외

Language is not monolithic. While benchmarks, including those designed for multiple languages, are often used as proxies to evaluate the performance of Large Language Models (LLMs), they tend to overlook the nuances of w…

FairnessGSM8KHumanEvalMath

Evaluating the Usage of African-American Vernacular English in Large Language Models

2026-02-25 · Deja Dunlap, R. Thomas McCoy arxiv

In AI, most evaluations of natural language understanding tasks are conducted in standardized dialects such as Standard American English (SAE). In this work, we investigate how accurately large language models (LLMs) rep…

Natural Language UnderstandingSentiment Analysis

Finding A Voice: Evaluating African American Dialect Generation for Chatbot Technology

2025-01-07 · Sarah E. Finch, Ellie S. Paek, Sejung Kwon, Ikseon Choi 외

As chatbots become increasingly integrated into everyday tasks, designing systems that accommodate diverse user populations is crucial for fostering trust, engagement, and inclusivity. This study investigates the ability…

ChatbotDiversity

Side-by-side Comparison Amplifies Dialect Bias in Language Models

2026-05-23 · Kritee Kondapally, Claire J. Smerdon, Pooja C. Patel, Ogheneyoma Akoni 외 arxiv

Language models (LMs) can exhibit biases based on variations in their dialects, even in the absence of a dialect label, a behavior known as covert dialect bias. In this work, we quantify covert dialect bias in online dis…

Decision Making

Liquidity Risks in Lending Protocols: Evidence from Aave Protocol

2022-06-23 · Xiaotong Sun, Charalampos Stasinakis, Georgios Sermpinis

Lending Protocols (LPs), as blockchain-based lending systems, allow any agents to borrow and lend cryptocurrencies. However, liquidity risks could occur, especially when salient loans are initiated by a particular group …