paper-with-me

홈 › Papers

Digital Linguistic Bias in Spanish: Evidence from Lexical Variation in LLMs

2026-02-10 · Yoshifumi Kawasaki arxiv

This study examines the extent to which Large Language Models (LLMs) capture geographic lexical variation in Spanish, a language that exhibits substantial regional variation. Treating LLMs as virtual informants, we probe their dialectal knowledge using two survey-style question formats: Yes-No questions and multiple-choice questions. To this end, we exploited a large-scale, expert-curated database of Spanish lexical variation. Our evaluation covers more than 900 lexical items across 21 Spanish-speaking countries and is conducted at both the country and dialectal area levels. Across both evaluation formats, the results reveal systematic differences in how LLMs represent Spanish language varieties. Lexical variation associated with Spain, Equatorial Guinea, Mexico & Central America, and the La Plata River is recognized more accurately by the models, while the Chilean variety proves particularly difficult for the models to distinguish. Importantly, differences in the volume of country-level digital resources do not account for these performance patterns, suggesting that factors beyond data quantity shape dialectal representation in LLMs. By providing a fine-grained, large-scale evaluation of geographic lexical variation, this work advances empirical understanding of dialectal knowledge in LLMs and contributes new evidence to discussions of Digital Linguistic Bias in Spanish.

📄 PDF Abstract BibTeX arXiv:2602.09346

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Dialect and Gender Bias in YouTube's Spanish Captioning System

2026-02-27 · Iris Dania Jimenez, Christoph Kern arxiv

Spanish is the official language of twenty-one countries and is spoken by over 441 million people. Naturally, there are many variations in how Spanish is spoken across these countries. Media platforms such as YouTube rel…

Speech Recognition

PUCP-Metrix: An Open-source and Comprehensive Toolkit for Linguistic Analysis of Spanish Texts

2025-11-21 · Javier Alonso Villegas Luis, Marco Antonio Sobrevilla Cabezudo arxiv

Linguistic features remain essential for interpretability and tasks that involve style, structure, and readability, but existing Spanish tools offer limited coverage. We present PUCP-Metrix, an open-source and comprehens…

Text Detection

WordNet-QU: Development of a Lexical Database for Quechua Varieties

2022-10-01 · COLING 2022 10 · Nelsi Melgarejo, Rodolfo Zevallos, Hector Gomez, John E. Ortega

In the effort to minimize the risk of extinction of a language, linguistic resources are fundamental. Quechua, a low-resource language from South America, is a language spoken by millions but, despite several efforts in …

SESGO: Spanish Evaluation of Stereotypical Generative Outputs

2025-09-03 · Melissa Robles, Catalina Bernal, Denniss Raigoso, Mateo Dulce Rubio arxiv

This paper addresses the critical gap in evaluating bias in multilingual Large Language Models (LLMs), with a specific focus on Spanish language within culturally-aware Latin American contexts. Despite widespread global …

Modeling Topics and Sociolinguistic Variation in Code-Switched Discourse: Insights from Spanish-English and Spanish-Guaraní

2025-12-03 · Nemika Tyagi, Nelvin Licona Guevara, Olga Kellert arxiv

This study presents an LLM-assisted annotation pipeline for the sociolinguistic and topical analysis of bilingual discourse in two typologically distinct contexts: Spanish-English and Spanish-Guaraní. Using large languag…