paper-with-me

Papers

Large Language Models Discriminate Against Speakers of German Dialects

2025-09-17 · Minh Duc Bui, Carolin Holtermann, Valentin Hofmann, Anne Lauscher, Katharina von der Wense arxiv

Dialects represent a significant component of human culture and are found across all regions of the world. In Germany, more than 40% of the population speaks a regional dialect (Adler and Hansen, 2022). However, despite cultural importance, individuals speaking dialects often face negative societal stereotypes. We examine whether such stereotypes are mirrored by large language models (LLMs). We draw on the sociolinguistic literature on dialect perception to analyze traits commonly associated with dialect speakers. Based on these traits, we assess the dialect naming bias and dialect usage bias expressed by LLMs in two tasks: an association task and a decision task. To assess a model's dialect usage bias, we construct a novel evaluation corpus that pairs sentences from seven regional German dialects (e.g., Alemannic and Bavarian) with their standard German counterparts. We find that: (1) in the association task, all evaluated LLMs exhibit significant dialect naming and dialect usage bias against German dialect speakers, reflected in negative adjective associations; (2) all models reproduce these dialect naming and dialect usage biases in their decision making; and (3) contrary to prior work showing minimal bias with explicit demographic mentions, we find that explicitly labeling linguistic demographics--German dialect speakers--amplifies bias more than implicit cues like dialect usage.

📄 PDF Abstract BibTeX arXiv:2509.13835

Code (0)

등록된 구현이 없습니다.

Tasks

Decision Making

Similar Papers 제목 키워드 기반

What about em? How Commercial Machine Translation Fails to Handle (Neo-)Pronouns

2023-05-25 · Anne Lauscher, Debora Nozza, Archie Crowley, Ehm Miltersen 외

As 3rd-person pronoun usage shifts to include novel forms, e.g., neopronouns, we need more research on identity-inclusive NLP. Exclusion is particularly harmful in one of the most popular NLP applications, machine transl…

Machine TranslationTranslation

Open Source Automatic Speech Recognition for German

2018-07-26 · Benjamin Milde, Arne Köhn

High quality Automatic Speech Recognition (ASR) is a prerequisite for speech-based applications and research. While state-of-the-art ASR software is freely available, the language dependent acoustic models are lacking fo…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Diversityspeech-recognition+1

Grammatical Gender, Neo-Whorfianism, and Word Embeddings: A Data-Driven Approach to Linguistic Relativity

2019-10-22 · Katharina Kann

The relation between language and thought has occupied linguists for at least a century. Neo-Whorfianism, a weak version of the controversial Sapir-Whorf hypothesis, holds that our thoughts are subtly influenced by the g…

Experimental DesignWord Embeddings

Neural Speech Synthesis in German

2021-10-03 · 14th International Conference on Advances in Human-oriented and Personalized Mechanisms, Technologies, and Services 2021 10 · Johannes Wirth, Pascal Puchtler, René Peinl

While many speech synthesis systems based on deep neural networks are thoroughly evaluated and released for free use in English, models for languages with far less active speakers like German are scarcely trained and mos…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

CMU-MOSEAS: A Multimodal Language Dataset for Spanish, Portuguese, German and French

2020-11-01 · EMNLP 2020 11 · AmirAli Bagher Zadeh, Yansheng Cao, Simon Hessner, Paul Pu Liang 외

Modeling multimodal language is a core research area in natural language processing. While languages such as English have relatively large multimodal language resources, other widely spoken languages across the globe hav…