paper-with-me

홈 › Papers

Can Small-Scale Data Poisoning Exacerbate Dialect-Linked Biases in Large Language Models?

2025-07-25 · Chaymaa Abbas, Mariette Awad, Razane Tajeddine arxiv

Style-conditioned data poisoning is identified as a covert vector for amplifying sociolinguistic bias in large language models. Using small poisoned budgets that pair dialectal prompts -- principally African American Vernacular English (AAVE) and a Southern dialect -- with toxic or stereotyped completions during instruction tuning, this work probes whether linguistic style can act as a latent trigger for harmful behavior. Across multiple model families and scales, poisoned exposure elevates toxicity and stereotype expression for dialectal inputs -- most consistently for AAVE -- while Standard American English remains comparatively lower yet not immune. A multi-metric audit combining classifier-based toxicity with an LLM-as-a-judge reveals stereotype-laden content even when lexical toxicity appears muted, indicating that conventional detectors under-estimate sociolinguistic harms. Additionally, poisoned models exhibit emergent jailbreaking despite the absence of explicit slurs in the poison, suggesting weakened alignment rather than memorization. These findings underscore the need for dialect-aware evaluation, content-level stereotype auditing, and training protocols that explicitly decouple style from toxicity to prevent bias amplification through seemingly minor, style-based contamination.

📄 PDF Abstract BibTeX arXiv:2507.19195

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AutoDetect: Designing an Autoencoder-based Detection Method for Poisoning Attacks on Object Detection Applications in the Military Domain

2025-09-03 · Alma M. Liezenga, Stefan Wijnja, Puck de Haan, Niels W. T. Brink 외 arxiv

Poisoning attacks pose an increasing threat to the security and robustness of Artificial Intelligence systems in the military domain. The widespread use of open-source datasets and pretrained models exacerbates this risk…

Anomaly DetectionObject Detection

Noise-Robust Morphological Disambiguation for Dialectal Arabic

2018-06-01 · NAACL 2018 6 · Nasser Zalmout, Alex Erdmann, er, Nizar Habash

User-generated text tends to be noisy with many lexical and orthographic inconsistencies, making natural language processing (NLP) tasks more challenging. The challenging nature of noisy text processing is exacerbated fo…

Lexical NormalizationMorphological AnalysisMorphological DisambiguationMorphological Tagging+1

Learning to Recognize Dialect Features

2020-10-23 · NAACL 2021 4 · Dorottya Demszky, Devyani Sharma, Jonathan H. Clark, Vinodkumar Prabhakaran 외

Building NLP systems that serve everyone requires accounting for dialect differences. But dialects are not monolithic entities: rather, distinctions between and within dialects are captured by the presence, absence, and …

Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples

2025-10-08 · Alexandra Souly, Javier Rando, Ed Chapman, Xander Davies 외 arxiv

Poisoning attacks can compromise the safety of large language models (LLMs) by injecting malicious documents into their training data. Existing work has studied pretraining poisoning assuming adversaries control a percen…

Side-by-side Comparison Amplifies Dialect Bias in Language Models

2026-05-23 · Kritee Kondapally, Claire J. Smerdon, Pooja C. Patel, Ogheneyoma Akoni 외 arxiv

Language models (LMs) can exhibit biases based on variations in their dialects, even in the absence of a dialect label, a behavior known as covert dialect bias. In this work, we quantify covert dialect bias in online dis…

Decision Making