paper-with-me

홈 › Papers

How far can bias go? -- Tracing bias from pretraining data to alignment

2024-11-28 · Marion Thaler, Abdullatif Köksal, Alina Leidinger, Anna Korhonen, Hinrich Schütze

As LLMs are increasingly integrated into user-facing applications, addressing biases that perpetuate societal inequalities is crucial. While much work has gone into measuring or mitigating biases in these models, fewer studies have investigated their origins. Therefore, this study examines the correlation between gender-occupation bias in pre-training data and their manifestation in LLMs, focusing on the Dolma dataset and the OLMo model. Using zero-shot prompting and token co-occurrence analyses, we explore how biases in training data influence model outputs. Our findings reveal that biases present in pre-training data are amplified in model outputs. The study also examines the effects of prompt types, hyperparameters, and instruction-tuning on bias expression, finding instruction-tuning partially alleviating representational bias while still maintaining overall stereotypical gender associations, whereas hyperparameters and prompting variation have a lesser effect on bias expression. Our research traces bias throughout the LLM development pipeline and underscores the importance of mitigating bias at the pretraining stage.

📄 PDF Abstract BibTeX arXiv:2411.19240

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A surprisingly simple technique to control the pretraining bias for better transfer: Expand or Narrow your representation

2023-04-11 · Florian Bordes, Samuel Lavoie, Randall Balestriero, Nicolas Ballas 외

Self-Supervised Learning (SSL) models rely on a pretext task to learn representations. Because this pretext task differs from the downstream tasks used to evaluate the performance of these models, there is an inherent mi…

Self-Supervised Learning

Context-Aware Counterfactual Data Augmentation for Gender Bias Mitigation in Language Models

2026-02-10 · Shweta Parihar, Liu Guangliang, Natalie Parde, Lu Cheng arxiv

A challenge in mitigating social bias in fine-tuned language models (LMs) is the potential reduction in language modeling capability, which can harm downstream performance. Counterfactual data augmentation (CDA), a widel…

Data Augmentation

The Narcissus Hypothesis: Descending to the Rung of Illusion

2025-09-22 · Riccardo Cadei, Christian Internò arxiv

Modern foundational models increasingly reflect not just world knowledge, but patterns of human preference embedded in their training data. We hypothesize that recursive alignment-via human feedback and model-generated c…

Pretraining Exposure Explains Popularity Judgments in Large Language Models

2026-05-12 · Jamshid Mozafari, Bhawna Piryani, Adam Jatowt arxiv

Large language models (LLMs) exhibit systematic preferences for well-known entities, a phenomenon often attributed to popularity bias. However, the extent to which these preferences reflect real-world popularity versus s…

Disentangled Knowledge Tracing for Alleviating Cognitive Bias

2025-03-04 · Yiyun Zhou, Zheqi Lv, Shengyu Zhang, Jingyuan Chen

In the realm of Intelligent Tutoring System (ITS), the accurate assessment of students' knowledge states through Knowledge Tracing (KT) is crucial for personalized learning. However, due to data bias, $\textit{i.e.}$, th…

Knowledge Tracing