paper-with-me

홈 › Papers

From n-gram to Attention: How Model Architectures Learn and Propagate Bias in Language Modeling

2025-05-18 · Mohsinul Kabir, Tasfia Tahsin, Sophia Ananiadou

Current research on bias in language models (LMs) predominantly focuses on data quality, with significantly less attention paid to model architecture and temporal influences of data. Even more critically, few studies systematically investigate the origins of bias. We propose a methodology grounded in comparative behavioral theory to interpret the complex interaction between training data and model architecture in bias propagation during language modeling. Building on recent work that relates transformers to n-gram LMs, we evaluate how data, model design choices, and temporal dynamics affect bias propagation. Our findings reveal that: (1) n-gram LMs are highly sensitive to context window size in bias propagation, while transformers demonstrate architectural robustness; (2) the temporal provenance of training data significantly affects bias; and (3) different model architectures respond differentially to controlled bias injection, with certain biases (e.g. sexual orientation) being disproportionately amplified. As language models become ubiquitous, our findings highlight the need for a holistic approach -- tracing bias to its origins across both data and model dimensions, not just symptoms, to mitigate harm.

📄 PDF Abstract BibTeX arXiv:2505.12381

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

[RE] Double-Hard Debias: Tailoring Word Embeddings for Gender Bias Mitigation

2021-04-14 · Haswanth Aekula, Sugam Garg, Animesh Gupta

Despite widespread use in natural language processing (NLP) tasks, word embeddings have been criticized for inheriting unintended gender bias from training corpora. programmer is more closely associated with man and home…

Word Embeddings

How Implicit Bias Accumulates and Propagates in LLM Long-term Memory

2026-02-02 · Yiming Ma, Lixu Wang, Lionel Z. Wang, Hongkun Yang 외 arxiv

Long-term memory mechanisms enable Large Language Models (LLMs) to maintain continuity and personalization across extended interaction lifecycles, but they also introduce new and underexplored risks related to fairness. …

Differentiable Convex Optimization Layers

2019-10-28 · NeurIPS 2019 12 · Akshay Agrawal, Brandon Amos, Shane Barratt, Stephen Boyd 외

Recent work has shown how to embed differentiable optimization problems (that is, problems whose solutions can be backpropagated through) as layers within deep learning architectures. This method provides a useful induct…

Inductive Bias

Leveraging Large Language Models to Measure Gender Representation Bias in Gendered Language Corpora

2024-06-19 · Erik Derner, Sara Sansalvador de la Fuente, Yoan Gutiérrez, Paloma Moreda 외

Large language models (LLMs) often inherit and amplify social biases embedded in their training data. A prominent social bias is gender bias. In this regard, prior work has mainly focused on gender stereotyping bias - th…

Multilingual NLP

Memory Contagion: Cross-Temporal Propagation of Evaluator Bias via Agent Memory

2026-06-22 · Zewen Liu arxiv

Large Language Model (LLM) agents increasingly rely on memory systems to maintain long-term coherence. Recent work shows that agent memories degrade during continuous consolidation. However, existing research assumes mem…