paper-with-me

홈 › Papers

Manipulating and Mitigating Generative Model Biases without Retraining

2024-04-03 · Jordan Vice, Naveed Akhtar, Richard Hartley, Ajmal Mian

Text-to-image (T2I) generative models have gained increased popularity in the public domain. While boasting impressive user-guided generative abilities, their black-box nature exposes users to intentionally- and intrinsically-biased outputs. Bias manipulation (and mitigation) techniques typically rely on careful tuning of learning parameters and training data to adjust decision boundaries to influence model bias characteristics, which is often computationally demanding. We propose a dynamic and computationally efficient manipulation of T2I model biases by exploiting their rich language embedding spaces without model retraining. We show that leveraging foundational vector algebra allows for a convenient control over language model embeddings to shift T2I model outputs and control the distribution of generated classes. As a by-product, this control serves as a form of precise prompt engineering to generate images which are generally implausible using regular text prompts. We demonstrate a constructive application of our technique by balancing the frequency of social classes in generated images, effectively balancing class distributions across three social bias dimensions. We also highlight a negative implication of bias manipulation by framing our method as a backdoor attack with severity control using semantically-null input triggers, reporting up to 100% attack success rate. Key-words: Text-to-Image Models, Generative Models, Bias, Prompt Engineering, Backdoor Attacks

📄 PDF Abstract BibTeX arXiv:2404.02530

Code (0)

등록된 구현이 없습니다.

Tasks

Backdoor AttackLanguage ModellingmodelPrompt Engineering

Similar Papers 제목 키워드 기반

Sensing and Steering Stereotypes: Extracting and Applying Gender Representation Vectors in LLMs

2025-02-27 · Hannah Cyberey, Yangfeng Ji, David Evans

Large language models (LLMs) are known to perpetuate stereotypes and exhibit biases. Various strategies have been proposed to mitigate these biases, but most work studies biases in LLMs as a black-box problem without con…

How far can bias go? -- Tracing bias from pretraining data to alignment

2024-11-28 · Marion Thaler, Abdullatif Köksal, Alina Leidinger, Anna Korhonen 외

As LLMs are increasingly integrated into user-facing applications, addressing biases that perpetuate societal inequalities is crucial. While much work has gone into measuring or mitigating biases in these models, fewer s…

Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs

2025-07-09 · Itay Itzhak, Yonatan Belinkov, Gabriel Stanovsky arxiv

Large language models (LLMs) exhibit cognitive biases -- systematic tendencies of irrational decision-making, similar to those seen in humans. Prior work has found that these biases vary across models and can be amplifie…

Mitigating Label Biases for In-context Learning

2023-05-28 · Yu Fei, Yifan Hou, Zeming Chen, Antoine Bosselut

Various design settings for in-context learning (ICL), such as the choice and order of the in-context examples, can bias a model toward a particular prediction without being reflective of an understanding of the task. Wh…

In-Context Learningtext-classificationText Classification

A Multi-Agent Framework for Mitigating Dialect Biases in Privacy Policy Question-Answering Systems

2025-06-03 · Đorđe Klisura, Astrid R Bernaga Torres, Anna Karen Gárate-Escamilla, Rajesh Roshan Biswal 외

Privacy policies inform users about data collection and usage, yet their complexity limits accessibility for diverse populations. Existing Privacy Policy Question Answering (QA) systems exhibit performance disparities ac…

Question Answering