paper-with-me

홈 › Papers

Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-Tuning

2026-04-07 · Yanbei Jiang, Amr Keleg, Ryandito Diandaru, Jey Han Lau, Lea Frermann, Biaoyan Fang, Fajri Koto arxiv

While the real world is inherently stochastic, Large Language Models (LLMs) are predominantly evaluated on single-round inference against fixed ground truths. In this work, we shift the lens to distribution alignment: assessing whether LLMs, when prompted repeatedly, can generate outputs that adhere to a desired target distribution, e.g. reflecting real-world statistics or a uniform distribution. We formulate distribution alignment using the attributes of gender, race, and sentiment within occupational contexts. Our empirical analysis reveals that off-the-shelf LLMs and standard alignment techniques, including prompt engineering and Direct Preference Optimization, fail to reliably control output distributions. To bridge this gap, we propose a novel fine-tuning framework that couples Steering Token Calibration with Semantic Alignment. We introduce a hybrid objective function combining Kullback-Leibler divergence to anchor the probability mass of latent steering tokens and Kahneman-Tversky Optimization to bind these tokens to semantically consistent responses. Experiments across six diverse datasets demonstrate that our approach significantly outperforms baselines, achieving precise distributional control in attribute generation tasks.

📄 PDF Abstract BibTeX arXiv:2604.05756

Code (0)

등록된 구현이 없습니다.

Tasks

Prompt Engineering

Similar Papers 제목 키워드 기반

Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile Critics

2020-05-08 · ICML 2020 1 · Arsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry Vetrov

The overestimation bias is one of the major impediments to accurate off-policy learning. This paper investigates a novel way to alleviate the overestimation bias in a continuous control setting. Our method---Truncated Qu…

continuous-controlContinuous Control

Finetuning Text-to-Image Diffusion Models for Fairness

2023-11-11 · Xudong Shen, Chao Du, Tianyu Pang, Min Lin 외

The rapid adoption of text-to-image diffusion models in society underscores an urgent need to address their biases. Without interventions, these biases could propagate a skewed worldview and restrict opportunities for mi…

Fairness

Expert-Guided Extinction of Toxic Tokens for Debiased Generation

2024-05-29 · Xueyao Sun, Kaize Shi, Haoran Tang, Guandong Xu 외

Large language models (LLMs) can elicit social bias during generations, especially when inference with toxic prompts. Controlling the sensitive attributes in generation encounters challenges in data distribution, general…

FairnessRetrieval

disco: a toolkit for Distributional Control of Generative Models

2023-03-08 · Germán Kruszewski, Jos Rozen, Marc Dymetman

Pre-trained language models and other generative models have revolutionized NLP and beyond. However, these models tend to reproduce undesirable biases present in their training data. Also, they may overlook patterns that…

Can Prompt Modifiers Control Bias? A Comparative Analysis of Text-to-Image Generative Models

2024-06-09 · Philip Wootaek Shin, Jihyun Janice Ahn, Wenpeng Yin, Jack Sampson 외

It has been shown that many generative models inherit and amplify societal biases. To date, there is no uniform/systematic agreed standard to control/adjust for these biases. This study examines the presence and manipula…

DiversityEthicsImage GenerationPrompt Engineering+2