paper-with-me

홈 › Papers

How Quantization Shapes Bias in Large Language Models

2025-08-25 · Federico Marcuzzi, Xuefei Ning, Roy Schwartz, Iryna Gurevych arxiv

This work presents a comprehensive evaluation of how quantization affects model bias, with particular attention to its impact on individual demographic subgroups. We focus on weight and activation quantization strategies and examine their effects across a broad range of bias types, including stereotypes, fairness, toxicity, and sentiment. We employ both probability- and generated text-based metrics across 13 benchmarks and evaluate models that differ in architecture family and reasoning ability. Our findings show that quantization has a nuanced impact on bias: while it can reduce model toxicity and does not significantly impact sentiment, it tends to slightly increase stereotypes and unfairness in generative tasks, especially under aggressive compression. These trends are generally consistent across demographic categories and subgroups, and model types, although their magnitude depends on the specific setting. Overall, our results highlight the importance of carefully balancing efficiency and ethical considerations when applying quantization in practice.

📄 PDF Abstract BibTeX arXiv:2508.18088

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fair-GPTQ: Bias-Aware Quantization for Large Language Models

2025-09-18 · Irina Proskurina, Guillaume Metzler, Julien Velcin arxiv

The high memory demands of generative language models have drawn attention to quantization, which reduces memory usage by mapping model weights to lower-precision integers. However, recent empirical studies show that, wh…

Text Generation

Investigating Social Bias Changes in Quantized Language Models

2026-02-05 · Stanley Z. Hua, Sanae Lotfi, Irene Y. Chen arxiv

Post-training quantization reduces the memory needed to run large language models but alters their social biases in ways that aggregate metrics fail to capture. We present the first large-scale study of 50 quantized mode…

Softmax Bias Correction for Quantized Generative Models

2023-09-04 · Nilesh Prasad Pandey, Marios Fournarakis, Chirag Patel, Markus Nagel

Post-training quantization (PTQ) is the go-to compression technique for large generative models, such as stable diffusion or large language models. PTQ methods commonly keep the softmax activation in higher precision as …

Language ModelingLanguage ModellingQuantization

You Never Know: Quantization Induces Inconsistent Biases in Vision-Language Foundation Models

2024-10-26 · Eric Slyman, Anirudh Kanneganti, Sanghyun Hong, Stefan Lee

We study the impact of a standard practice in compressing foundation vision-language models - quantization - on the models' ability to produce socially-fair outputs. In contrast to prior findings with unimodal models tha…

Quantization

Understanding the Effect of Model Compression on Social Bias in Large Language Models

2023-12-09 · Gustavo Gonçalves, Emma Strubell

Large Language Models (LLMs) trained with self-supervision on vast corpora of web text fit to the social biases of that text. Without intervention, these social biases persist in the model's predictions in downstream tas…

Knowledge DistillationModel CompressionQuantization