paper-with-me

홈 › Papers

QuantiBias: Benchmarking Quantization-Induced Bias in LLMs

2026-07-23 · Emilio Ferrara arxiv

Almost every large language model that reaches a broad audience is quantized: trained in full precision, then compressed for efficiency. This step is assumed harmless and its safety is rarely re-checked. We find its principal side effect is increased bias that standard safety evaluation misses. Holding the model, its training, and the prompts fixed, a quantized model still refuses harmful requests, still avoids over-refusing benign prompts, and still selects the unbiased multiple-choice answer. Yet asked an open-ended question, the same model volunteers stereotypes in all eight languages we probe, in roughly one in four open-ended answers under an independent judge (~24% to ~27% across the compression ladder): it passes every standard check and still reaches users measurably more biased. The selective gap is a robust finding; whether open-ended bias further increases with compression is less certain, sensitive to the judge that scores it. We address both with \textbf{QuantiBias}, a benchmark that pairs a generative, multilingual stereotype probe with the refusal and multiple-choice controls that isolate open-ended generation, contrasts each build with and without reasoning, and rates the content severity of what it generates. Across two backbone models (Qwen and Gemma), a five-family screen, and eight benchmarks, quantizers allocate their extra precision by capability data that carries no bias-prevention signal, and reasoning before answering roughly halves the effect on some families while doing nothing on others. A quantized build must be re-evaluated for open-ended bias, not only on the short-form safeguards it already passes.

📄 PDF Abstract BibTeX arXiv:2607.21063

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Decomposing MXFP4 quantization error for LLM reinforcement learning: reducible bias, recoverable deadzone, and an irreducible floor

2026-05-19 · Xiaocan Li, Shiliang Wu, Zheng Shen arxiv

MXFP4 arithmetic can dramatically accelerate reinforcement learning (RL) post-training of large language models (LLMs), yet the quantization error introduces severe accuracy degradation. Existing work treats the quantiza…

Reinforcement Learning

Rethinking Perplexity: Revealing the Impact of Input Length on Perplexity Evaluation in LLMs

2026-02-04 · Letian Cheng, Junyan Wang, Yan Gao, Elliott Wen 외 arxiv

Perplexity is a widely adopted metric for assessing the predictive quality of large language models (LLMs) and often serves as a reference metric for downstream evaluations. However, recent evidence shows that perplexity…

Debias-SparseGPT: Bias-Aware Pruning for Large Language Models

2026-09-02 · Irina Proskurina, Guillaume Metzler, Antoine Gourru, Julien Velcin hf

Model compression techniques such as pruning and quantization facilitate the efficient deployment and acceleration of Large Language Models (LLMs). However, recent studies show that weight sparsification methods, such as…

Computational EfficiencyModel Compression

Does Compression Preserve Uncertainty? A Unified Benchmark for Quantized and Sparse LLMs via Conformal Prediction

2026-06-01 · Yujia Tong, Yuxi Wang, Yunyang Wan, Tian Zhang 외 arxiv

Model compression techniques such as quantization and pruning are widely used to reduce the deployment cost of large language models (LLMs), with existing evaluations focusing almost exclusively on accuracy preservation.…

Model Compression

Investigating Social Bias Changes in Quantized Language Models

2026-02-05 · Stanley Z. Hua, Sanae Lotfi, Irene Y. Chen arxiv

Post-training quantization reduces the memory needed to run large language models but alters their social biases in ways that aggregate metrics fail to capture. We present the first large-scale study of 50 quantized mode…