paper-with-me

홈 › Papers

Exposing Diversity Bias in Deep Generative Models: Statistical Origins and Correction of Diversity Error

2026-02-16 · Farzan Farnia, Mohammad Jalali, Azim Ospanov arxiv

Deep generative models have achieved great success in producing high-quality samples, making them a central tool across machine learning applications. Beyond sample quality, an important yet less systematically studied question is whether trained generative models faithfully capture the diversity of the underlying data distribution. In this work, we address this question by directly comparing the diversity of samples generated by state-of-the-art models with that of test samples drawn from the target data distribution, using recently proposed reference-free entropy-based diversity scores, Vendi and RKE. Across multiple benchmark datasets, we find that test data consistently attains substantially higher Vendi and RKE diversity scores than the generated samples, suggesting a systematic downward diversity bias in modern generative models. To understand the origin of this bias, we analyze the finite-sample behavior of entropy-based diversity scores and show that their expected values increase with sample size, implying that diversity estimated from finite training sets could inherently underestimate the diversity of the true distribution. As a result, optimizing the generators to minimize divergence to empirical data distributions would induce a loss of diversity. Finally, we discuss potential diversity-aware regularization and guidance strategies based on Vendi and RKE as principled directions for mitigating this bias, and provide empirical evidence suggesting their potential to improve the results.

📄 PDF Abstract BibTeX arXiv:2602.14682

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Theoretical and Empirical Insights into the Origins of Degree Bias in Graph Neural Networks

2024-04-04 · Arjun Subramonian, Jian Kang, Yizhou Sun

Graph Neural Networks (GNNs) often perform better for high-degree nodes than low-degree nodes on node classification tasks. This degree bias can reinforce social marginalization by, e.g., privileging celebrities and othe…

Node Classification

On the Origins of Bias in NLP through the Lens of the Jim Code

2023-05-16 · Fatma Elsafoury, Gavin Abercrombie

In this paper, we trace the biases in current natural language processing (NLP) models back to their origins in racism, sexism, and homophobia over the last 500 years. We review literature from critical race theory, gend…

Ethics

Bridging Semantic Understanding and Popularity Bias with LLMs

2026-01-14 · Renqiang Luo, Dong Zhang, Yupeng Gao, Wen Shi 외 arxiv

Semantic understanding of popularity bias is a crucial yet underexplored challenge in recommender systems, where popular items are often favored at the expense of niche content. Most existing debiasing methods treat the …

How did prebiotic polymers become informational foldamers?

2016-04-28

A mystery about the origins of life is which molecular structures $-$ and what spontaneous processes $-$ drove the autocatalytic transition from simple chemistry to biology? Using the HP lattice model of polymer sequence…

Predictive Biases in Natural Language Processing Models: A Conceptual Framework and Overview

2019-11-09 · ACL 2020 6 · Deven Shah, H. Andrew Schwartz, Dirk Hovy

An increasing number of works in natural language processing have addressed the effect of bias on the predicted outcomes, introducing mitigation techniques that act on different parts of the standard NLP pipeline (data a…

Selection bias