paper-with-me

홈 › Papers

Evaluating the Diversity of AI-Generated Content with Diversity Profiles

2026-08-18 · Xiuyuan Hu, Xuege Hou, Guoqing Liu, Yang Zhao, Jieran Li, Dongbiao Sun, José Miguel Hernández-Lobato, Hao Zhang, Xue Liu arxiv

Diversity is a fundamental criterion for evaluating generative artificial intelligence (AI) systems, yet its measurement remains inherently ambiguous. Existing approaches typically represent generated samples in an embedding space, compute pairwise distances or similarities, and aggregate them into a single scalar score. Such scalar summaries are convenient, but they often encode different inductive biases and may yield contradictory rankings of the same sample sets. In this paper, we argue that diversity evaluation for AI-generated content is intrinsically under-specified when reduced to a single number. We first review representative diversity metrics, and then diagnose their limitations from two complementary perspectives: an axiomatic analysis showing that no representative scalar metric satisfies all desirable properties simultaneously, and an empirical analysis showing that high-dimensional representation spaces can induce concentrated, modality-dependent distance distributions. To address these issues, we propose diversity profiles: curve-valued, condition-aware summaries that evaluate a parameterized diversity family across a range of thresholds, scales, exponents, or orders under a specified representation and distance or kernel function. Diversity profiles reveal whether a comparison is robust across resolutions or instead depends on an arbitrary parameter choice. We instantiate profiles for several representative metric families and demonstrate their practical use in generative AI evaluation. Overall, diversity profiles provide a more transparent and resolution-aware framework for comparing the diversity of AI-generated content.

📄 PDF Abstract BibTeX arXiv:2608.17731

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluating the Evaluation of Diversity in Natural Language Generation

2020-04-06 · EACL 2021 2 · Guy Tevet, Jonathan Berant

Despite growing interest in natural language generation (NLG) models that produce diverse outputs, there is currently no principled method for evaluating the diversity of an NLG system. In this work, we propose a framewo…

DiversityText Generation

Evaluating the Diversity and Quality of LLM Generated Content

2025-04-16 · Alexander Shypula, Shuo Li, Botong Zhang, Vishakh Padmakumar 외

Recent work suggests that preference-tuning techniques--including Reinforcement Learning from Human Preferences (RLHF) methods like PPO and GRPO, as well as alternatives like DPO--reduce diversity, creating a dilemma giv…

DiversitySynthetic Data Generation

Enhancing Diversity of LLM-Generated Educational Tasks

2025-12-29 · Manh Hung Nguyen, Sebastian Tschiatschek, Adish Singla arxiv

Large language models (LLMs) have shown the potential for generating educational content at scale, assisting educators in creating practice tasks or synthesizing data for training educational models. However, LLMs suffer…

Benchmarking Linguistic Diversity of Large Language Models

2024-12-13 · Yanzhu Guo, Guokan Shang, Chloé Clavel

The development and evaluation of Large Language Models (LLMs) has primarily focused on their task-solving capabilities, with recent models even surpassing human performance in some areas. However, this focus often negle…

BenchmarkingDiversityText Generation

The Procedural Content Generation Benchmark: An Open-source Testbed for Generative Challenges in Games

2025-03-27 · Ahmed Khalifa, Roberto Gallotta, Matthew Barthet, Antonios Liapis 외

This paper introduces the Procedural Content Generation Benchmark for evaluating generative algorithms on different game content creation tasks. The benchmark comes with 12 game-related problems with multiple variants on…

Diversity