paper-with-me

홈 › Papers

On the Statistical Complexity of Estimating VENDI Scores from Empirical Data

2024-10-29 · Azim Ospanov, Farzan Farnia

Reference-free evaluation metrics for generative models have recently been studied in the machine learning community. As a reference-free metric, the VENDI score quantifies the diversity of generative models using matrix-based entropy from information theory. The VENDI score is usually computed through the eigendecomposition of an $n \times n$ kernel matrix for $n$ generated samples. However, due to the high computational cost of eigendecomposition for large $n$, the score is often computed on sample sizes limited to a few tens of thousands. In this paper, we explore the statistical convergence of the VENDI score and demonstrate that for kernel functions with an infinite feature map dimension, the evaluated score for a limited sample size may not converge to the matrix-based entropy statistic. We introduce an alternative statistic called the $t$-truncated VENDI statistic. We show that the existing Nystr\"om method and the FKEA approximation method for the VENDI score will both converge to the defined truncated VENDI statistic given a moderate sample size. We perform several numerical experiments to illustrate the concentration of the empirical VENDI score around the truncated VENDI statistic and discuss how this statistic correlates with the visual diversity of image data.

📄 PDF Abstract BibTeX arXiv:2410.21719

Code (0)

등록된 구현이 없습니다.

Tasks

Diversity

Similar Papers 제목 키워드 기반

Cousins Of The Vendi Score: A Family Of Similarity-Based Diversity Metrics For Science And Machine Learning

2023-10-19 · Amey P. Pasarkar, Adji Bousso Dieng

Measuring diversity accurately is important for many scientific fields, including machine learning (ML), ecology, and chemistry. The Vendi Score was introduced as a generic similarity-based diversity metric that extends …

DiversityMemorizationSensitivity

Exposing Diversity Bias in Deep Generative Models: Statistical Origins and Correction of Diversity Error

2026-02-16 · Farzan Farnia, Mohammad Jalali, Azim Ospanov arxiv

Deep generative models have achieved great success in producing high-quality samples, making them a central tool across machine learning applications. Beyond sample quality, an important yet less systematically studied q…

Conditional Vendi Score: An Information-Theoretic Approach to Diversity Evaluation of Prompt-based Generative Models

2024-11-05 · Mohammad Jalali, Azim Ospanov, Amin Gohari, Farzan Farnia

Text-conditioned generation models are commonly evaluated based on the quality of the generated data and its alignment with the input text prompt. On the other hand, several applications of prompt-based generative models…

Diversity

Towards a Scalable Reference-Free Evaluation of Generative Models

2024-07-03 · Azim Ospanov, Jingwei Zhang, Mohammad Jalali, Xuenan Cao 외

While standard evaluation scores for generative models are mostly reference-based, a reference-dependent assessment of generative models could be generally difficult due to the unavailability of applicable reference data…

Diversity

Quality-Weighted Vendi Scores And Their Application To Diverse Experimental Design

2024-05-03 · Quan Nguyen, Adji Bousso Dieng

Experimental design techniques such as active search and Bayesian optimization are widely used in the natural sciences for data collection and discovery. However, existing techniques tend to favor exploitation over explo…

Bayesian OptimizationDiversityDrug DiscoveryExperimental Design