paper-with-me

홈 › Papers

A Theoretical Framework for Statistical Evaluability of Generative Models

2026-04-07 · Shashaank Aiyer, Yishay Mansour, Shay Moran, Han Shao arxiv

Statistical evaluation aims to estimate the generalization performance of a model using held-out i.i.d. test data sampled from the ground-truth distribution. In supervised learning settings such as classification, performance metrics such as error rate are well-defined, and test error reliably approximates population error given sufficiently large datasets. In contrast, evaluation is more challenging for generative models due to their open-ended nature: it is unclear which metrics are appropriate and whether such metrics can be reliably evaluated from finite samples. In this work, we introduce a theoretical framework for evaluating generative models and establish evaluability results for commonly used metrics. We study two categories of metrics: test-based metrics, including integral probability metrics (IPMs), and Rényi divergences. We show that IPMs with respect to any bounded test class can be evaluated from finite samples up to multiplicative and additive approximation errors. Moreover, when the test class has finite fat-shattering dimension, IPMs can be evaluated with arbitrary precision. In contrast, Rényi and KL divergences are not evaluable from finite samples, as their values can be critically determined by rare events. We also analyze the potential and limitations of perplexity as an evaluation method.

📄 PDF Abstract BibTeX arXiv:2604.05324

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Aggregating Incomplete Rankings

2024-02-26 · Yasunori Okumura

This study considers the method to derive a ranking of alternatives by aggregating the rankings submitted by several individuals who may not evaluate all of them. The collection of subsets of alternatives that individual…

The AI Evaluability Gap: The Missing Layer for Managing Risk and Sustaining Value

2026-06-19 · Vishal Srivastava, Tanmay Sah arxiv

Organizations deploying AI face two fundamental governance challenges: managing AI risk and sustaining AI value. Both depend on evidence whose sufficiency cannot be taken for granted. We call the shared underlying challe…

A Corpus of eRulemaking User Comments for Measuring Evaluability of Arguments

2018-05-01 · LREC 2018 5 · Joonsuk Park, Claire Cardie
Argument Mining

Scale-Sensitive Shattering: Learnability and Evaluability at Optimal Scale

2026-05-13 · Shashaank Aiyer, Yishay Mansour, Shay Moran, Han Shao 외 arxiv

We study the optimal scale at which real-valued function classes exhibit uniform convergence and learnability. Our main result establishes a scale-sensitive generalization of the fundamental theorem of PAC learning: for …

Generative Models and Statistical Validation

2026-05-28 · Sascha Diefenbacher, Sofia Palacios Schweitzer, Gregor Kasieczka arxiv

Generative machine learning has become an essential tool in theoretical and experimental physics, especially in the context of fast surrogates and density estimators. In this work, we first introduce the underlying frame…