paper-with-me

홈 › Papers

What is Your Metric Telling You? Evaluating Classifier Calibration under Context-Specific Definitions of Reliability

2022-05-23 · John Kirchenbauer, Jacob Oaks, Eric Heim

Classifier calibration has received recent attention from the machine learning community due both to its practical utility in facilitating decision making, as well as the observation that modern neural network classifiers are poorly calibrated. Much of this focus has been towards the goal of learning classifiers such that their output with largest magnitude (the "predicted class") is calibrated. However, this narrow interpretation of classifier outputs does not adequately capture the variety of practical use cases in which classifiers can aid in decision making. In this work, we argue that more expressive metrics must be developed that accurately measure calibration error for the specific context in which a classifier will be deployed. To this end, we derive a number of different metrics using a generalization of Expected Calibration Error (ECE) that measure calibration error under different definitions of reliability. We then provide an extensive empirical evaluation of commonly used neural network architectures and calibration techniques with respect to these metrics. We find that: 1) definitions of ECE that focus solely on the predicted class fail to accurately measure calibration error under a selection of practically useful definitions of reliability and 2) many common calibration techniques fail to improve calibration performance uniformly across ECE metrics derived from these diverse definitions of reliability.

📄 PDF Abstract BibTeX arXiv:2205.11454

Code (0)

등록된 구현이 없습니다.

Tasks

Classifier calibrationDecision Making

Similar Papers 제목 키워드 기반

A Personal Storytelling about Your Favorite Data

2015-09-01 · WS 2015 9 · Cyril Labb{\'e}, Claudia Roncancio, Damien Bras
Text Generation

Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition

2024-07-05 · Aditya K Surikuchi, Raquel Fernández, Sandro Pezzelle

Visual storytelling consists in generating a natural language story given a temporally ordered sequence of images. This task is not only challenging for models, but also very difficult to evaluate with automatic metrics …

Visual GroundingVisual Storytelling

What's Wrong with Your Synthetic Tabular Data? Using Explainable AI to Evaluate Generative Models

2025-04-29 · Jan Kapar, Niklas Koenen, Martin Jullum

Evaluating synthetic tabular data is challenging, since they can differ from the real data in so many ways. There exist numerous metrics of synthetic data quality, ranging from statistical distances to predictive perform…

counterfactualFeature ImportanceSynthetic Data Evaluation

RoViST:Learning Robust Metrics for Visual Storytelling

2022-05-08 · Eileen Wang, Caren Han, Josiah Poon

Visual storytelling (VST) is the task of generating a story paragraph that describes a given image sequence. Most existing storytelling approaches have evaluated their models using traditional natural language generation…

SentenceText GenerationVisual GroundingVisual Storytelling

RoViST: Learning Robust Metrics for Visual Storytelling

2022-07-01 · Findings (NAACL) 2022 7 · Eileen Wang, Caren Han, Josiah Poon

Visual storytelling (VST) is the task of generating a story paragraph that describes a given image sequence. Most existing storytelling approaches have evaluated their models using traditional natural language generation…

SentenceText GenerationVisual GroundingVisual Storytelling