paper-with-me

Papers

Improving the Validity and Practical Usefulness of AI/ML Evaluations Using an Estimands Framework

2024-06-14 · Olivier Binette, Jerome P. Reiter

Commonly, AI or machine learning (ML) models are evaluated on benchmark datasets. This practice supports innovative methodological research, but benchmark performance can be poorly correlated with performance in real-world applications -- a construct validity issue. To improve the validity and practical usefulness of evaluations, we propose using an estimands framework adapted from international clinical trials guidelines. This framework provides a systematic structure for inference and reporting in evaluations, emphasizing the importance of a well-defined estimation target. We illustrate our proposal on examples of commonly used evaluation methodologies - involving cross-validation, clustering evaluation, and LLM benchmarking - that can lead to incorrect rankings of competing models (rank reversals) with high probability, even when performance differences are large. We demonstrate how the estimands framework can help uncover underlying issues, their causes, and potential solutions. Ultimately, we believe this framework can improve the validity of evaluations through better-aligned inference, and help decision-makers and model users interpret reported results more effectively.

📄 PDF Abstract BibTeX arXiv:2406.10366

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

Quantifying the Internal Validity of Weighted Estimands

2024-04-22 · Alexandre Poirier, Tymon Słoczyński

In this paper we study a class of weighted estimands, which we define as parameters that can be expressed as weighted averages of the underlying heterogeneous treatment effects. The popular ordinary least squares (OLS), …

Diagnostic

Learning sources of variability from high-dimensional observational studies

2023-07-26 · Eric W. Bridgeford, Jaewon Chung, Brian Gilbert, Sambit Panda 외

Causal inference studies whether the presence of a variable influences an observed outcome. As measured by quantities such as the "average treatment effect," this paradigm is employed across numerous biological fields, f…

Causal Inference

Statistical Issues and Recommendations for Clinical Trials Conducted During the COVID-19 Pandemic

2020-05-21 · R. Daniel Meyer, Bohdana Ratitch, Marcel Wolbers, Olga Marchenko 외

The COVID-19 pandemic has had and continues to have major impacts on planned and ongoing clinical trials. Its effects on trial data create multiple potential statistical issues. The scale of impact is unprecedented, but …

Preference-based Conditional Treatment Effects and Policy Learning

2026-02-03 · Dovid Parnas, Mathieu Even, Julie Josse, Uri Shalit arxiv

We introduce a new preference-based framework for conditional treatment effect estimation and policy learning, built on the Conditional Preference-based Treatment Effect (CPTE). CPTE requires only that outcomes be ranked…

Automatic Debiased Machine Learning for Smooth Functionals of Nonparametric M-Estimands

2025-01-21 · Lars van der Laan, Aurelien Bibaut, Nathan Kallus, Alex Luedtke

We propose a unified framework for automatic debiased machine learning (autoDML) to perform inference on smooth functionals of infinite-dimensional M-estimands, defined as population risk minimizers over Hilbert spaces. …

Causal InferenceModel Selectionvalid