paper-with-me

홈 › Papers

Principles and Guidelines for Randomized Controlled Trials in AI Evaluation

2026-05-03 · Christopher Kelly, Angelica Chowdhury, Alexandra Campili, Bimpe Ayoola, Devin Barbour, Thomas Chen Dawson, Ze Shen Chin, Rokas Gipiškis arxiv

This work establishes a foundational framework for standardizing AI evaluation RCTs (sometimes called human uplift studies). Drawing on established experimental practices from disciplines with established RCT traditions, including software engineering, economics, clinical and health sciences, and psychology, we adopt the (Shadish et al., 2002) four-validity framework and extend it with a fifth principle on transparency, repeatability, and verification adapted from the Transparency and Openness Promotion (TOP) Guidelines (Center for Open Science, 2025). We operationalize all five principles into 33 guidelines adapted for AI evaluation RCT contexts, expressed as requirements with rationales, implementation instructions, and evidence bases. We position the principles and guidelines as serving three key roles for AI evaluation RCTs: a design tool for planning studies, an evaluation rubric for assessing existing work, and a blueprint for standard setting as the field converges on norms. Our framework extends prior work by centering evaluation on human performance rather than model output alone, formalizing causal inference through RCT methodology for AI contexts, integrating heterogeneity analysis and practical significance assessment, implementing a graded transparency and repeatability framework, and addressing AI-specific challenges including model versioning, human-AI interaction dynamics, contamination and spillover effects, and equitable impact assessment.

📄 PDF Abstract BibTeX arXiv:2605.02050

Code (0)

등록된 구현이 없습니다.

Tasks

Causal Inference

Similar Papers 제목 키워드 기반

Evaluating the Ability of Large Language Models to Identify Adherence to CONSORT Reporting Guidelines in Randomized Controlled Trials: A Methodological Evaluation Study

2025-11-17 · Zhichao He, Mouxiao Bian, Jianhong Zhu, Jiayuan Chen 외 arxiv

The Consolidated Standards of Reporting Trials statement is the global benchmark for transparent and high-quality reporting of randomized controlled trials. Manual verification of CONSORT adherence is a laborious, time-i…

Machine Learning Assisted Adjustment Boosts Efficiency of Exact Inference in Randomized Controlled Trials

2024-03-05 · Han Yu, Alan D. Hutson, Xiaoyi Ma

In this work, we proposed a novel inferential procedure assisted by machine learning based adjustment for randomized control trials. The method was developed under the Rosenbaum's framework of exact tests in randomized e…

Comparison of Methods that Combine Multiple Randomized Trials to Estimate Heterogeneous Treatment Effects

2023-03-28 · Carly Lupton Brantner, Trang Quynh Nguyen, Tengjie Tang, Congwen Zhao 외

Individualized treatment decisions can improve health outcomes, but using data to make these decisions in a reliable, precise, and generalizable way is challenging with a single dataset. Leveraging multiple randomized co…

Towards Understanding of Medical Randomized Controlled Trials by Conclusion Generation

2019-10-03 · WS 2019 11 · Alexander Te-Wei Shieh, Yung-Sung Chuang, Shang-Yu Su, Yun-Nung Chen

Randomized controlled trials (RCTs) represent the paramount evidence of clinical medicine. Using machines to interpret the massive amount of RCTs has the potential of aiding clinical decision-making. We propose a RCT con…

Decision MakingLanguage ModelingLanguage ModellingSentence+1

Optimality of Matched-Pair Designs in Randomized Controlled Trials

2022-06-15 · Yuehao Bai

In randomized controlled trials (RCTs), treatment is often assigned by stratified randomization. I show that among all stratified randomization schemes which treat all units with probability one half, a certain matched-p…