paper-with-me

홈 › Papers

The Validity Gap in Health AI Evaluation: A Cross-Sectional Analysis of Benchmark Composition

2026-03-18 · Alvin Rajkomar, Pavan Sudarshan, Angela Lai, Lily Peng arxiv

Background: Clinical trials rely on transparent inclusion criteria to ensure generalizability. In contrast, benchmarks validating health-related large language models (LLMs) rarely characterize the "patient" or "query" populations they contain. Without defined composition, aggregate performance metrics may misrepresent model readiness for clinical use. Methods: We analyzed 18,707 consumer health queries across six public benchmarks using LLMs as automated coding instruments to apply a standardized 16-field taxonomy profiling context, topic, and intent. Results: We identified a structural "validity gap." While benchmarks have evolved from static retrieval to interactive dialogue, clinical composition remains misaligned with real-world needs. Although 42% of the corpus referenced objective data, this was polarized toward wellness-focused wearable signals (17.7%); complex diagnostic inputs remained rare, including laboratory values (5.2%), imaging (3.8%), and raw medical records (0.6%). Safety-critical scenarios were effectively absent: suicide/self-harm queries comprised <0.7% of the corpus and chronic disease management only 5.5%. Benchmarks also neglected vulnerable populations (pediatrics/older adults <11%) and global health needs. Conclusions: Evaluation benchmarks remain misaligned with real-world clinical needs, lacking raw clinical artifacts, adequate representation of vulnerable populations, and longitudinal chronic care scenarios. The field must adopt standardized query profiling--analogous to clinical trial reporting--to align evaluation with the full complexity of clinical practice.

📄 PDF Abstract BibTeX arXiv:2603.18294

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Conformal Prediction Intervals with Temporal Dependence

2022-05-25 · Zhen Lin, Shubhendu Trivedi, Jimeng Sun

Cross-sectional prediction is common in many domains such as healthcare, including forecasting tasks using electronic health records, where different patients form a cross-section. We focus on the task of constructing va…

Conformal PredictionPredictionPrediction Intervalsregression+4

Evaluating Intersectional Fairness across Clinical Machine Learning Use Cases using Fairlogue and the All of Us Research Program

2026-04-07 · Nick Souligne, Vignesh Subbian arxiv

Intersectional biases in healthcare data can produce compound disparities in clinical machine learning models, yet most fairness evaluations assess demographic attributes independently. FairLogue, a toolkit for intersect…

Forward-Selected Panel Data Approach for Program Evaluation

2019-08-16 · Zhentao Shi, Jingyi Huang

Policy evaluation is central to economic data analysis, but economists mostly work with observational data in view of limited opportunities to carry out controlled experiments. In the potential outcome framework, the pan…

counterfactual

FairLogue: A Toolkit for Intersectional Fairness Analysis in Clinical Machine Learning Models

2026-04-06 · Nick Souligne, Vignesh Subbian arxiv

Objective: Algorithmic fairness is essential for equitable and trustworthy machine learning in healthcare. Most fairness tools emphasize single-axis demographic comparisons and may miss compounded disparities affecting i…

Unified and robust Lagrange multiplier type tests for cross-sectional independence in large panel data models

2023-02-28 · Zhenhong Huang, Zhaoyuan Li, Jianfeng Yao

This paper revisits the Lagrange multiplier type test for the null hypothesis of no cross-sectional dependence in large panel data models. We propose a unified test procedure and its power enhancement version, which show…