paper-with-me

홈 › Papers

Leveraging Computerized Adaptive Testing for Cost-effective Evaluation of Large Language Models in Medical Benchmarking

2026-02-28 · Tianpeng Zheng, Zhehan Jiang, Jiayi Liu, Shicong Feng arxiv

The rapid proliferation of large language models (LLMs) in healthcare creates an urgent need for scalable and psychometrically sound evaluation methods. Conventional static benchmarks are costly to administer repeatedly, vulnerable to data contamination, and lack calibrated measurement properties for fine-grained performance tracking. We propose and validate a computerized adaptive testing (CAT) framework grounded in item response theory (IRT) for efficient assessment of standardized medical knowledge in LLMs. The study comprises a two-phase design: a Monte Carlo simulation to identify optimal CAT configurations and an empirical evaluation of 38 LLMs using a human-calibrated medical item bank. Each model completed both the full item bank and an adaptive test that dynamically selected items based on real-time ability estimates and terminated upon reaching a predefined reliability threshold (standard error <= 0.3). Results show that CAT-derived proficiency estimates achieved a near-perfect correlation with full-bank estimates (r = 0.988) while using only 1.3 percent of the items. Evaluation time was reduced from several hours to minutes per model, with substantial reductions in token usage and computational cost, while preserving inter-model performance rankings. This work establishes a psychometric framework for rapid, low-cost benchmarking of foundational medical knowledge in LLMs. The proposed adaptive methodology is intended as a standardized pre-screening and continuous monitoring tool and is not a substitute for real-world clinical validation or safety-oriented prospective studies.

📄 PDF Abstract BibTeX arXiv:2603.23506

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Balancing Test Accuracy and Security in Computerized Adaptive Testing

2023-05-18 · Wanyong Feng, Aritra Ghosh, Stephen Sireci, Andrew S. Lan

Computerized adaptive testing (CAT) is a form of personalized testing that accurately measures students' knowledge levels while reducing test length. Bilevel optimization-based CAT (BOBCAT) is a recent framework that lea…

Bilevel OptimizationQuestion Selection

An Intelligent Testing Strategy for Vocabulary Assessment of Chinese Second Language Learners

2019-08-01 · WS 2019 8 · Wei Zhou, Renfen Hu, Feipeng Sun, Ronghuai Huang

Vocabulary is one of the most important parts of language competence. Testing of vocabulary knowledge is central to research on reading and language. However, it usually costs a large amount of time and human labor to bu…

BOBCAT: Bilevel Optimization-Based Computerized Adaptive Testing

2021-08-17 · Aritra Ghosh, Andrew Lan

Computerized adaptive testing (CAT) refers to a form of tests that are personalized to every student/test taker. CAT methods adaptively select the next most informative question/item for each student given their response…

Bilevel OptimizationQuestion Selection

Addressing Selection Bias in Computerized Adaptive Testing: A User-Wise Aggregate Influence Function Approach

2023-08-23 · Soonwoo Kwon, Sojung Kim, SeungHyun Lee, Jin-Young Kim 외

Computerized Adaptive Testing (CAT) is a widely used, efficient test mode that adapts to the examinee's proficiency level in the test domain. CAT requires pre-trained item profiles, for CAT iteratively assesses the stude…

Diagnosticparameter estimationSelection bias

Survey of Computerized Adaptive Testing: A Machine Learning Perspective

2024-03-31 · Qi Liu, Yan Zhuang, Haoyang Bi, Zhenya Huang 외

Computerized Adaptive Testing (CAT) provides an efficient and tailored method for assessing the proficiency of examinees, by dynamically adjusting test questions based on their performance. Widely adopted across diverse …

cognitive diagnosisQuestion SelectionSociologySurvey