paper-with-me

홈 › Papers

Medical Large Language Model Benchmarks Should Prioritize Construct Validity

2025-03-12 · Ahmed Alaa, Thomas Hartvigsen, Niloufar Golchini, Shiladitya Dutta, Frances Dean, Inioluwa Deborah Raji, Travis Zack

Medical large language models (LLMs) research often makes bold claims, from encoding clinical knowledge to reasoning like a physician. These claims are usually backed by evaluation on competitive benchmarks; a tradition inherited from mainstream machine learning. But how do we separate real progress from a leaderboard flex? Medical LLM benchmarks, much like those in other fields, are arbitrarily constructed using medical licensing exam questions. For these benchmarks to truly measure progress, they must accurately capture the real-world tasks they aim to represent. In this position paper, we argue that medical LLM benchmarks should (and indeed can) be empirically evaluated for their construct validity. In the psychological testing literature, "construct validity" refers to the ability of a test to measure an underlying "construct", that is the actual conceptual target of evaluation. By drawing an analogy between LLM benchmarks and psychological tests, we explain how frameworks from this field can provide empirical foundations for validating benchmarks. To put these ideas into practice, we use real-world clinical data in proof-of-concept experiments to evaluate popular medical LLM benchmarks and report significant gaps in their construct validity. Finally, we outline a vision for a new ecosystem of medical LLM evaluation centered around the creation of valid benchmarks.

📄 PDF Abstract BibTeX arXiv:2503.10694

Code (0)

등록된 구현이 없습니다.

Tasks

Clinical KnowledgeLanguage ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

Eir: Thai Medical Large Language Models

2024-09-13 · Yutthakorn Thiprak, Rungtam Ngodngamthaweesuk, Songtam Ngodngamtaweesuk

We present Eir-8B, a large language model with 8 billion parameters, specifically designed to enhance the accuracy of handling medical tasks in the Thai language. This model focuses on providing clear and easy-to-underst…

Language ModellingLarge Language ModelMedQAMMLU

A data- and compute-efficient chest X-ray foundation model beyond aggressive scaling

2026-02-26 · Chong Wang, Yabin Zhang, Yunhe Gao, Maya Varma 외 arxiv

Foundation models for medical imaging are typically pretrained on increasingly large datasets, following a "scale-at-all-costs" paradigm. However, this strategy faces two critical challenges: large-scale medical datasets…

Representation LearningSemantic Segmentation

Harnessing Large Language Models for Biomedical Named Entity Recognition

2025-12-28 · Jian Chen, Leilei Su, Cong Sun arxiv

Background and Objective: Biomedical Named Entity Recognition (BioNER) is a foundational task in medical informatics, crucial for downstream applications like drug discovery and clinical trial matching. However, adapting…

Drug Discovery

Not Every AI Problem is a Data Problem: We Should Be Intentional About Data Scaling

2025-01-23 · Tanya Rodchenko, Natasha Noy, Nino Scherrer, Jennifer Prendki

While Large Language Models require more and more data to train and scale, rather than looking for any data to acquire, we should consider what types of tasks are more likely to benefit from data scaling. We should be in…

Repetition over Diversity: High-Signal Data Filtering for Sample-Efficient German Language Modeling

2026-04-30 · Ansar Aynetdinov, Patrick Haller, Alan Akbik arxiv

Recent research has shown that filtering massive English web corpora into high-quality subsets significantly improves training efficiency. However, for high-resource non-English languages like German, French, or Japanese…