paper-with-me

홈 › Papers

A Scalable Framework for Evaluating Health Language Models

2025-03-30 · Neil Mallinar, A. Ali Heydari, Xin Liu, Anthony Z. Faranesh, Brent Winslow, Nova Hammerquist, Benjamin Graef, Cathy Speed, Mark Malhotra, Shwetak Patel, Javier L. Prieto, Daniel McDuff, Ahmed A. Metwally

Large language models (LLMs) have emerged as powerful tools for analyzing complex datasets. Recent studies demonstrate their potential to generate useful, personalized responses when provided with patient-specific health information that encompasses lifestyle, biomarkers, and context. As LLM-driven health applications are increasingly adopted, rigorous and efficient one-sided evaluation methodologies are crucial to ensure response quality across multiple dimensions, including accuracy, personalization and safety. Current evaluation practices for open-ended text responses heavily rely on human experts. This approach introduces human factors and is often cost-prohibitive, labor-intensive, and hinders scalability, especially in complex domains like healthcare where response assessment necessitates domain expertise and considers multifaceted patient data. In this work, we introduce Adaptive Precise Boolean rubrics: an evaluation framework that streamlines human and automated evaluation of open-ended questions by identifying gaps in model responses using a minimal set of targeted rubrics questions. Our approach is based on recent work in more general evaluation settings that contrasts a smaller set of complex evaluation targets with a larger set of more precise, granular targets answerable with simple boolean responses. We validate this approach in metabolic health, a domain encompassing diabetes, cardiovascular disease, and obesity. Our results demonstrate that Adaptive Precise Boolean rubrics yield higher inter-rater agreement among expert and non-expert human evaluators, and in automated assessments, compared to traditional Likert scales, while requiring approximately half the evaluation time of Likert-based methods. This enhanced efficiency, particularly in automated evaluation and non-expert contributions, paves the way for more extensive and cost-effective evaluation of LLMs in health.

📄 PDF Abstract BibTeX arXiv:2503.23339

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI

2024-12-17 · Deep Bhatt, Surya Ayyagari, Anuruddh Mishra

Diagnostic errors in healthcare persist as a critical challenge, with increasing numbers of patients turning to online resources for health information. While AI-powered healthcare chatbots show promise, there exists no …

BenchmarkingChatbotDiagnostic

Health-SCORE: Towards Scalable Rubrics for Improving Health-LLMs

2026-01-26 · Zhichao Yang, Sepehr Janghorbani, Dongxu Zhang, Jun Han 외 arxiv

Rubrics are essential for evaluating open-ended LLM responses, especially in safety-critical domains such as healthcare. However, creating high-quality and domain-specific rubrics typically requires significant human exp…

Reinforcement Learning

PsyEval: A Suite of Mental Health Related Tasks for Evaluating Large Language Models

2023-11-15 · Haoan Jin, Siyuan Chen, Dilawaier Dilixiati, Yewei Jiang 외

Evaluating Large Language Models (LLMs) in the mental health domain poses distinct challenged from other domains, given the subtle and highly subjective nature of symptoms that exhibit significant variability among indiv…

Language ModellingLarge Language ModelModel Optimization

CSEval: A Framework for Evaluating Clinical Semantics in Text-to-Image Generation

2026-02-12 · Robert Cronshaw, Konstantinos Vilouras, Junyu Yan, Yuning Du 외 arxiv

Text-to-image generation has been increasingly applied in medical domains for various purposes such as data augmentation and education. Evaluating the quality and clinical reliability of these generated images is essenti…

Text-to-Image GenerationData Augmentation

Adaptive Trust Metrics for Multi-LLM Systems: Enhancing Reliability in Regulated Industries

2026-01-07 · Tejaswini Bollikonda arxiv

Large Language Models (LLMs) are increasingly deployed in sensitive domains such as healthcare, finance, and law, yet their integration raises pressing concerns around trust, accountability, and reliability. This paper e…