paper-with-me

Papers

MedRiskEval: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings

2025-07-09 · Jean-Philippe Corbeil, Minseon Kim, Maxime Griot, Sheela Agarwal, Alessandro Sordoni, Francois Beaulieu, Paul Vozila arxiv

As the performance of large language models (LLMs) continues to advance, their adoption in the medical domain is increasing. However, most existing risk evaluations largely focused on general safety benchmarks. In the medical applications, LLMs may be used by a wide range of users, ranging from general users and patients to clinicians, with diverse levels of expertise and the model's outputs can have a direct impact on human health which raises serious safety concerns. In this paper, we introduce MedRiskEval, a medical risk evaluation benchmark tailored to the medical domain. To fill the gap in previous benchmarks that only focused on the clinician perspective, we introduce a new patient-oriented dataset called PatientSafetyBench containing 466 samples across 5 critical risk categories. Leveraging our new benchmark alongside existing datasets, we evaluate a variety of open- and closed-source LLMs. To the best of our knowledge, this work establishes an initial foundation for safer deployment of LLMs in healthcare.

📄 PDF Abstract BibTeX arXiv:2507.07248

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts

2025-09-09 · Rochana Prih Hastuti, Rian Adam Rajagede, Mansour Al Ghanim, Mengxin Zheng 외 arxiv

As large language models (LLMs) are adapted to sensitive domains such as medicine, their fluency raises safety risks, particularly regarding provenance and accountability. Watermarking embeds detectable patterns to mitig…

Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications

2025-10-20 · Xiao Ye, Jacob Dineen, Zhaonan Li, Zhikun Xu 외 arxiv

Medical Large language models achieve strong scores on standard benchmarks; however, the transfer of those results to safe and reliable performance in clinical workflows remains a challenge. This survey reframes evaluati…

Beyond Accuracy: Risk-Sensitive Evaluation of Hallucinated Medical Advice

2026-02-07 · Savan Doshi arxiv

Large language models are increasingly being used in patient-facing medical question answering, where hallucinated outputs can vary widely in potential harm. However, existing hallucination standards and evaluation metri…

Question Answering

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

2025-08-06 · Wenting Chen, Guo Yu, Yiu-Fai Cheung, Meidan Ding 외 arxiv

Large language models (LLMs) show significant potential in healthcare, prompting numerous benchmarks to evaluate their capabilities. However, concerns persist regarding the reliability of these benchmarks, which often la…

Ensuring Safety and Trust: Analyzing the Risks of Large Language Models in Medicine

2024-11-20 · Yifan Yang, Qiao Jin, Robert Leaman, Xiaoyu Liu 외

The remarkable capabilities of Large Language Models (LLMs) make them increasingly compelling for adoption in real-world healthcare applications. However, the risks associated with using LLMs in medical applications have…

FairnessSafety Alignment