paper-with-me

홈 › Papers

Evaluating Large Language Models on Rare Disease Diagnosis: A Case Study using House M.D

2025-11-14 · Arsh Gupta, Ajay Narayanan Sridhar, Bonam Mingole, Amulya Yadav arxiv

Large language models (LLMs) have demonstrated capabilities across diverse domains, yet their performance on rare disease diagnosis from narrative medical cases remains underexplored. We introduce a novel dataset of 176 symptom-diagnosis pairs extracted from House M.D., a medical television series validated for teaching rare disease recognition in medical education. We evaluate four state-of-the-art LLMs such as GPT 4o mini, GPT 5 mini, Gemini 2.5 Flash, and Gemini 2.5 Pro on narrative-based diagnostic reasoning tasks. Results show significant variation in performance, ranging from 16.48% to 38.64% accuracy, with newer model generations demonstrating a 2.3 times improvement. While all models face substantial challenges with rare disease diagnosis, the observed improvement across architectures suggests promising directions for future development. Our educationally validated benchmark establishes baseline performance metrics for narrative medical reasoning and provides a publicly accessible evaluation framework for advancing AI-assisted diagnosis research.

📄 PDF Abstract BibTeX arXiv:2511.10912

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MIMIC-RD: Can LLMs differentially diagnose rare diseases in real-world clinical settings?

2025-12-18 · Zilal Eiz AlDin, John Wu, Jeffrey Paul Fung, Jennifer King 외 arxiv

Despite rare diseases affecting 1 in 10 Americans, their differential diagnosis remains challenging. Due to their impressive recall abilities, large language models (LLMs) have been recently explored for differential dia…

A specialized reasoning large language model for accelerating rare disease diagnosis: a randomized AI physician assistance trial

2026-06-23 · Haichao Chen, Songchi Zhou, Zhengyun Zhao, Shikai Hu 외 arxiv

Rare diseases affect millions of individuals worldwide, yet timely diagnosis remains a major public health challenge due to scarcity of specialized clinical expertise. While large language models (LLMs) show promise to s…

RareBench: Can LLMs Serve as Rare Diseases Specialists?

2024-02-09 · Xuanzhong Chen, Xiaohao Mao, Qihan Guo, Lun Wang 외

Generalist Large Language Models (LLMs), such as GPT-4, have shown considerable promise in various domains, including medical diagnosis. Rare diseases, affecting approximately 300 million people worldwide, often have uns…

DiagnosticMedical Diagnosis

AutoRD: An Automatic and End-to-End System for Rare Disease Knowledge Graph Construction Based on Ontologies-enhanced Large Language Models

2024-03-01 · Lang Cao, Jimeng Sun, Adam Cross

Rare diseases affect millions worldwide but often face limited research focus due to their low prevalence. This results in prolonged diagnoses and a lack of approved therapies. Recent advancements in Large Language Model…

graph constructionKnowledge GraphsManagementMedical Diagnosis+2

MMRareBench: A Rare-Disease Multimodal and Multi-Image Medical Benchmark

2026-04-12 · Junzhi Ning, Jiashi Lin, Yingying Fang, Wei Li 외 arxiv

Multimodal large language models (MLLMs) have advanced clinical tasks for common conditions, but their performance on rare diseases remains largely untested. In rare-disease scenarios, clinicians often lack prior clinica…

Clinical Knowledge