paper-with-me

홈 › Papers

Evaluating the Pre-Consultation Ability of LLMs using Diagnostic Guidelines

2026-01-07 · Jean Seo, Gibaeg Kim, Kihun Shin, Seungseop Lim, Hyunkyung Lee, Wooseok Han, Jongwon Lee, Eunho Yang arxiv

We introduce EPAG, a benchmark dataset and framework designed for Evaluating the Pre-consultation Ability of LLMs using diagnostic Guidelines. LLMs are evaluated directly through HPI-diagnostic guideline comparison and indirectly through disease diagnosis. In our experiments, we observe that small open-source models fine-tuned with a well-curated, task-specific dataset can outperform frontier LLMs in pre-consultation. Additionally, we find that increased amount of HPI (History of Present Illness) does not necessarily lead to improved diagnostic performance. Further experiments reveal that the language of pre-consultation influences the characteristics of the dialogue. By open-sourcing our dataset and evaluation pipeline on https://github.com/seemdog/EPAG, we aim to contribute to the evaluation and further development of LLM applications in real-world clinical settings.

📄 PDF Abstract BibTeX arXiv:2601.03627

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DLawBench: Evaluating LLMs Through Multi-Turn Legal Consultation

2026-06-11 · Li Zhang, Yuzhen Shi, Yiran Hu, Jingwen Zhang 외 arxiv

Lawyer-client consultation is a critical starting point for legal services. Effective legal assistance hinges on eliciting sufficient and truthful information from clients in order to devise strategies that best protect …

Legal Reasoning

LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis

2026-02-10 · Shihao Xu, Tiancheng Zhou, Jiatong Ma, Yanli Ding 외 arxiv

Mental disorders are highly prevalent worldwide, but the shortage of psychiatrists and the inherent subjectivity of interview-based diagnosis create substantial barriers to timely and consistent mental-health assessment.…

LexRubric: A Rubric-Guided Diagnostic Benchmark for Open-Ended Legal Tasks

2026-06-08 · Yifan Chen, Haitao Li, Yiran Hu, Kaisong Song 외 arxiv

As large language models (LLMs) are increasingly applied to real-world legal tasks, evaluating the reliability of their open-ended legal responses has become essential. These tasks require context-sensitive answers and a…

Legal Reasoning

AIPatient Arena: EHR-grounded evaluation of large language models in end-to-end clinical consultation workflows

2026-06-16 · Jiahui Niu, Huizi Yu, Wenkong Wang, Guangxin Dai 외 arxiv

Large language models (LLMs) are increasingly considered for use in clinical consultation tasks, yet most medical evaluations remain static, single-turn, or narrowly outcome-based, limiting their ability to reflect the s…

Knowledge Graphs

Are Multimodal LLMs Ready for Clinical Dermatology? A Real-World Evaluation in Dermatology

2026-05-01 · Roy Jiang, Hyunjae Kim, Zhenyue Qin, Morten Lee 외 arxiv

Multimodal large language models (MLLMs) have demonstrated promise on publicly available dermatology benchmarks. However, benchmark performance may not generalize to real-world dermatologic decision-making. To quantify t…