paper-with-me

홈 › Papers

Limitations of Large Language Models in Clinical Problem-Solving Arising from Inflexible Reasoning

2025-02-05 · Jonathan Kim, Anna Podlasek, Kie Shidara, Feng Liu, Ahmed Alaa, Danilo Bernardo

Large Language Models (LLMs) have attained human-level accuracy on medical question-answer (QA) benchmarks. However, their limitations in navigating open-ended clinical scenarios have recently been shown, raising concerns about the robustness and generalizability of LLM reasoning across diverse, real-world medical tasks. To probe potential LLM failure modes in clinical problem-solving, we present the medical abstraction and reasoning corpus (M-ARC). M-ARC assesses clinical reasoning through scenarios designed to exploit the Einstellung effect -- the fixation of thought arising from prior experience, targeting LLM inductive biases toward inflexible pattern matching from their training data rather than engaging in flexible reasoning. We find that LLMs, including current state-of-the-art o1 and Gemini models, perform poorly compared to physicians on M-ARC, often demonstrating lack of commonsense medical reasoning and a propensity to hallucinate. In addition, uncertainty estimation analyses indicate that LLMs exhibit overconfidence in their answers, despite their limited accuracy. The failure modes revealed by M-ARC in LLM medical reasoning underscore the need to exercise caution when deploying these models in clinical settings.

📄 PDF Abstract BibTeX arXiv:2502.04381

Code (0)

등록된 구현이 없습니다.

Tasks

ARC

Similar Papers 제목 키워드 기반

Knowledge Augmented Complex Problem Solving with Large Language Models: A Survey

2025-05-06 · Da Zheng, Lun Du, Junwei Su, Yuchen Tian 외

Problem-solving has been a fundamental driver of human progress in numerous domains. With advancements in artificial intelligence, Large Language Models (LLMs) have emerged as powerful tools capable of tackling complex p…

Mathematical Reasoning

Large Language Models as Students Who Think Aloud: Overly Coherent, Verbose, and Confident

2026-02-01 · Conrad Borchers, Jill-Jênn Vie, Roger Azevedo arxiv

Large language models (LLMs) are increasingly embedded in AI-based tutoring systems. Can they faithfully model novice reasoning and metacognitive judgments? Existing evaluations emphasize problem-solving accuracy, overlo…

GCoder: Improving Large Language Model for Generalized Graph Problem Solving

2024-10-24 · Qifan Zhang, Xiaobin Hong, Jianheng Tang, Nuo Chen 외

Large Language Models (LLMs) have demonstrated strong reasoning abilities, making them suitable for complex tasks such as graph computation. Traditional reasoning steps paradigm for graph problems is hindered by unverifi…

Language ModelingLanguage ModellingLarge Language Model

Emulating Human Cognitive Processes for Expert-Level Medical Question-Answering with Large Language Models

2023-10-17 · Khushboo Verma, Marina Moore, Stephanie Wottrich, Karla Robles López 외

In response to the pressing need for advanced clinical problem-solving tools in healthcare, we introduce BooksMed, a novel framework based on a Large Language Model (LLM). BooksMed uniquely emulates human cognitive proce…

Decision MakingLanguage ModelingLanguage ModellingLarge Language Model+2

Small Language Models for Emergency Departments Decision Support: A Benchmark Study

2025-10-05 · Zirui Wang, Jiajun Wu, Braden Teitge, Jessalyn Holodinsky 외 arxiv

Large language models (LLMs) have become increasingly popular in medical domains to assist physicians with a variety of clinical and operational tasks. Given the fast-paced and high-stakes environment of emergency depart…