paper-with-me

홈 › Papers

Advances in LLM Reasoning Enable Flexibility in Clinical Problem-Solving

2026-01-17 · Kie Shidara, Preethi Prem, Jonathan Kim, Anna Podlasek, Feng Liu, Ahmed Alaa, Danilo Bernardo arxiv

Large Language Models (LLMs) have achieved high accuracy on medical question-answer (QA) benchmarks, yet their capacity for flexible clinical reasoning has been debated. Here, we asked whether advances in reasoning LLMs improve their cognitive flexibility in clinical reasoning. We assessed reasoning models from the OpenAI, Grok, Gemini, Claude, and DeepSeek families on the medicine abstraction and reasoning corpus (mARC), an adversarial medical QA benchmark which utilizes the Einstellung effect to induce inflexible overreliance on learned heuristic patterns in contexts where they become suboptimal. We found that strong reasoning models avoided Einstellung-based traps more often than weaker reasoning models, achieving human-level performance on mARC. On questions most commonly missed by physicians, the top 5 performing models answered 55% to 70% correctly with high confidence, indicating that these models may be less susceptible than humans to Einstellung effects. Our results indicate that strong reasoning models demonstrate improved flexibility in medical reasoning, achieving performance on par with humans on mARC.

📄 PDF Abstract BibTeX arXiv:2601.11866

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL

2026-01-14 · Yifei Shen, Yilun Zhao, Justice Ou, Tinglin Huang 외 arxiv

Real-world clinical text-to-SQL requires reasoning over heterogeneous EHR tables, temporal windows, and patient-similarity cohorts to produce executable queries. We introduce CLINSQL, a benchmark of 633 expert-annotated …

A Multi-granularity Concept Sparse Activation and Hierarchical Knowledge Graph Fusion Framework for Rare Disease Diagnosis

2025-07-11 · Mingda Zhang, Na Zhao, Jianglong Qin, Guoyu Ye 외 arxiv

Despite advances from medical large language models in healthcare, rare-disease diagnosis remains hampered by insufficient knowledge-representation depth, limited concept understanding, and constrained clinical reasoning…

Multimodal Model for Computational Pathology:Representation Learning and Image Compression

2026-03-19 · Peihang Wu, Zehong Chen, Lijian Xu arxiv

Whole slide imaging (WSI) has transformed digital pathology by enabling computational analysis of gigapixel histopathology images. Recent foundation model advances have accelerated progress in computational pathology, fa…

Representation LearningImage CompressionFew-Shot Learning

ArgMed-Agents: Explainable Clinical Decision Reasoning with LLM Disscusion via Argumentation Schemes

2024-03-10 · Shengxin Hong, Liang Xiao, Xin Zhang, Jianxia Chen

There are two main barriers to using large language models (LLMs) in clinical reasoning. Firstly, while LLMs exhibit significant promise in Natural Language Processing (NLP) tasks, their performance in complex reasoning …

ClinicalGPT-R1: Pushing reasoning capability of generalist disease diagnosis with large language model

2025-04-13 · Wuyang Lan, Wenzheng Wang, Changwei Ji, Guoxing Yang 외

Recent advances in reasoning with large language models (LLMs)has shown remarkable reasoning capabilities in domains such as mathematics and coding, yet their application to clinical diagnosis remains underexplored. Here…

DiagnosticLanguage ModelingLanguage ModellingLarge Language Model