paper-with-me

홈 › Papers

Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies

2025-07-05 · Mael Jullien, Marco Valentino, Leonardo Ranaldi, Andre Freitas arxiv

Recent works on large language models (LLMs) have demonstrated the impact of prompting strategies and fine-tuning techniques on their reasoning capabilities. Yet, their effectiveness on clinical natural language inference (NLI) remains underexplored. This study presents the first controlled evaluation of how prompt structure and efficient fine-tuning jointly shape model performance in clinical NLI. We inspect four classes of prompting strategies to elicit reasoning in LLMs at different levels of abstraction, and evaluate their impact on a range of clinically motivated reasoning types. For each prompting strategy, we construct high-quality demonstrations using a frontier model to distil multi-step reasoning capabilities into smaller models (4B parameters) via Low-Rank Adaptation (LoRA). Across different language models fine-tuned on the NLI4CT benchmark, we found that prompt type alone accounts for up to 44% of the variance in macro-F1. Moreover, LoRA fine-tuning yields consistent gains of +8 to 12 F1, raises output alignment above 97%, and narrows the performance gap to GPT-4o-mini to within 7.1%. Additional experiments on reasoning generalisation reveal that LoRA improves performance in 75% of the models on MedNLI and TREC Clinical Trials Track. Overall, these findings demonstrate that (i) prompt structure is a primary driver of clinical reasoning performance, (ii) compact models equipped with strong prompts and LoRA can rival frontier-scale systems, and (iii) reasoning-type-aware evaluation is essential to uncover prompt-induced trade-offs. Our results highlight the promise of combining prompt design and lightweight adaptation for more efficient and trustworthy clinical NLP systems, providing insights on the strengths and limitations of widely adopted prompting and parameter-efficient techniques in highly specialised domains.

📄 PDF Abstract BibTeX arXiv:2507.04142

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language Inference

Similar Papers 제목 키워드 기반

ClinicalGPT-R1: Pushing reasoning capability of generalist disease diagnosis with large language model

2025-04-13 · Wuyang Lan, Wenzheng Wang, Changwei Ji, Guoxing Yang 외

Recent advances in reasoning with large language models (LLMs)has shown remarkable reasoning capabilities in domains such as mathematics and coding, yet their application to clinical diagnosis remains underexplored. Here…

DiagnosticLanguage ModelingLanguage ModellingLarge Language Model

Reliability of Large Language Model Generated Clinical Reasoning in Assisted Reproductive Technology: Blinded Comparative Evaluation Study

2025-10-17 · Dou Liu, Ying Long, Sophia Zuoqiu, Di Liu 외 arxiv

Creating high-quality clinical Chains-of-Thought (CoTs) is crucial for explainable medical Artificial Intelligence (AI) while constrained by data scarcity. Although Large Language Models (LLMs) can synthesize medical dat…

A Vision-language Framework for Comparative Reasoning in Radiology

2026-06-04 · Tengfei Zhang, Ziheng Zhao, Xiaoman Zhang, Lisong Dai 외 arxiv

Medical imaging artificial intelligence has achieved strong performance in isolated image interpretation, but remains poorly aligned with radiological practice, where diagnosis and follow-up rely on comparison across pri…

Visual Question Answering

Dissecting Role Cognition in Medical LLMs via Neuronal Ablation

2025-10-28 · Xun Liang, Huayi Lai, Hanyu Wang, Wentao Zhang 외 arxiv

Large language models (LLMs) have gained significant traction in medical decision support systems, particularly in the context of medical question answering and role-playing simulations. A common practice, Prompt-Based R…

Question Answering

Multi-Task Training with In-Domain Language Models for Diagnostic Reasoning

2023-06-07 · Brihat Sharma, Yanjun Gao, Timothy Miller, Matthew M. Churpek 외

Generative artificial intelligence (AI) is a promising direction for augmenting clinical diagnostic decision support and reducing diagnostic errors, a leading contributor to medical errors. To further the development of …

DiagnosticLanguage ModelingLanguage Modelling