paper-with-me

홈 › Papers

Prompt engineering does not universally improve Large Language Model performance across clinical decision-making tasks

2025-12-28 · Mengdi Chai, Ali R. Zomorrodi arxiv

Large Language Models (LLMs) have demonstrated promise in medical knowledge assessments, yet their practical utility in real-world clinical decision-making remains underexplored. In this study, we evaluated the performance of three state-of-the-art LLMs-ChatGPT-4o, Gemini 1.5 Pro, and LIama 3.3 70B-in clinical decision support across the entire clinical reasoning workflow of a typical patient encounter. Using 36 case studies, we first assessed LLM's out-of-the-box performance across five key sequential clinical decision-making tasks under two temperature settings (default vs. zero): differential diagnosis, essential immediate steps, relevant diagnostic testing, final diagnosis, and treatment recommendation. All models showed high variability by task, achieving near-perfect accuracy in final diagnosis, poor performance in relevant diagnostic testing, and moderate performance in remaining tasks. Furthermore, ChatGPT performed better under the zero temperature, whereas LIama showed stronger performance under the default temperature. Next, we assessed whether prompt engineering could enhance LLM performance by applying variations of the MedPrompt framework, incorporating targeted and random dynamic few-shot learning. The results demonstrate that prompt engineering is not a one-size-fit-all solution. While it significantly improved the performance on the task with lowest baseline accuracy (relevant diagnostic testing), it was counterproductive for others. Another key finding was that the targeted dynamic few-shot prompting did not consistently outperform random selection, indicating that the presumed benefits of closely matched examples may be counterbalanced by loss of broader contextual diversity. These findings suggest that the impact of prompt engineering is highly model and task-dependent, highlighting the need for tailored, context-aware strategies for integrating LLMs into healthcare.

📄 PDF Abstract BibTeX arXiv:2512.22966

Code (0)

등록된 구현이 없습니다.

Tasks

Prompt EngineeringFew-Shot Learning

Similar Papers 제목 키워드 기반

Instruction Tuning and CoT Prompting for Contextual Medical QA with LLMs

2025-06-13 · Chenqian Le, Ziheng Gong, Chihang Wang, Haowei Ni 외

Large language models (LLMs) have shown great potential in medical question answering (MedQA), yet adapting them to biomedical reasoning remains challenging due to domain-specific complexity and limited supervision. In t…

Medical Question AnsweringMedQAMultiple-choicePrompt Engineering+1

Selection of Prompt Engineering Techniques for Code Generation through Predicting Code Complexity

2024-09-24 · Chung-Yu Wang, Alireza DaghighFarsoodeh, Hung Viet Pham

Large Language Models (LLMs) have demonstrated impressive performance in software engineering tasks. However, improving their accuracy in generating correct and reliable code remains challenging. Numerous prompt engineer…

Code GenerationContrastive LearningHumanEvalmbpp+1

Visual prompt engineering for video models

2026-07-28 · Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer, Neha Kalibhat 외 arxiv

In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential technique for improving language model performance. Since video models are currently bec…

Prompt EngineeringVisual ReasoningImage Editing

P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks

2021-10-14 · Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Lam Tam 외

Prompt tuning, which only tunes continuous prompts with a frozen language model, substantially reduces per-task storage and memory usage at training. However, in the context of NLU, prior work reveals that prompt tuning …

Language ModelingLanguage Modelling

A Survey of Prompt Engineering Methods in Large Language Models for Different NLP Tasks

2024-07-17 · Shubham Vatsal, Harsh Dubey

Large language models (LLMs) have shown remarkable performance on many different Natural Language Processing (NLP) tasks. Prompt engineering plays a key role in adding more to the already existing abilities of LLMs to ac…

Prompt Engineering