paper-with-me

홈 › Papers

Diagnostic Reasoning Prompts Reveal the Potential for Large Language Model Interpretability in Medicine

2023-08-13 · Thomas Savage, Ashwin Nayak, Robert Gallo, Ekanath Rangan, Jonathan H Chen

One of the major barriers to using large language models (LLMs) in medicine is the perception they use uninterpretable methods to make clinical decisions that are inherently different from the cognitive processes of clinicians. In this manuscript we develop novel diagnostic reasoning prompts to study whether LLMs can perform clinical reasoning to accurately form a diagnosis. We find that GPT4 can be prompted to mimic the common clinical reasoning processes of clinicians without sacrificing diagnostic accuracy. This is significant because an LLM that can use clinical reasoning to provide an interpretable rationale offers physicians a means to evaluate whether LLMs can be trusted for patient care. Novel prompting methods have the potential to expose the black box of LLMs, bringing them one step closer to safe and effective use in medicine.

📄 PDF Abstract BibTeX arXiv:2308.06834

Code (0)

등록된 구현이 없습니다.

Tasks

DiagnosticLanguage ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

Symbolic Grounding Reveals Representational Bottlenecks in Abstract Visual Reasoning

2026-04-23 · Mohit Vaishnav, Tanel Tammet arxiv

Vision--language models (VLMs) often fail on abstract visual reasoning benchmarks such as Bongard problems, raising the question of whether the main bottleneck lies in reasoning or representation. We study this on Bongar…

Visual GroundingVisual Reasoning

Brittleness and Promise: Knowledge Graph Based Reward Modeling for Diagnostic Reasoning

2025-09-22 · Saksham Khatwani, He Cheng, Majid Afshar, Dmitriy Dligach 외 arxiv

Large language models (LLMs) show promise for diagnostic reasoning but often lack reliable, knowledge grounded inference. Knowledge graphs (KGs), such as the Unified Medical Language System (UMLS), offer structured biome…

Question AnsweringKnowledge Graphs

Evaluating ChatGPT's Performance in Classifying Pneumonia from Chest X-Ray Images

2025-10-22 · Pragna Prahallad, Pranathi Prahallad arxiv

In this study, we evaluate the ability of OpenAI's gpt-4o model to classify chest X-ray images as either NORMAL or PNEUMONIA in a zero-shot setting, without any prior fine-tuning. A balanced test set of 400 images (200 f…

Visual Reasoning

Persona is a Double-edged Sword: Mitigating the Negative Impact of Role-playing Prompts in Zero-shot Reasoning Tasks

2024-08-16 · Junseok Kim, Nakyeong Yang, Kyomin Jung

Recent studies demonstrate that prompting a role-playing persona to an LLM improves reasoning capability. However, assigning an adequate persona is difficult since LLMs are extremely sensitive to assigned prompts; thus, …

Position

Extending Beacon to Hindi: Cultural Adaptation Drives Cross-Lingual Sycophancy

2026-01-19 · Sarthak Sattigeri arxiv

Sycophancy, the tendency of language models to prioritize agreement with user preferences over principled reasoning, has been identified as a persistent alignment failure in English-language evaluations. However, it rema…