paper-with-me

Papers Medical Question Answering

“Medical Question Answering” 태그가 달린 논문 139편 · 필터 해제

From RAG to Agentic: Validating Islamic-Medicine Responses with LLM Agents

2025-06-18 · Mohammad Amaan Sayeed, Mohammed Talha Alam, Raza Imam, Shahab Saquib Sohail 외

Centuries-old Islamic medical texts like Avicenna's Canon of Medicine and the Prophetic Tibb-e-Nabawi encode a wealth of preventive care, nutrition, and holistic therapies, yet remain inaccessible to many and underutiliz…

Language ModelingLanguage ModellingMedical Question AnsweringNutrition+4

Instruction Tuning and CoT Prompting for Contextual Medical QA with LLMs

2025-06-13 · Chenqian Le, Ziheng Gong, Chihang Wang, Haowei Ni 외

Large language models (LLMs) have shown great potential in medical question answering (MedQA), yet adapting them to biomedical reasoning remains challenging due to domain-specific complexity and limited supervision. In t…

Medical Question AnsweringMedQAMultiple-choicePrompt Engineering+1

MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models

2025-06-12 · Yu Huang, Zelin Peng, Yichen Zhao, Piao Yang 외

Medical image segmentation is crucial for clinical diagnosis, yet existing models are limited by their reliance on explicit human instructions and lack the active reasoning capabilities to understand complex clinical que…

Image SegmentationMedical DiagnosisMedical Image SegmentationMedical Question Answering+4

ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning

2025-06-11 · Yu Sun, Xingyu Qian, Weiwen Xu, Hao Zhang 외

Though reasoning-based large language models (LLMs) have excelled in mathematics and programming, their capabilities in knowledge-intensive medical question answering remain underexplored. To address this, we introduce R…

Medical Question AnsweringQuestion Answering

Med-REFL: Medical Reasoning Enhancement via Self-Corrected Fine-grained Reflection

2025-06-11 · Zongxian Yang, Jiayu Qian, Zegao Peng, Haoyu Zhang 외

Large reasoning models have recently made significant strides in mathematical and code reasoning, yet their success has not transferred smoothly to the medical domain. While multiple factors contribute to this disparity,…

Medical Question AnsweringMedQAQuestion Answering

Improving Reliability and Explainability of Medical Question Answering through Atomic Fact Checking in Retrieval-Augmented LLMs

2025-05-30 · Juraj Vladika, Annika Domres, Mai Nguyen, Rebecca Moser 외

Large language models (LLMs) exhibit extensive medical knowledge but are prone to hallucinations and inaccurate citations, which pose a challenge to their clinical adoption and regulatory compliance. Current methods, suc…

Fact CheckingHallucinationLong Form Question AnsweringMedical Question Answering+2

ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases

2025-05-30 · Yuchong Li, Xiaojun Zeng, Chihua Fang, Jian Yang 외

Hepato-pancreato-biliary (HPB) disorders represent a global public health challenge due to their high morbidity and mortality. Although large language models (LLMs) have shown promising performance in general medical que…

Medical Question AnsweringMultiple-choiceQuestion Answering

MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering

2025-05-29 · Yuexing Hao, Kumail Alhamoud, Hyewon Jeong, Haoran Zhang 외

Large Language Models (LLMs) have demonstrated remarkable performance on various medical question-answering (QA) benchmarks, including standardized medical exams. However, correct answers alone do not ensure correct logi…

Medical Question AnsweringQuestion AnsweringSentence

ER-REASON: A Benchmark Dataset for LLM-Based Clinical Reasoning in the Emergency Room

2025-05-28 · Nikita Mehandru, Niloufar Golchini, David Bamman, Travis Zack 외

Large language models (LLMs) have been extensively evaluated on medical question answering tasks based on licensing exams. However, real-world evaluations often depend on costly human annotators, and existing benchmarks …

Medical Question AnsweringQuestion Answering

AMQA: An Adversarial Dataset for Benchmarking Bias of LLMs in Medicine and Healthcare

2025-05-26 · Ying Xiao, Jie Huang, Ruijuan He, Jing Xiao 외

Large language models (LLMs) are reaching expert-level accuracy on medical diagnosis questions, yet their mistakes and the biases behind them pose life-critical risks. Bias linked to race, sex, and socioeconomic status i…

BenchmarkingMedical DiagnosisMedical Question AnsweringQuestion Answering

Task Specific Pruning with LLM-Sieve: How Many Parameters Does Your Task Really Need?

2025-05-23 · Waleed Reda, Abhinav Jangda, Krishna Chintalapudi

As Large Language Models (LLMs) are increasingly being adopted for narrow tasks - such as medical question answering or sentiment analysis - and deployed in resource-constrained settings, a key question arises: how many …

Medical Question AnsweringQuantizationQuestion AnsweringSentiment Analysis

Collaboration among Multiple Large Language Models for Medical Question Answering

2025-05-22 · Kexin Shang, Chia-Hsuan Chang, Christopher C. Yang

Empowered by vast internal knowledge reservoir, the new generation of large language models (LLMs) demonstrate untapped potential to tackle medical tasks. However, there is insufficient effort made towards summoning up a…

Medical Question AnsweringMultiple-choiceQuestion Answering

Leveraging Online Data to Enhance Medical Knowledge in a Small Persian Language Model

2025-05-21 · Mehrdad ghassabi, Pedram Rostami, Hamidreza Baradaran Kashani, Amirhossein Poursina 외

The rapid advancement of language models has demonstrated the potential of artificial intelligence in the healthcare industry. However, small language models struggle with specialized domains in low-resource languages li…

Language ModelingLanguage ModellingMedical Question AnsweringPatient QA+2

What Does Neuro Mean to Cardio? Investigating the Role of Clinical Specialty Data in Medical LLMs

2025-05-15 · Xinlan Yan, Di wu, Yibin Lei, Christof Monz 외

In this paper, we introduce S-MedQA, an English medical question-answering (QA) dataset for benchmarking large language models in fine-grained clinical specialties. We use S-MedQA to check the applicability of a popular …

AllBenchmarkingMedical Question AnsweringMedQA+1

Building a Human-Verified Clinical Reasoning Dataset via a Human LLM Hybrid Pipeline for Trustworthy Medical AI

2025-05-11 · Chao Ding, Mouxiao Bian, Pengcheng Chen, Hongliang Zhang 외

Despite strong performance in medical question-answering, the clinical adoption of Large Language Models (LLMs) is critically hampered by their opaque 'black-box' reasoning, limiting clinician trust. This challenge is co…

Medical Question AnsweringQuestion Answering

Talk Before You Retrieve: Agent-Led Discussions for Better RAG in Medical QA

2025-04-30 · Xuanzhao Dong, Wenhui Zhu, Hao Wang, Xiwen Chen 외

Medical question answering (QA) is a reasoning-intensive task that remains challenging for large language models (LLMs) due to hallucinations and outdated domain knowledge. Retrieval-Augmented Generation (RAG) provides a…

Information RetrievalMedical Question AnsweringQuestion AnsweringRAG+2

Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding

2025-04-30 · Trilok Padhi, Ramneet Kaur, Adam D. Cobb, Manoj Acharya 외

We introduce a novel approach for calibrating uncertainty quantification (UQ) tailored for multi-modal large language models (LLMs). Existing state-of-the-art UQ methods rely on consistency among multiple responses gener…

Medical Question AnsweringQuestion AnsweringUncertainty QuantificationVisual Question Answering

Walk the Talk? Measuring the Faithfulness of Large Language Model Explanations

2025-04-19 · Katie Matton, Robert Osazuwa Ness, John Guttag, Emre Kiciman

Large language models (LLMs) are capable of generating plausible explanations of how they arrived at an answer to a question. However, these explanations can misrepresent the model's "reasoning" process, i.e., they can b…

Language ModelingLanguage ModellingLarge Language ModelMedical Question Answering+1

Exploring the Role of Knowledge Graph-Based RAG in Japanese Medical Question Answering with Small-Scale LLMs

2025-04-15 · Yingjian Chen, Feiyang Li, Xingyu Song, Tianxiao Li 외

Large language models (LLMs) perform well in medical QA, but their effectiveness in Japanese contexts is limited due to privacy constraints that prevent the use of commercial models like GPT-4 in clinical settings. As a …

Medical Question AnsweringQuestion AnsweringRAGRetrieval-augmented Generation

PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel Optimization

2025-04-10 · Yang Jiao, Xiaodong Wang, Kai Yang

Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of applications, e.g., medical question-answering, mathematical sciences, and code generation. However, they also exhibit inherent…

Anomaly DetectionBilevel OptimizationCode GenerationMedical Question Answering+3
1–20 / 139 다음 →