Papers Medical Question Answering
“Medical Question Answering” 태그가 달린 논문 139편 · 필터 해제
From RAG to Agentic: Validating Islamic-Medicine Responses with LLM Agents
Centuries-old Islamic medical texts like Avicenna's Canon of Medicine and the Prophetic Tibb-e-Nabawi encode a wealth of preventive care, nutrition, and holistic therapies, yet remain inaccessible to many and underutiliz…
Language ModelingLanguage ModellingMedical Question AnsweringNutrition+4Instruction Tuning and CoT Prompting for Contextual Medical QA with LLMs
Large language models (LLMs) have shown great potential in medical question answering (MedQA), yet adapting them to biomedical reasoning remains challenging due to domain-specific complexity and limited supervision. In t…
Medical Question AnsweringMedQAMultiple-choicePrompt Engineering+1MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models
Medical image segmentation is crucial for clinical diagnosis, yet existing models are limited by their reliance on explicit human instructions and lack the active reasoning capabilities to understand complex clinical que…
Image SegmentationMedical DiagnosisMedical Image SegmentationMedical Question Answering+4ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning
Though reasoning-based large language models (LLMs) have excelled in mathematics and programming, their capabilities in knowledge-intensive medical question answering remain underexplored. To address this, we introduce R…
Medical Question AnsweringQuestion AnsweringMed-REFL: Medical Reasoning Enhancement via Self-Corrected Fine-grained Reflection
Large reasoning models have recently made significant strides in mathematical and code reasoning, yet their success has not transferred smoothly to the medical domain. While multiple factors contribute to this disparity,…
Medical Question AnsweringMedQAQuestion AnsweringImproving Reliability and Explainability of Medical Question Answering through Atomic Fact Checking in Retrieval-Augmented LLMs
Large language models (LLMs) exhibit extensive medical knowledge but are prone to hallucinations and inaccurate citations, which pose a challenge to their clinical adoption and regulatory compliance. Current methods, suc…
Fact CheckingHallucinationLong Form Question AnsweringMedical Question Answering+2ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
Hepato-pancreato-biliary (HPB) disorders represent a global public health challenge due to their high morbidity and mortality. Although large language models (LLMs) have shown promising performance in general medical que…
Medical Question AnsweringMultiple-choiceQuestion AnsweringMedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering
Large Language Models (LLMs) have demonstrated remarkable performance on various medical question-answering (QA) benchmarks, including standardized medical exams. However, correct answers alone do not ensure correct logi…
Medical Question AnsweringQuestion AnsweringSentenceER-REASON: A Benchmark Dataset for LLM-Based Clinical Reasoning in the Emergency Room
Large language models (LLMs) have been extensively evaluated on medical question answering tasks based on licensing exams. However, real-world evaluations often depend on costly human annotators, and existing benchmarks …
Medical Question AnsweringQuestion AnsweringAMQA: An Adversarial Dataset for Benchmarking Bias of LLMs in Medicine and Healthcare
Large language models (LLMs) are reaching expert-level accuracy on medical diagnosis questions, yet their mistakes and the biases behind them pose life-critical risks. Bias linked to race, sex, and socioeconomic status i…
BenchmarkingMedical DiagnosisMedical Question AnsweringQuestion AnsweringTask Specific Pruning with LLM-Sieve: How Many Parameters Does Your Task Really Need?
As Large Language Models (LLMs) are increasingly being adopted for narrow tasks - such as medical question answering or sentiment analysis - and deployed in resource-constrained settings, a key question arises: how many …
Medical Question AnsweringQuantizationQuestion AnsweringSentiment AnalysisCollaboration among Multiple Large Language Models for Medical Question Answering
Empowered by vast internal knowledge reservoir, the new generation of large language models (LLMs) demonstrate untapped potential to tackle medical tasks. However, there is insufficient effort made towards summoning up a…
Medical Question AnsweringMultiple-choiceQuestion AnsweringLeveraging Online Data to Enhance Medical Knowledge in a Small Persian Language Model
The rapid advancement of language models has demonstrated the potential of artificial intelligence in the healthcare industry. However, small language models struggle with specialized domains in low-resource languages li…
Language ModelingLanguage ModellingMedical Question AnsweringPatient QA+2What Does Neuro Mean to Cardio? Investigating the Role of Clinical Specialty Data in Medical LLMs
In this paper, we introduce S-MedQA, an English medical question-answering (QA) dataset for benchmarking large language models in fine-grained clinical specialties. We use S-MedQA to check the applicability of a popular …
AllBenchmarkingMedical Question AnsweringMedQA+1Building a Human-Verified Clinical Reasoning Dataset via a Human LLM Hybrid Pipeline for Trustworthy Medical AI
Despite strong performance in medical question-answering, the clinical adoption of Large Language Models (LLMs) is critically hampered by their opaque 'black-box' reasoning, limiting clinician trust. This challenge is co…
Medical Question AnsweringQuestion AnsweringTalk Before You Retrieve: Agent-Led Discussions for Better RAG in Medical QA
Medical question answering (QA) is a reasoning-intensive task that remains challenging for large language models (LLMs) due to hallucinations and outdated domain knowledge. Retrieval-Augmented Generation (RAG) provides a…
Information RetrievalMedical Question AnsweringQuestion AnsweringRAG+2Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding
We introduce a novel approach for calibrating uncertainty quantification (UQ) tailored for multi-modal large language models (LLMs). Existing state-of-the-art UQ methods rely on consistency among multiple responses gener…
Medical Question AnsweringQuestion AnsweringUncertainty QuantificationVisual Question AnsweringWalk the Talk? Measuring the Faithfulness of Large Language Model Explanations
Large language models (LLMs) are capable of generating plausible explanations of how they arrived at an answer to a question. However, these explanations can misrepresent the model's "reasoning" process, i.e., they can b…
Language ModelingLanguage ModellingLarge Language ModelMedical Question Answering+1Exploring the Role of Knowledge Graph-Based RAG in Japanese Medical Question Answering with Small-Scale LLMs
Large language models (LLMs) perform well in medical QA, but their effectiveness in Japanese contexts is limited due to privacy constraints that prevent the use of commercial models like GPT-4 in clinical settings. As a …
Medical Question AnsweringQuestion AnsweringRAGRetrieval-augmented GenerationPR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel Optimization
Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of applications, e.g., medical question-answering, mathematical sciences, and code generation. However, they also exhibit inherent…
Anomaly DetectionBilevel OptimizationCode GenerationMedical Question Answering+3