paper-with-me

홈 › Papers

Fusion-Augmented Large Language Models: Boosting Diagnostic Trustworthiness via Model Consensus

2025-10-16 · Md Kamrul Siam, Md Jobair Hossain Faruk, Jerry Q. Cheng, Huanying Gu arxiv

This study presents a novel multi-model fusion framework leveraging two state-of-the-art large language models (LLMs), ChatGPT and Claude, to enhance the reliability of chest X-ray interpretation on the CheXpert dataset. From the full CheXpert corpus of 224,316 chest radiographs, we randomly selected 234 radiologist-annotated studies to evaluate unimodal performance using image-only prompts. In this setting, ChatGPT and Claude achieved diagnostic accuracies of 62.8% and 76.9%, respectively. A similarity-based consensus approach, using a 95% output similarity threshold, improved accuracy to 77.6%. To assess the impact of multimodal inputs, we then generated synthetic clinical notes following the MIMIC-CXR template and evaluated a separate subset of 50 randomly selected cases paired with both images and synthetic text. On this multimodal cohort, performance improved to 84% for ChatGPT and 76% for Claude, while consensus accuracy reached 91.3%. Across both experimental conditions, agreement-based fusion consistently outperformed individual models. These findings highlight the utility of integrating complementary modalities and using output-level consensus to improve the trustworthiness and clinical utility of AI-assisted radiological diagnosis, offering a practical path to reduce diagnostic errors with minimal computational overhead.

📄 PDF Abstract BibTeX arXiv:2510.16057

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Language-Aware Token Boosting: LLM Language Confusion Reduction Without Tuning

2026-06-08 · Trapoom Ukarapol, Pakhapoom Sarapat, Nut Chukamphaeng arxiv

Large language models (LLMs) sometimes exhibit language confusion when generating non-English text. Existing approaches typically rely on fine-tuning to mitigate this issue. In contrast, we propose a tuning-free paradigm…

X-ray Insights Unleashed: Pioneering the Enhancement of Multi-Label Long-Tail Data

2025-12-24 · Xinquan Yang, Jinheng Xie, Yawen Huang, Yuexiang Li 외 arxiv

Long-tailed pulmonary anomalies in chest radiography present formidable diagnostic challenges. Despite the recent strides in diffusion-based methods for enhancing the representation of tailed lesions, the paucity of rare…

Incremental Learning

CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA

2026-07-23 · Hanseok Oh, Parishad BehnamGhader, Benno Krojer, Hyunji Lee 외 arxiv

Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions. K…

Visual Question AnsweringReferring ExpressionObject RecognitionVisual Grounding

BRAINS: A Retrieval-Augmented System for Alzheimer's Detection and Monitoring

2025-11-04 · Rajan Das Gupta, Md Kishor Morol, Nafiz Fahad, Md Tanzib Hosain 외 arxiv

As the global burden of Alzheimer's disease (AD) continues to grow, early and accurate detection has become increasingly critical, especially in regions with limited access to advanced diagnostic tools. We propose BRAINS…

Alzheimer's Disease Detection

A Survey of Knowledge-Intensive NLP with Pre-Trained Language Models

2022-02-17 · Da Yin, Li Dong, Hao Cheng, Xiaodong Liu 외

With the increasing of model capacity brought by pre-trained language models, there emerges boosting needs for more knowledgeable natural language processing (NLP) models with advanced functionalities including providing…

Language ModelingLanguage Modelling