paper-with-me

홈 › Papers

A Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray Interpretation

2024-01-22 · Zhihong Chen, Maya Varma, Justin Xu, Magdalini Paschali, Dave Van Veen, Andrew Johnston, Alaa Youssef, Louis Blankemeier, Christian Bluethgen, Stephan Altmayer, Jeya Maria Jose Valanarasu, Mohamed Siddig Eltayeb Muneer, Eduardo Pontes Reis, Joseph Paul Cohen, Cameron Olsen, Tanishq Mathew Abraham, Emily B. Tsai, Christopher F. Beaulieu, Jenia Jitsev, Sergios Gatidis, Jean-Benoit Delbrouck, Akshay S. Chaudhari, Curtis P. Langlotz

Over 1.4 billion chest X-rays (CXRs) are performed annually due to their cost-effectiveness as an initial diagnostic test. This scale of radiological studies provides a significant opportunity to streamline CXR interpretation and documentation. While foundation models are a promising solution, the lack of publicly available large-scale datasets and benchmarks inhibits their iterative development and real-world evaluation. To overcome these challenges, we constructed a large-scale dataset (CheXinstruct), which we utilized to train a vision-language foundation model (CheXagent). We systematically demonstrated competitive performance across eight distinct task types on our novel evaluation benchmark (CheXbench). Beyond technical validation, we assessed the real-world utility of CheXagent in directly drafting radiology reports. Our clinical assessment with eight radiologists revealed a 36% time saving for residents using CheXagent-drafted reports, while attending radiologists showed no significant time difference editing resident-drafted or CheXagent-drafted reports. The CheXagent-drafted reports improved the writing efficiency of both radiology residents and attending radiologists in 81% and 61% of cases, respectively, without loss of quality. Overall, we demonstrate that CheXagent can effectively perform a variety of CXR interpretation tasks and holds potential to assist radiologists in routine clinical workflows.

📄 PDF Abstract BibTeX arXiv:2401.12208

Code (1)

Stanford-AIMI/CheXagent pytorch

Tasks

BenchmarkingDiagnosticFairnessLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

Knowledge-enhanced Visual-Language Pre-training on Chest Radiology Images

2023-02-27 · Xiaoman Zhang, Chaoyi Wu, Ya zhang, Yanfeng Wang 외

While multi-modal foundation models pre-trained on large-scale data have been successful in natural language understanding and vision recognition, their use in medical domains is still limited due to the fine-grained nat…

Natural Language UnderstandingRepresentation Learning

A Generative Foundation Model for Chest Radiography

2025-09-04 · Yuanfeng Ji, Dan Lin, Xiyue Wang, Lu Zhang 외 arxiv

The scarcity of well-annotated diverse medical images is a major hurdle for developing reliable AI models in healthcare. Substantial technical advances have been made in generative foundation models for natural images. H…

Data Augmentation

MedDChest: A Content-Aware Multimodal Foundational Vision Model for Thoracic Imaging

2025-11-06 · Mahmoud Soliman, Islam Osman, Mohamed S. Shehata, Rasika Rajapakshe arxiv

The performance of vision models in medical imaging is often hindered by the prevailing paradigm of fine-tuning backbones pre-trained on out-of-domain natural images. To address this fundamental domain gap, we propose Me…

Data Augmentation

Benchmarking Chest X-ray Diagnosis Models Across Multinational Datasets

2025-05-21 · Qinmei Xu, Yiheng Li, Xianghao Zhan, Ahmet Gorkem Er 외

Foundation models leveraging vision-language pretraining have shown promise in chest X-ray (CXR) interpretation, yet their real-world performance across diverse populations and diagnostic tasks remains insufficiently eva…

BenchmarkingDiagnostic

Less Could Be Better: Parameter-efficient Fine-tuning Advances Medical Vision Foundation Models

2024-01-22 · Chenyu Lian, Hong-Yu Zhou, Yizhou Yu, Liansheng Wang

Parameter-efficient fine-tuning (PEFT) that was initially developed for exploiting pre-trained large language models has recently emerged as an effective approach to perform transfer learning on computer vision tasks. Ho…

parameter-efficient fine-tuningTransfer Learning