paper-with-me

홈 › Papers

A Generalist Learner for Multifaceted Medical Image Interpretation

2024-05-13 · Hong-Yu Zhou, Subathra Adithan, Julián Nicolás Acosta, Eric J. Topol, Pranav Rajpurkar

Current medical artificial intelligence systems are often limited to narrow applications, hindering their widespread adoption in clinical practice. To address this limitation, we propose MedVersa, a generalist learner that enables flexible learning and tasking for medical image interpretation. By leveraging a large language model as a learnable orchestrator, MedVersa can learn from both visual and linguistic supervision, support multimodal inputs, and perform real-time task specification. This versatility allows MedVersa to adapt to various clinical scenarios and perform multifaceted medical image analysis. We introduce MedInterp, the largest multimodal dataset to date for medical image interpretation, consisting of over 13 million annotated instances spanning 11 tasks across 3 modalities, to support the development of MedVersa. Our experiments demonstrate that MedVersa achieves state-of-the-art performance in 9 tasks, sometimes outperforming specialist counterparts by over 10%. MedVersa is the first to showcase the viability of multimodal generative medical AI in implementing multimodal outputs, inputs, and dynamic task specification, highlighting its potential as a multifunctional system for comprehensive medical image analysis. This generalist approach to medical image interpretation paves the way for more adaptable and efficient AI-assisted clinical decision-making.

📄 PDF Abstract BibTeX arXiv:2405.07988

Code (0)

등록된 구현이 없습니다.

Tasks

Decision MakingLanguage ModellingLarge Language ModelMedical Image Analysis

Similar Papers 제목 키워드 기반

Diffusion Model as a Generalist Segmentation Learner

2026-04-27 · Haoxiao Wang, Antao Xiang, Haiyang Sun, Peilin Sun 외 arxiv

Diffusion models are primarily trained for image synthesis, yet their denoising trajectories encode rich, spatially aligned visual priors. In this paper, we demonstrate that these priors can be utilized for text-conditio…

Semantic Segmentation

Towards Generalist Biomedical AI

2023-07-26 · Tao Tu, Shekoofeh Azizi, Danny Driess, Mike Schaekermann 외

Medicine is inherently multimodal, with rich data modalities spanning text, imaging, genomics, and more. Generalist biomedical artificial intelligence (AI) systems that flexibly encode, integrate, and interpret this data…

Medical Question AnsweringQuestion Answeringscientific discoveryTransfer Learning+1

Attention-based Dynamic Subspace Learners for Medical Image Analysis

2022-06-18 · Sukesh Adiga V, Jose Dolz, Herve Lombaert

Learning similarity is a key aspect in medical image analysis, particularly in recommendation systems or in uncovering the interpretation of anatomical data in images. Most existing methods learn such similarities in the…

ClusteringImage ClusteringImage RetrievalMedical Image Analysis+3

Med-Flamingo: a Multimodal Medical Few-shot Learner

2023-07-27 · Michael Moor, Qian Huang, Shirley Wu, Michihiro Yasunaga 외

Medicine, by its nature, is a multifaceted domain that requires the synthesis of information across various modalities. Medical generative vision-language models (VLMs) make a first step in this direction and promise man…

Medical Visual Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Super-Generalist: Towards Comprehensive and Accurate Medical Image Understanding via Generalist-Specialist Synergy

2026-07-10 · Shaoteng Zhang, Weiwei Cao, Wanxing Chang, Yutong Xie 외 arxiv

Medical images require comprehensive and accurate interpretation to support the diagnosis of diverse clincial conditions. Recent vision-language generalist models offer broad task coverage and promising zero-shot capabil…