paper-with-me

홈 › Papers

Med-Flamingo: a Multimodal Medical Few-shot Learner

2023-07-27 · Michael Moor, Qian Huang, Shirley Wu, Michihiro Yasunaga, Cyril Zakka, Yash Dalmia, Eduardo Pontes Reis, Pranav Rajpurkar, Jure Leskovec

Medicine, by its nature, is a multifaceted domain that requires the synthesis of information across various modalities. Medical generative vision-language models (VLMs) make a first step in this direction and promise many exciting clinical applications. However, existing models typically have to be fine-tuned on sizeable down-stream datasets, which poses a significant limitation as in many medical applications data is scarce, necessitating models that are capable of learning from few examples in real-time. Here we propose Med-Flamingo, a multimodal few-shot learner adapted to the medical domain. Based on OpenFlamingo-9B, we continue pre-training on paired and interleaved medical image-text data from publications and textbooks. Med-Flamingo unlocks few-shot generative medical visual question answering (VQA) abilities, which we evaluate on several datasets including a novel challenging open-ended VQA dataset of visual USMLE-style problems. Furthermore, we conduct the first human evaluation for generative medical VQA where physicians review the problems and blinded generations in an interactive app. Med-Flamingo improves performance in generative medical VQA by up to 20\% in clinician's rating and firstly enables multimodal medical few-shot adaptations, such as rationale generation. We release our model, code, and evaluation app under https://github.com/snap-stanford/med-flamingo.

📄 PDF Abstract BibTeX arXiv:2307.15189

Code (4)

snap-stanford/med-flamingo 공식 구현 pytorch
MindSpore-scientific-2/code-5/tree/main/med-flamingo mindspore
MindSpore-scientific-2/code-8/tree/main/med-flamingo mindspore
MindSpore-scientific-2/code-9/tree/main/med-flamingo mindspore

Tasks

Medical Visual Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Flamingo: a Visual Language Model for Few-Shot Learning

2022-04-29 · DeepMind 2022 4 · Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 외

Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research. We introduce Flamingo, a family of Visual Language Mode…

Few-Shot LearningGenerative Visual Question AnsweringLanguage ModelingLanguage Modelling+12

On Large Visual Language Models for Medical Imaging Analysis: An Empirical Study

2024-02-21 · Minh-Hao Van, Prateek Verma, Xintao Wu

Recently, large language models (LLMs) have taken the spotlight in natural language processing. Further, integrating LLMs with vision enables the users to explore emergent abilities with multimodal data. Visual language …

Otter: A Multi-Modal Model with In-Context Instruction Tuning

2023-05-05 · Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang 외

Large language models (LLMs) have demonstrated significant universal capabilities as few/zero-shot learners in various tasks due to their pre-training on vast amounts of text data, as exemplified by GPT-3, which boosted …

GPUIn-Context LearningInstruction FollowingVisual Question Answering+2

Lightweight In-Context Tuning for Multimodal Unified Models

2023-10-08 · Yixin Chen, Shuai Zhang, Boran Han, Jiaya Jia

In-context learning (ICL) involves reasoning from given contextual examples. As more modalities comes, this procedure is becoming more challenging as the interleaved input modalities convolutes the understanding process.…

Image CaptioningIn-Context LearningQuestion AnsweringVisual Entailment+2

COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training

2024-01-01 · Alex Jinpeng Wang, Linjie Li, Kevin Qinghong Lin, JianFeng Wang 외

In the evolution of Vision-Language Pre-training, shifting from short-text comprehension to encompassing extended textual contexts is pivotal. Recent autoregressive vision-language models like \cite{flamingo, palme}, lev…

Language ModellingReading ComprehensionText Generation