paper-with-me

홈 › Papers

On Large Visual Language Models for Medical Imaging Analysis: An Empirical Study

2024-02-21 · Minh-Hao Van, Prateek Verma, Xintao Wu

Recently, large language models (LLMs) have taken the spotlight in natural language processing. Further, integrating LLMs with vision enables the users to explore emergent abilities with multimodal data. Visual language models (VLMs), such as LLaVA, Flamingo, or CLIP, have demonstrated impressive performance on various visio-linguistic tasks. Consequently, there are enormous applications of large models that could be potentially used in the biomedical imaging field. Along that direction, there is a lack of related work to show the ability of large models to diagnose the diseases. In this work, we study the zero-shot and few-shot robustness of VLMs on the medical imaging analysis tasks. Our comprehensive experiments demonstrate the effectiveness of VLMs in analyzing biomedical images such as brain MRIs, microscopic images of blood cells, and chest X-rays.

📄 PDF Abstract BibTeX arXiv:2402.14162

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

MoME: Mixture of Visual Language Medical Experts for Medical Imaging Segmentation

2025-10-30 · Arghavan Rezvani, Xiangyi Yan, Anthony T. Wu, Kun Han 외 arxiv

In this study, we propose MoME, a Mixture of Visual Language Medical Experts, for Medical Image Segmentation. MoME adapts the successful Mixture of Experts (MoE) paradigm, widely used in Large Language Models (LLMs), for…

Medical Image Segmentation

Advancing Medical Imaging with Language Models: A Journey from N-grams to ChatGPT

2023-04-11 · Mingzhe Hu, Shaoyan Pan, Yuheng Li, Xiaofeng Yang

In this paper, we aimed to provide a review and tutorial for researchers in the field of medical imaging using language models to improve their tasks at hand. We began by providing an overview of the history and concepts…

DiagnosticImage CaptioningQuestion AnsweringVisual Question Answering

UniBiomed: A Universal Foundation Model for Grounded Biomedical Image Interpretation

2025-04-30 · Linshan Wu, Yuxiang Nie, Sunan He, Jiaxin Zhuang 외

Multi-modal interpretation of biomedical images opens up novel opportunities in biomedical image analysis. Conventional AI approaches typically rely on disjointed training, i.e., Large Language Models (LLMs) for clinical…

DiagnosticLarge Language ModelQuestion AnsweringText Generation+1

OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical Tasks

2025-11-02 · Zhihao Peng, Cheng Wang, Shengyuan Liu, Zhiying Liang 외 arxiv

Brain imaging analysis is crucial for diagnosing and treating brain disorders, and multimodal large language models (MLLMs) are increasingly supporting it. However, current brain imaging visual question-answering (VQA) b…

Residual-based Language Models are Free Boosters for Biomedical Imaging

2024-03-26 · Zhixin Lai, Jing Wu, Suiyao Chen, Yucheng Zhou 외

In this study, we uncover the unexpected efficacy of residual-based large language models (LLMs) as part of encoders for biomedical imaging tasks, a domain traditionally devoid of language or textual data. The approach d…