paper-with-me

Papers

Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine

2024-12-12 · Xiaoshuang Huang, Lingdong Shen, Jia Liu, Fangxin Shang, Hongxiang Li, Haifeng Huang, Yehui Yang

In recent years, Multimodal Large Language Models (MLLM) have achieved notable advancements, demonstrating the feasibility of developing an intelligent biomedical assistant. However, current biomedical MLLMs predominantly focus on image-level understanding and restrict interactions to textual commands, thus limiting their capability boundaries and the flexibility of usage. In this paper, we introduce a novel end-to-end multimodal large language model for the biomedical domain, named MedPLIB, which possesses pixel-level understanding. Excitingly, it supports visual question answering (VQA), arbitrary pixel-level prompts (points, bounding boxes, and free-form shapes), and pixel-level grounding. We propose a novel Mixture-of-Experts (MoE) multi-stage training strategy, which divides MoE into separate training phases for a visual-language expert model and a pixel-grounding expert model, followed by fine-tuning using MoE. This strategy effectively coordinates multitask learning while maintaining the computational cost at inference equivalent to that of a single expert model. To advance the research of biomedical MLLMs, we introduce the Medical Complex Vision Question Answering Dataset (MeCoVQA), which comprises an array of 8 modalities for complex medical imaging question answering and image region understanding. Experimental results indicate that MedPLIB has achieved state-of-the-art outcomes across multiple medical visual language tasks. More importantly, in zero-shot evaluations for the pixel grounding task, MedPLIB leads the best small and large models by margins of 19.7 and 15.6 respectively on the mDice metric. The codes, data, and model checkpoints will be made publicly available at https://github.com/ShawnHuang497/MedPLIB.

📄 PDF Abstract BibTeX arXiv:2412.09278

Code (1)

shawnhuang497/medplib 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingLarge Language ModelMixture-of-ExpertsMultimodal Large Language ModelQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

MoE 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

MRIR: Integrating Multimodal Insights for Diffusion-based Realistic Image Restoration

2024-07-04 · Yuhong Zhang, Hengsheng Zhang, Xinning Chai, Rong Xie 외

Realistic image restoration is a crucial task in computer vision, and the use of diffusion-based models for image restoration has garnered significant attention due to their ability to produce realistic results. However,…

DenoisingImage RestorationLanguage ModellingLarge Language Model+1

SafePLUG: Empowering Multimodal LLMs with Pixel-Level Insight and Temporal Grounding for Traffic Accident Understanding

2025-08-09 · Zihao Sheng, Zilin Huang, Yansong Qu, Jiancong Chen 외 arxiv

Multimodal large language models (MLLMs) have achieved remarkable progress across a range of vision-language tasks and demonstrate strong potential for traffic accident understanding. However, existing MLLMs in this doma…

Question Answering

VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

2024-11-07 · Shehan Munasinghe, Hanan Gani, Wenqi Zhu, Jiale Cao 외

Fine-grained alignment between videos and text is challenging due to complex spatial and temporal dynamics in videos. Existing video-based Large Multimodal Models (LMMs) handle basic conversations but struggle with preci…

DecoderLanguage ModelingLanguage ModellingLarge Language Model+3

VideoGLaMM : A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

2025-01-01 · CVPR 2025 1 · Shehan Munasinghe, Hanan Gani, Wenqi Zhu, Jiale Cao 외

Fine-grained alignment between videos and text is challenging due to complex spatial and temporal dynamics in videos. Existing video-based Large Multimodal Models (LMMs) handle basic conversations but struggle with p…

Large Language ModelVideo SegmentationVideo Semantic SegmentationVisual Grounding

DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding

2025-05-27 · Weihao Xuan, Junjue Wang, Heli Qi, Zihang Chen 외

Multimodal large language models have demonstrated remarkable capabilities in visual understanding, but their application to long-term Earth observation analysis remains limited, primarily focusing on single-temporal or …

BenchmarkingChange DetectionEarth ObservationQuestion Answering