paper-with-me

Papers

Parameter-Efficient Fine-Tuning Medical Multimodal Large Language Models for Medical Visual Grounding

2024-10-31 · Jinlong He, Pengfei Li, Gang Liu, Shenjun Zhong

Multimodal Large Language Models (MLLMs) inherit the superior text understanding capabilities of LLMs and extend these capabilities to multimodal scenarios. These models achieve excellent results in the general domain of multimodal tasks. However, in the medical domain, the substantial training costs and the requirement for extensive medical data pose challenges to the development of medical MLLMs. Furthermore, due to the free-text form of answers, tasks such as visual grounding that need to produce output in a prescribed form become difficult for MLLMs. So far, there have been no medical MLLMs works in medical visual grounding area. For the medical vision grounding task, which involves identifying locations in medical images based on short text descriptions, we propose Parameter-efficient Fine-tuning medical multimodal large language models for Medcial Visual Grounding (PFMVG). To validate the performance of the model, we evaluate it on a public benchmark dataset for medical visual grounding, where it achieves competitive results, and significantly outperforming GPT-4v. Our code will be open sourced after peer review.

📄 PDF Abstract BibTeX arXiv:2410.23822

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuningVisual Grounding

Similar Papers 제목 키워드 기반

Can LLMs' Tuning Methods Work in Medical Multimodal Domain?

2024-03-11 · Jiawei Chen, Yue Jiang, Dingkang Yang, Mingcheng Li 외

While Large Language Models (LLMs) excel in world knowledge understanding, adapting them to specific subfields requires precise adjustments. Due to the model's vast scale, traditional global fine-tuning methods for large…

Transfer LearningWorld Knowledge

LDP: Parameter-Efficient Fine-Tuning of Multimodal LLM for Medical Report Generation

2025-12-11 · Tianyu Zhou, Junyi Tang, Zehui Li, Dahong Qian 외 arxiv

Colonoscopic polyp diagnosis is pivotal for early colorectal cancer detection, yet traditional automated reporting suffers from inconsistencies and hallucinations due to the scarcity of high-quality multimodal medical da…

parameter-efficient fine-tuningMedical Report Generation

Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models

2024-04-16 · Songtao Jiang, Tuo Zheng, Yan Zhang, Yeying Jin 외

Recent advancements in general-purpose or domain-specific multimodal large language models (LLMs) have witnessed remarkable progress for medical decision-making. However, they are designated for specific classification o…

image-classificationImage ClassificationMedical Question AnsweringMixture-of-Experts+2

PeFoMed: Parameter Efficient Fine-tuning of Multimodal Large Language Models for Medical Imaging

2024-01-05 · Jinlong He, Pengfei Li, Gang Liu, Genrong He 외

Multimodal large language models (MLLMs) represent an evolutionary expansion in the capabilities of traditional large language models, enabling them to tackle challenges that surpass the scope of purely text-based applic…

Medical Report GenerationMedical Visual Question Answeringparameter-efficient fine-tuningQuestion Answering+4

LLaVA-Ultra: Large Chinese Language and Vision Assistant for Ultrasound

2024-10-19 · Xuechen Guo, Wenhao Chai, Shi-Yan Li, Gaoang Wang

Multimodal Large Language Model (MLLM) has recently garnered attention as a prominent research focus. By harnessing powerful LLM, it facilitates a transition of conversational generative AI from unimodal text to performi…

Instruction FollowingKnowledge DistillationLanguage ModelingLanguage Modelling+6