paper-with-me

Papers

LDP: Parameter-Efficient Fine-Tuning of Multimodal LLM for Medical Report Generation

2025-12-11 · Tianyu Zhou, Junyi Tang, Zehui Li, Dahong Qian, Suncheng Xiang arxiv

Colonoscopic polyp diagnosis is pivotal for early colorectal cancer detection, yet traditional automated reporting suffers from inconsistencies and hallucinations due to the scarcity of high-quality multimodal medical data. To bridge this gap, we propose LDP, a novel framework leveraging multimodal large language models (MLLMs) for professional polyp diagnosis report generation. Specifically, we curate MMEndo, a multimodal endoscopic dataset comprising expert-annotated colonoscopy image-text pairs. We fine-tune the Qwen2-VL-7B backbone using Parameter-Efficient Fine-Tuning (LoRA) and align it with clinical standards via Direct Preference Optimization (DPO). Extensive experiments show that our LDP outperforms existing baselines on both automated metrics and rigorous clinical expert evaluations (achieving a Physician Score of 7.2/10), significantly reducing training computational costs by 833x compared to full fine-tuning. The proposed solution offers a scalable, clinically viable path for primary healthcare, with additional validation on the IU-XRay dataset confirming its robustness.

📄 PDF Abstract BibTeX arXiv:2512.10750

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuningMedical Report Generation

Similar Papers 제목 키워드 기반

PeFoMed: Parameter Efficient Fine-tuning of Multimodal Large Language Models for Medical Imaging

2024-01-05 · Jinlong He, Pengfei Li, Gang Liu, Genrong He 외

Multimodal large language models (MLLMs) represent an evolutionary expansion in the capabilities of traditional large language models, enabling them to tackle challenges that surpass the scope of purely text-based applic…

Medical Report GenerationMedical Visual Question Answeringparameter-efficient fine-tuningQuestion Answering+4

AMRG: Extend Vision Language Models for Automatic Mammography Report Generation

2025-08-12 · Nak-Jun Sung, Donghyun Lee, Bo Hwa Choi, Chae Jung Park arxiv

Mammography report generation is a critical yet underexplored task in medical AI, characterized by challenges such as multiview image reasoning, high-resolution visual cues, and unstructured radiologic language. In this …

parameter-efficient fine-tuning

Scaling medical imaging report generation with multimodal reinforcement learning

2026-01-23 · Qianchu Liu, Sheng Zhang, Guanghui Qin, Yu Gu 외 arxiv

Frontier models have demonstrated remarkable capabilities in understanding and reasoning with natural-language text, but they still exhibit major competency gaps in multimodal understanding and reasoning especially in hi…

Reinforcement Learning

Hallucination-Aware Multimodal Benchmark for Gastrointestinal Image Analysis with Large Vision-Language Models

2025-05-11 · Bidur Khanal, Sandesh Pokhrel, Sanjay Bhandari, Ramesh Rana 외

Vision-Language Models (VLMs) are becoming increasingly popular in the medical domain, bridging the gap between medical images and clinical language. Existing VLMs demonstrate an impressive ability to comprehend medical …

DescriptiveDiagnosticHallucination

Parameter-Efficient Fine-Tuning Medical Multimodal Large Language Models for Medical Visual Grounding

2024-10-31 · Jinlong He, Pengfei Li, Gang Liu, Shenjun Zhong

Multimodal Large Language Models (MLLMs) inherit the superior text understanding capabilities of LLMs and extend these capabilities to multimodal scenarios. These models achieve excellent results in the general domain of…

parameter-efficient fine-tuningVisual Grounding