paper-with-me

홈 › Papers

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation

2025-02-14 · Tianwei Lin, Wenqiao Zhang, Sijing Li, Yuqian Yuan, Binhe Yu, Haoyuan Li, Wanggui He, Hao Jiang, Mengze Li, Xiaohui Song, Siliang Tang, Jun Xiao, Hui Lin, Yueting Zhuang, Beng Chin Ooi

We present HealthGPT, a powerful Medical Large Vision-Language Model (Med-LVLM) that integrates medical visual comprehension and generation capabilities within a unified autoregressive paradigm. Our bootstrapping philosophy is to progressively adapt heterogeneous comprehension and generation knowledge to pre-trained large language models (LLMs). This is achieved through a novel heterogeneous low-rank adaptation (H-LoRA) technique, which is complemented by a tailored hierarchical visual perception approach and a three-stage learning strategy. To effectively learn the HealthGPT, we devise a comprehensive medical domain-specific comprehension and generation dataset called VL-Health. Experimental results demonstrate exceptional performance and scalability of HealthGPT in medical visual unified tasks. Our project can be accessed at https://github.com/DCDmllm/HealthGPT.

📄 PDF Abstract BibTeX arXiv:2502.09838

Code (1)

dcdmllm/healthgpt 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingPhilosophy

Similar Papers 제목 키워드 기반

DIYHealth Suite: Dataset, Model, and Benchmark for Health Management at Home

2026-05-01 · Changshuo Liu, Junran Wu, Zhongle Xie, Wenqiao Zhang 외 arxiv

Generative AI is reshaping healthcare, yet most existing advances rely on hospital-grade devices, which limits their accessibility and potential for health management outside clinical settings. With the proliferation of …

Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing Bias

2023-05-31 · NeurIPS 2023 11 · Zhongwei Wan, Che Liu, Mi Zhang, Jie Fu 외

The scarcity of data presents a critical obstacle to the efficacy of medical visionlanguage pre-training (VLP). A potential solution lies in the combination of datasets from various language communities. Nevertheless, th…

Disentanglement

Unifying Segment Anything in Microscopy with Multimodal Large Language Model

2025-05-16 · Manyu Li, Ruian He, Zixian Zhang, Weimin Tan 외

Accurate segmentation of regions of interest in biomedical images holds substantial value in image analysis. Although several foundation models for biomedical segmentation have currently achieved excellent performance on…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model

UniDCP: Unifying Multiple Medical Vision-language Tasks via Dynamic Cross-modal Learnable Prompts

2023-12-18 · Chenlu Zhan, Yufei Zhang, Yu Lin, Gaoang Wang 외

Medical vision-language pre-training (Med-VLP) models have recently accelerated the fast-growing medical diagnostics application. However, most Med-VLP models learn task-specific representations independently from scratc…

Language ModelingLanguage Modelling

MEDBind: Unifying Language and Multimodal Medical Data Embeddings

2024-03-19 · Yuan Gao, SangWook Kim, David E Austin, Chris McIntosh

Medical vision-language pretraining models (VLPM) have achieved remarkable progress in fusing chest X-rays (CXR) with clinical texts, introducing image-text data binding approaches that enable zero-shot learning and down…

Language ModelingLanguage ModellingLarge Language ModelRetrieval+2