paper-with-me

홈 › Papers

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning

2024-11-17 · Wenke Huang, Jian Liang, Zekun Shi, Didi Zhu, Guancheng Wan, He Li, Bo Du, DaCheng Tao, Mang Ye

Multimodal Large Language Model (MLLM) have demonstrated strong generalization capabilities across diverse distributions and tasks, largely due to extensive pre-training datasets. Fine-tuning MLLM has become a common practice to improve performance on specific downstream tasks. However, during fine-tuning, MLLM often faces the risk of forgetting knowledge acquired during pre-training, which can result in a decline in generalization abilities. To balance the trade-off between generalization and specialization, we propose measuring the parameter importance for both pre-trained and fine-tuning distributions, based on frozen pre-trained weight magnitude and accumulated fine-tuning gradient values. We further apply an importance-aware weight allocation strategy, selectively updating relatively important parameters for downstream tasks. We conduct empirical evaluations on both image captioning and visual question-answering tasks using various MLLM architectures. The comprehensive experimental analysis demonstrates the effectiveness of the proposed solution, highlighting the efficiency of the crucial modules in enhancing downstream specialization performance while mitigating generalization degradation in MLLM Fine-Tuning.

📄 PDF Abstract BibTeX arXiv:2411.10928

Code (2)

MindCode-4/code-14/tree/main/CAJ mindspore
pwc-1/Paper-9/tree/main/6/CAJ/src/models mindspore

Tasks

Image CaptioningLanguage ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelQuestion AnsweringVisual Question Answering

Similar Papers 제목 키워드 기반

Keeping Yourself is Important in Downstream Tuning Multimodal Large Language Model

2025-03-06 · Wenke Huang, Jian Liang, Xianda Guo, Yiyang Fang 외

Multi-modal Large Language Models (MLLMs) integrate visual and linguistic reasoning to address complex tasks such as image captioning and visual question answering. While MLLMs demonstrate remarkable versatility, MLLMs a…

General KnowledgeImage CaptioningLanguage ModelingLanguage Modelling+4

MambaTrans: Multimodal Fusion Image Translation via Large Language Model Priors for Downstream Visual Tasks

2025-08-11 · Yushen Xu, Xiaosong Li, Zhenyu Kuang, Xiaoqi Cheng 외 arxiv

The goal of multimodal image fusion is to integrate complementary information from infrared and visible images, generating multimodal fused images for downstream tasks. Existing downstream pre-training models are typical…

Semantic SegmentationObject Detection

Don't Lose Yourself: Boosting Multimodal Recommendation via Reducing Node-neighbor Discrepancy in Graph Convolutional Network

2024-12-25 · Zheyu Chen, Jinfeng Xu, Haibo Hu

The rapid expansion of multimedia contents has led to the emergence of multimodal recommendation systems. It has attracted increasing attention in recommendation systems because its full utilization of data from differen…

Multimodal RecommendationRecommendation Systems

LetsMT!: Cloud-Based Platform for Do-It-Yourself Machine Translation

2012-07-01 · ACL 2012 7 · Andrejs Vasi{\c{l}}jevs, Raivis Skadi{\c{n}}{\v{s}}, J{\"o}rg Tiedemann
Machine TranslationTranslation

Exploring the Diversity and Invariance in Yourself for Visual Pre-Training Task

2021-06-01 · Longhui Wei, Lingxi Xie, Wengang Zhou, Houqiang Li 외

Recently, self-supervised learning methods have achieved remarkable success in visual pre-training task. By simply pulling the different augmented views of each image together or other novel mechanisms, they can learn mu…

DiversitySelf-Supervised Learning