paper-with-me

홈 › Papers

Enhancing Multi-modal Models with Heterogeneous MoE Adapters for Fine-tuning

2025-03-26 · Sashuai Zhou, Hai Huang, Yan Xia

Multi-modal models excel in cross-modal tasks but are computationally expensive due to their billions of parameters. Parameter-efficient fine-tuning (PEFT) offers a solution by adding small trainable components while freezing pre-trained parameters. However, existing methods primarily focus on uni-modal processing, overlooking the critical modal fusion needed for multi-modal tasks. To fill this gap, we propose heterogeneous mixture of experts adapters that extend the traditional PEFT framework to support multi-modal expert combinations and improve information interaction. Additionally, our approach modifies the affine linear expert design to enable efficient modal fusion in a low-rank space, achieving competitive performance with only 5-8\% of the parameters fine-tuned. Experiments across eight downstream tasks, including visual-audio and text-visual, demonstrate the superior performance of the approach.

📄 PDF Abstract BibTeX arXiv:2503.20633

Code (0)

등록된 구현이 없습니다.

Tasks

Mixture-of-Expertsparameter-efficient fine-tuning

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Federated Cross-Modal Retrieval with Missing Modalities via Semantic Routing and Adapter Personalization

2026-04-24 · Hefeng Zhou, Xuan Liu, Sicheng Chen, Wutong Zhang 외 arxiv

Federated cross-modal retrieval faces severe challenges from heterogeneous client data, particularly non-IID semantic distributions and missing modalities. Under such heterogeneity, a single global model is often insuffi…

Cross-Modal Retrieval

FAME-ViL: Multi-Tasking Vision-Language Model for Heterogeneous Fashion Tasks

2023-03-04 · CVPR 2023 1 · Xiao Han, Xiatian Zhu, Licheng Yu, Li Zhang 외

In the fashion domain, there exists a variety of vision-and-language (V+L) tasks, including cross-modal retrieval, text-guided image retrieval, multi-modal classification, and image captioning. They differ drastically in…

Cross-Modal RetrievalImage CaptioningImage RetrievalLanguage Modeling+3

FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations

2024-09-09 · Ziyao Wang, Zheyu Shen, Yexiao He, Guoheng Sun 외

The rapid development of Large Language Models (LLMs) has been pivotal in advancing AI, with pre-trained LLMs being adaptable to diverse downstream tasks through fine-tuning. Federated learning (FL) further enhances fine…

Federated LearningPrivacy Preserving

Enhancing Model Performance: Another Approach to Vision-Language Instruction Tuning

2024-07-25 · Vedanshu, MM Tripathi, Bhavnesh Jaint

The integration of large language models (LLMs) with vision-language (VL) tasks has been a transformative development in the realm of artificial intelligence, highlighting the potential of LLMs as a versatile general-pur…

Chatbot

HierAdaptMR: Cross-Center Cardiac MRI Reconstruction with Hierarchical Feature Adapters

2025-08-18 · Ruru Xu, Ilkay Oksuz arxiv

Deep learning-based cardiac MRI reconstruction faces significant domain shift challenges when deployed across multiple clinical centers with heterogeneous scanner configurations and imaging protocols. We propose HierAdap…

MRI Reconstruction