paper-with-me

홈 › Papers

MoDE: Effective Multi-task Parameter Efficient Fine-Tuning with a Mixture of Dyadic Experts

2024-08-02 · Lin Ning, Harsh Lara, Meiqi Guo, Abhinav Rastogi

Parameter-efficient fine-tuning techniques like Low-Rank Adaptation (LoRA) have revolutionized the adaptation of large language models (LLMs) to diverse tasks. Recent efforts have explored mixtures of LoRA modules for multi-task settings. However, our analysis reveals redundancy in the down-projection matrices of these architectures. This observation motivates our proposed method, Mixture of Dyadic Experts (MoDE), which introduces a novel design for efficient multi-task adaptation. This is done by sharing the down-projection matrix across tasks and employing atomic rank-one adapters, coupled with routers that allow more sophisticated task-level specialization. Our design allows for more fine-grained mixing, thereby increasing the model's ability to jointly handle multiple tasks. We evaluate MoDE on the Supernatural Instructions (SNI) benchmark consisting of a diverse set of 700+ tasks and demonstrate that it outperforms state-of-the-art multi-task parameter-efficient fine-tuning (PEFT) methods, without introducing additional parameters. Our findings contribute to a deeper understanding of parameter efficiency in multi-task LLM adaptation and provide a practical solution for deploying high-performing, lightweight models.

📄 PDF Abstract BibTeX arXiv:2408.01505

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Jointly Reparametrized Multi-Layer Adaptation for Efficient and Private Tuning

2023-05-30 · Umang Gupta, Aram Galstyan, Greg Ver Steeg

Efficient finetuning of pretrained language transformers is becoming increasingly prevalent for solving natural language processing tasks. While effective, it can still require a large number of tunable parameters. This …

One Adapter for All Programming Languages? Adapter Tuning for Code Search and Summarization

2023-03-28 · Deze Wang, Boxing Chen, Shanshan Li, Wei Luo 외

As pre-trained models automate many code intelligence tasks, a widely used paradigm is to fine-tune a model on the task dataset for each programming language. A recent study reported that multilingual fine-tuning benefit…

AllCode SearchCode Summarization

Parameter Efficient Multi-task Model Fusion with Partial Linearization

2023-10-07 · Anke Tang, Li Shen, Yong Luo, Yibing Zhan 외

Large pre-trained models have enabled significant advances in machine learning and served as foundation components. Model fusion methods, such as task arithmetic, have been proven to be powerful and scalable to incorpora…

parameter-efficient fine-tuningTask Arithmetic

TASO: Task-Aligned Sparse Optimization for Parameter-Efficient Model Adaptation

2025-09-22 · Daiye Miao, Yufang Liu, Jie Wang, Changzhi Sun 외 arxiv

LoRA has become one of the most widely used parameter-efficient fine-tuning methods due to its simplicity and effectiveness. However, numerous studies have shown that LoRA often introduces substantial parameter redundanc…

parameter-efficient fine-tuning

HyperPELT: Unified Parameter-Efficient Language Model Tuning for Both Language and Vision-and-Language Tasks

2022-03-08 · Zhengkun Zhang, Wenya Guo, Xiaojun Meng, Yasheng Wang 외

The workflow of pretraining and fine-tuning has emerged as a popular paradigm for solving various NLP and V&L (Vision-and-Language) downstream tasks. With the capacity of pretrained models growing rapidly, how to perform…

Language ModelingLanguage ModellingMulti-Task Learningparameter-efficient fine-tuning+1