paper-with-me

홈 › Papers

Scalable Multi-Domain Adaptation of Language Models using Modular Experts

2024-10-14 · Peter Schafhalter, Shun Liao, Yanqi Zhou, Chih-Kuan Yeh, Arun Kandoor, James Laudon

Domain-specific adaptation is critical to maximizing the performance of pre-trained language models (PLMs) on one or multiple targeted tasks, especially under resource-constrained use cases, such as edge devices. However, existing methods often struggle to balance domain-specific performance, retention of general knowledge, and efficiency for training and inference. To address these challenges, we propose Modular Domain Experts (MoDE). MoDE is a mixture-of-experts architecture that augments a general PLMs with modular, domain-specialized experts. These experts are trained independently and composed together via a lightweight training process. In contrast to standard low-rank adaptation methods, each MoDE expert consists of several transformer layers which scale better with more training examples and larger parameter counts. Our evaluation demonstrates that MoDE achieves comparable target performances to full parameter fine-tuning while achieving 1.65% better retention performance. Moreover, MoDE's architecture enables flexible sharding configurations and improves training speeds by up to 38% over state-of-the-art distributed training configurations.

📄 PDF Abstract BibTeX arXiv:2410.10181

Code (0)

등록된 구현이 없습니다.

Tasks

Domain AdaptationGeneral KnowledgeMixture-of-Experts

Similar Papers 제목 키워드 기반

Domain-Specific Data Generation Framework for RAG Adaptation

2025-10-13 · Chris Xing Tian, Weihao Xie, Zhen Chen, Zhengyuan Yi 외 arxiv

Retrieval-Augmented Generation (RAG) combines the language understanding and reasoning power of large language models (LLMs) with external retrieval to enable domain-grounded responses. Effectively adapting RAG systems t…

CATCH: A Modular Cross-domain Adaptive Template with Hook

2025-10-30 · Xinjin Li, Yulie Lu, Jinghan Cao, Yu Ma 외 arxiv

Recent advances in Visual Question Answering (VQA) have demonstrated impressive performance in natural image domains, with models like LLaVA leveraging large language models (LLMs) for open-ended reasoning. However, thei…

Visual Question AnsweringDomain Adaptation

Plug-and-Play Transformer Modules for Test-Time Adaptation

2024-01-06 · Xiangyu Chang, Sk Miraj Ahmed, Srikanth V. Krishnamurthy, Basak Guler 외

Parameter-efficient tuning (PET) methods such as LoRA, Adapter, and Visual Prompt Tuning (VPT) have found success in enabling adaptation to new domains by tuning small modules within a transformer model. However, the num…

Domain AdaptationTest-time AdaptationVisual Prompt Tuning

Graft: Integrating the Domain Knowledge via Efficient Parameter Synergy for MLLMs

2025-06-30 · Yang Dai, Jianxiang An, Tianwei Lin, Hongyang He 외

Multimodal Large Language Models (MLLMs) have achieved success across various domains. However, their applicability tends to degrade when confronted with different types of data inputs, especially for MLLMs that have bee…

Modular Adaptation for Cross-Domain Few-Shot Learning

2021-04-01 · Xiao Lin, Meng Ye, Yunye Gong, Giedrius Buracas 외

Adapting pre-trained representations has become the go-to recipe for learning new downstream tasks with limited examples. While literature has demonstrated great successes via representation learning, in this work, we sh…

Cross-Domain Few-Shotcross-domain few-shot learningFew-Shot LearningRepresentation Learning