paper-with-me

홈 › Papers

Rethinking Parameter Sharing for LLM Fine-Tuning with Multiple LoRAs

2025-09-29 · Hao Ban, Kaiyi Ji arxiv

Large language models are often adapted using parameter-efficient techniques such as Low-Rank Adaptation (LoRA), formulated as $y = W_0x + BAx$, where $W_0$ is the pre-trained parameters and $x$ is the input to the adapted layer. While multi-adapter extensions often employ multiple LoRAs, prior studies suggest that the inner $A$ matrices are highly similar during training and thus suitable for sharing. We revisit this phenomenon and find that this similarity is largely attributable to the identical initialization rather than shared knowledge, with $B$ playing a more critical role in knowledge encoding and transfer. Motivated by these insights, we propose \textbf{ALoRA}, an asymmetric multi-LoRA design with multiple $A$ matrices and a single shared $B$ in multi-task fine-tuning, and \textbf{Fed-ALoRA}, which shares $B$ across clients in federated fine-tuning under both homogeneous and heterogeneous settings, through a novel matrix decomposition strategy to accommodate heterogeneous ranks across clients. Experiments on commonsense reasoning, math reasoning, multi-task NLP dataset, and federated NLP dataset demonstrate that our methods achieve more balanced performance across tasks with comparable or superior average accuracy relative to existing multi-LoRA approaches. The code is available at https://github.com/OptMN-Lab/ALoRA.

📄 PDF Abstract BibTeX arXiv:2509.25414

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Rethinking embedding coupling in pre-trained language models

2020-10-24 · ICLR 2021 1 · Hyung Won Chung, Thibault Févry, Henry Tsai, Melvin Johnson 외

We re-evaluate the standard practice of sharing weights between input and output embeddings in state-of-the-art pre-trained language models. We show that decoupled embeddings provide increased modeling flexibility, allow…

Cross-Lingual Natural Language InferenceCross-Lingual NERCross-Lingual Paraphrase IdentificationCross-Lingual Question Answering+2

Rethinking Parameter Sharing as Graph Coloring for Structured Compression

2025-11-10 · Boyang Zhang, Daning Cheng, Yunquan Zhang arxiv

Modern deep models have massive parameter sizes, leading to high inference-time memory usage that limits practical deployment. Parameter sharing, a form of structured compression, effectively reduces redundancy, but exis…

Parameter Transfer Unit for Deep Neural Networks

2018-04-23 · Yinghua Zhang, Yu Zhang, Qiang Yang

Parameters in deep neural networks which are trained on large-scale databases can generalize across multiple domains, which is referred as "transferability". Unfortunately, the transferability is usually defined as discr…

How do languages influence each other? Studying cross-lingual data sharing during LM fine-tuning

2023-05-22 · Rochelle Choenni, Dan Garrette, Ekaterina Shutova

Multilingual large language models (MLLMs) are jointly trained on data from many different languages such that representation of individual languages can benefit from other languages' data. Impressive performance on zero…

Cross-Lingual TransferZero-Shot Cross-Lingual Transfer

Task Adaptive Parameter Sharing for Multi-Task Learning

2022-03-30 · CVPR 2022 1 · Matthew Wallingford, Hao Li, Alessandro Achille, Avinash Ravichandran 외

Adapting pre-trained models with broad capabilities has become standard practice for learning a wide range of downstream tasks. The typical approach of fine-tuning different models for each task is performant, but incurs…

Multi-Task Learning