paper-with-me

홈 › Papers

Will it Merge? On The Causes of Model Mergeability

2026-01-10 · Adir Rahamim, Asaf Yehudai, Boaz Carmeli, Leshem Choshen, Yosi Mass, Yonatan Belinkov arxiv

Model merging has emerged as a promising technique for combining multiple fine-tuned models into a single multitask model without retraining. However, the factors that determine whether merging will succeed or fail remain poorly understood. In this work, we investigate why specific models are merged better than others. To do so, we propose a concrete, measurable definition of mergeability. We investigate several potential causes for high or low mergeability, highlighting the base model knowledge as a dominant factor: Models fine-tuned on instances that the base model knows better are more mergeable than models fine-tuned on instances that the base model struggles with. Based on our mergeability definition, we explore a simple weighted merging technique that better preserves weak knowledge in the base model.

📄 PDF Abstract BibTeX arXiv:2601.06672

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Demystifying Mergeability: Interpretable Properties to Predict Model Merging Success

2026-01-29 · Luca Zhou, Bo Zhao, Rose Yu, Emanuele Rodolà arxiv

Model merging combines knowledge from separately fine-tuned models, yet the factors driving its success remain poorly understood. While recent work treats mergeability as an intrinsic property of the models, we show with…

Predicting Mergeability of Parameter-Efficient Fine-Tuning Updates

2026-06-17 · Lin Tang, Wei Zhang, Jing Li, Hongyu Chen 외 arxiv

Low-rank adaptation (LoRA) makes it cheap to train many domain- and task-specific language model adapters, but whether two adapters can be merged is usually discovered only after both have been fully trained and evaluate…

parameter-efficient fine-tuningInstruction Following

When Privacy Hurts Mergeability: Geometry-Aware Model Merging under Differential Privacy

2026-08-27 · Jin Liu, Junkang Liu, Ning Xi, Yinbin Miao 외 arxiv

Model merging promises to construct a single multi-task model from independently fine-tuned task models without accessing the original task data. This makes it attractive when task data cannot be centralized, but release…

MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent

2025-11-24 · Yuxia Fu, Zhizhen Zhang, Yuqi Zhang, Zijian Wang 외 arxiv

Recent Vision-Language-Action (VLA) models reformulate vision-language models by tuning them with millions of robotic demonstrations. While they perform well when fine-tuned for a single embodiment or task family, extend…

AFA-LoRA: Enabling Non-Linear Adaptations in LoRA with Activation Function Annealing

2025-12-27 · Jiacheng Li, Jianchao Tan, Zhidong Yang, Feiye Huo 외 arxiv

Low-Rank Adaptation (LoRA) is a widely adopted parameter-efficient fine-tuning (PEFT) method. However, its linear adaptation process limits its expressive power. This means there is a gap between the expressive power of …

parameter-efficient fine-tuningReinforcement Learning