paper-with-me

홈 › Papers

Demystifying Mergeability: Interpretable Properties to Predict Model Merging Success

2026-01-29 · Luca Zhou, Bo Zhao, Rose Yu, Emanuele Rodolà arxiv

Model merging combines knowledge from separately fine-tuned models, yet the factors driving its success remain poorly understood. While recent work treats mergeability as an intrinsic property of the models, we show with an architecture-agnostic framework that it fundamentally depends on both the merging method and the partner tasks. Using L1-regularized linear optimization over a set of interpretable pairwise metrics (e.g., gradient L_2 distance), we uncover properties correlating with post-merge normalized accuracy across five merging methods. We find architecture- and method-specific variation in success drivers (64.0% average top-5 metric overlap; 79.3% sign agreement), with certain methods, notably TIES, exhibiting distinct ``fingerprints'' that diverge from the broader consensus. Crucially, however, gradient alignment metrics consistently emerge as the most fundamental signals of compatibility. These findings provide a diagnostic foundation for understanding mergeability and motivate future merge-aware fine-tuning strategies.

📄 PDF Abstract BibTeX arXiv:2601.22285

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Will it Merge? On The Causes of Model Mergeability

2026-01-10 · Adir Rahamim, Asaf Yehudai, Boaz Carmeli, Leshem Choshen 외 arxiv

Model merging has emerged as a promising technique for combining multiple fine-tuned models into a single multitask model without retraining. However, the factors that determine whether merging will succeed or fail remai…

Predicting Mergeability of Parameter-Efficient Fine-Tuning Updates

2026-06-17 · Lin Tang, Wei Zhang, Jing Li, Hongyu Chen 외 arxiv

Low-rank adaptation (LoRA) makes it cheap to train many domain- and task-specific language model adapters, but whether two adapters can be merged is usually discovered only after both have been fully trained and evaluate…

parameter-efficient fine-tuningInstruction Following

When Privacy Hurts Mergeability: Geometry-Aware Model Merging under Differential Privacy

2026-08-27 · Jin Liu, Junkang Liu, Ning Xi, Yinbin Miao 외 arxiv

Model merging promises to construct a single multi-task model from independently fine-tuned task models without accessing the original task data. This makes it attractive when task data cannot be centralized, but release…

MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent

2025-11-24 · Yuxia Fu, Zhizhen Zhang, Yuqi Zhang, Zijian Wang 외 arxiv

Recent Vision-Language-Action (VLA) models reformulate vision-language models by tuning them with millions of robotic demonstrations. While they perform well when fine-tuned for a single embodiment or task family, extend…

An Empirical Study and Theoretical Explanation on Task-Level Model-Merging Collapse

2026-03-10 · Yuan Cao, Dezhi Ran, Yuzhe Guo, Mengzhou Wu 외 arxiv

Model merging unifies independently fine-tuned LLMs from the same base, enabling reuse and integration of parallel development efforts without retraining. However, in practice we observe that merging does not always succ…