paper-with-me

홈 › Papers

An Empirical Study and Theoretical Explanation on Task-Level Model-Merging Collapse

2026-03-10 · Yuan Cao, Dezhi Ran, Yuzhe Guo, Mengzhou Wu, Simin Chen, Linyi Li, Wei Yang, Tao Xie arxiv

Model merging unifies independently fine-tuned LLMs from the same base, enabling reuse and integration of parallel development efforts without retraining. However, in practice we observe that merging does not always succeed: certain combinations of task-specialist models suffer from catastrophic performance degradation after merging. We refer to this failure mode as merging collapse. Intuitively, collapse arises when the learned representations or parameter adjustments for different tasks are fundamentally incompatible, so that merging forces destructive interference rather than synergy. In this paper, we identify and characterize the phenomenon of task-level merging collapse, where certain task combinations consistently trigger huge performance degradation across all merging methods. Through extensive experiments and statistical analysis, we demonstrate that representational incompatibility between tasks is strongly correlated with merging collapse, while parameter-space conflict metrics show minimal correlation, challenging conventional wisdom in model merging literature. We provide a theoretical explanation on this phenomenon through rate-distortion theory with a dimension-dependent bound, establishing fundamental limits on task mergeability regardless of methodology.

📄 PDF Abstract BibTeX arXiv:2603.09463

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Probing GNN Explainers: A Rigorous Theoretical and Empirical Analysis of GNN Explanation Methods

2021-06-16 · Chirag Agarwal, Marinka Zitnik, Himabindu Lakkaraju

As Graph Neural Networks (GNNs) are increasingly being employed in critical real-world applications, several methods have been proposed in recent literature to explain the predictions of these models. However, there has …

Fairness

On a Benefit of Masked Language Model Pretraining: Robustness to Simplicity Bias

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Despite the success of pretrained masked language models (MLM), why MLM pretraining is useful is still a question not fully answered. In this work we theoretically and empirically that MLM pretraining makes models robust…

Language ModelingLanguage Modelling

Understanding Masked Autoencoders via Hierarchical Latent Variable Models

2023-06-08 · CVPR 2023 1 · Lingjing Kong, Martin Q. Ma, Guangyi Chen, Eric P. Xing 외

Masked autoencoder (MAE), a simple and effective self-supervised learning framework based on the reconstruction of masked image regions, has recently achieved prominent success in a variety of vision tasks. Despite the e…

Self-Supervised Learning

Exploring Counterfactual Explanations Through the Lens of Adversarial Examples: A Theoretical and Empirical Analysis

2021-06-18 · Martin Pawelczyk, Chirag Agarwal, Shalmali Joshi, Sohini Upadhyay 외

As machine learning (ML) models become more widely deployed in high-stakes applications, counterfactual explanations have emerged as key tools for providing actionable model explanations in practice. Despite the growing …

counterfactualCounterfactual Explanation

Resilience of Bayesian Layer-Wise Explanations under Adversarial Attacks

2021-02-22 · Ginevra Carbone, Guido Sanguinetti, Luca Bortolussi

We consider the problem of the stability of saliency-based explanations of Neural Network predictions under adversarial attacks in a classification task. Saliency interpretations of deterministic Neural Networks are rema…

General Classification