paper-with-me

홈 › Papers

Sparsity-Aware Evolution for Model Merging

2026-02-09 · Huan Zhang, Yanjian Zhang, Guillaume Wisniewski, Nadi Tomeh, Bang Liu arxiv

We propose a sparsity-aware evolutionary (SAE) framework for model merging that involves iterative pruning-merging cycles to act as a novel mutation operator. We incorporate the sparsity constraints into the score function, which steers the evolutionary process to favor more sparse models, in addition to other conventional performance scores. Interestingly, the by-product of \textit{competition} for sparsity introduces an extra local \textit{attraction} and interplay into the evolutionary process: if one competitor has more zero elements, the other competitor's non-zero elements will occupy those positions, even though the less sparse competitor loses to the more sparse competitor in other positions. The proposed pipeline is evaluated on a variety of large-scale LLM benchmarks. Experiments demonstrate that our approach can improve model merging reliability across multiple benchmarks, and is easy to incorporate due to its simplicity and being orthogonal to most existing approaches.

📄 PDF Abstract BibTeX arXiv:2602.08218

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sparsity Curse: Understanding RLVR Model Parameter Space from Model Merging

2026-06-16 · Chenrui Wu, Zexi Li, Jiajun Bu, Jiangchuan Liu 외 arxiv

Reinforcement Learning with Verifiable Reward (RLVR) has emerged as a powerful post-training paradigm that surpasses Supervised Fine-Tuning (SFT) in eliciting reasoning intelligence and resisting catastrophic forgetting.…

Reinforcement Learning

Dynamic Topic Analysis in Academic Journals using Convex Non-negative Matrix Factorization Method

2025-03-23 · Yang Yang, Tong Zhang, Jian Wu, Lijie Su

With the rapid advancement of large language models, academic topic identification and topic evolution analysis are crucial for enhancing AI's understanding capabilities. Dynamic topic analysis provides a powerful approa…

Black-box Model Merging for Language-Model-as-a-Service with Massive Model Repositories

2025-09-16 · Shilian Chen, Jie Zhou, Tianyu Huai, Yujiang Lu 외 arxiv

Model merging refers to the process of integrating multiple distinct models into a unified model that preserves and combines the strengths and capabilities of the individual models. Most existing approaches rely on task …

ImPart: Importance-Aware Delta-Sparsification for Improved Model Compression and Merging in LLMs

2025-04-17 · Yan Yang, Yixia Li, Hongru Wang, Xuetao Wei 외

With the proliferation of task-specific large language models, delta compression has emerged as a method to mitigate the resource challenges of deploying numerous such models by effectively compressing the delta model pa…

Model CompressionQuantization

Bridging Training and Merging Through Momentum-Aware Optimization

2025-12-18 · Alireza Moayedikia, Alicia Troncoso arxiv

Training large neural networks and merging task-specific models both exploit low-rank structure and require parameter importance estimation, yet these challenges have been pursued in isolation. Current workflows compute …

Natural Language Understanding