paper-with-me

홈 › Papers

Merge to Mix: Mixing Datasets via Model Merging

2025-05-21 · Zhixu Silvia Tao, Kasper Vinken, Hao-Wei Yeh, Avi Cooper, Xavier Boix

Mixing datasets for fine-tuning large models (LMs) has become critical for maximizing performance on downstream tasks. However, composing effective dataset mixtures typically relies on heuristics and trial-and-error, often requiring multiple fine-tuning runs to achieve the desired outcome. We propose a novel method, $\textit{Merge to Mix}$, that accelerates composing dataset mixtures through model merging. Model merging is a recent technique that combines the abilities of multiple individually fine-tuned LMs into a single LM by using a few simple arithmetic operations. Our key insight is that merging models individually fine-tuned on each dataset in a mixture can effectively serve as a surrogate for a model fine-tuned on the entire mixture. Merge to Mix leverages this insight to accelerate selecting dataset mixtures without requiring full fine-tuning on each candidate mixture. Our experiments demonstrate that Merge to Mix surpasses state-of-the-art methods in dataset selection for fine-tuning LMs.

📄 PDF Abstract BibTeX arXiv:2505.16066

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Exploring Information Seeking Agent Consolidation

2026-01-31 · Guochen Yan, Jialong Wu, Zhengwei Tao, Bo Li 외 arxiv

Information-seeking agents have emerged as a powerful paradigm for knowledge-intensive tasks, yet today's systems remain specialized for the open web, documents, or local knowledge bases, hindering scalable and cross-dom…

Multi-task Code LLMs: Data Mix or Model Merge?

2026-01-28 · Mingzhi Zhu, Boris Sobolev, Rahul Krishna, Raju Pavuluri 외 arxiv

Recent research advocates deploying smaller, specialized code LLMs in agentic frameworks alongside frontier models, sparking interest in efficient strategies for multi-task learning that balance performance, constraints,…

Multi-Task LearningCode Generation

MergeMix: Optimizing Mid-Training Data Mixtures via Learnable Model Merging

2026-01-25 · Jiapeng Wang, Changxin Tian, Kunlong Chen, Ziqi Liu 외 arxiv

Optimizing data mixtures is essential for unlocking the full potential of large language models (LLMs), yet identifying the optimal composition remains computationally prohibitive due to reliance on heuristic trials or e…

On the Limits of Model Merging for Multilinguality in Pre-Training

2026-05-25 · Seth Aycock, Fedor Vitiugin, Aleksandr Umnov, Christof Monz 외 arxiv

Endowing models with consistent multilingual performance can be achieved by mixing pre-training data, or post-training approaches such as language-specific model merging. In this work, we test whether merging can be appl…

Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning

2024-10-14 · Aakanksha, Arash Ahmadian, Seraphina Goldfarb-Tarrant, Beyza Ermis 외

Large Language Models (LLMs) have been adopted and deployed worldwide for a broad variety of applications. However, ensuring their safe use remains a significant challenge. Preference training and safety measures often o…

Multi-Task Learning