paper-with-me

홈 › Papers

ES-Merging: Biological MLLM Merging via Embedding Space Signals

2026-03-15 · Wonbin Lee, Dongki Kim, Sung Ju Hwang arxiv

Biological multimodal large language models (MLLMs) have emerged as powerful foundation models for scientific discovery. However, existing models are specialized to a single modality, limiting their ability to solve inherently cross-modal scientific problems. While model merging is an efficient method to combine the different modalities into a unified MLLM, existing methods rely on input-agnostic parameter space heuristics that fail to faithfully capture modality specialization. To overcome this limitation, we propose the Embedding-Signal-based MLLM Merging (ES-Merging), a framework that estimates merging coefficients from embedding space signals, moving the merging paradigm from the parameter signals to the embedding signals. ES-Merging exploits coarse-grained and fine-grained signals from embedding space to estimate the layer-wise and element-wise merging coefficients, respectively, which are jointly combined for complementary coefficient estimation. Through extensive experiments, we demonstrate that ES-Merging outperforms existing merging methods not only on the cross-modal reasoning but also on the single-modal knowledge preservation, establishing that embedding space signals provide a principled and effective foundation for MLLM merging.

📄 PDF Abstract BibTeX arXiv:2603.14405

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization

2025-03-31 · CVPR 2025 1 · Yiyang Du, Xiaochen Wang, Chi Chen, Jiabo Ye 외

Recently, model merging methods have demonstrated powerful strengths in combining abilities on various tasks from multiple Large Language Models (LLMs). While previous model merging methods mainly focus on merging homoge…

Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?

2025-10-16 · Yijie Hu, Zihao Zhou, Kaizhu Huang, Xiaowei Huang 외 arxiv

Math reasoning has been one crucial ability of large language models (LLMs), where significant advancements have been achieved in recent years. However, most efforts focus on LLMs by curating high-quality annotation data…

SSAM: Singular Subspace Alignment for Merging Multimodal Large Language Models

2026-03-23 · Md Kaykobad Reza, Ameya Patil, Edward Ayrapetian, M. Salman Asif arxiv

Multimodal large language models (MLLMs) achieve strong performance by jointly processing inputs from multiple modalities, such as vision, audio, and language. However, building such models or extending them to new modal…

PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging

2026-04-18 · Zibo Shao, Baochen Xiong, Xiaoshan Yang, Yaguang Song 외 arxiv

Multimodal Large Language Models (MLLMs) rely on multimodal pre-training over diverse data sources, where different datasets often induce complementary cross-modal alignment capabilities. Model merging provides a cost-ef…

DRIFT: Transferring Reasoning Priors for Efficient MLLM Fine-Tuning

2025-10-16 · Chao Huang, Zeliang Zhang, Jiang Liu, Ximeng Sun 외 arxiv

Multimodal large language models (MLLMs) have made rapid progress, yet their reasoning ability often lags behind strong text-only LLMs. Bridging this gap typically requires large-scale multimodal reasoning data or reinfo…

Reinforcement LearningMultimodal Reasoning