paper-with-me

Papers

A correlation-permutation approach for speech-music encoders model merging

2025-06-13 · Fabian Ritter-Gutierrez, Yi-Cheng Lin, Jeremy H. M Wong, Hung-Yi Lee, Eng Siong Chng, Nancy F. Chen

Creating a unified speech and music model requires expensive pre-training. Model merging can instead create an unified audio model with minimal computational expense. However, direct merging is challenging when the models are not aligned in the weight space. Motivated by Git Re-Basin, we introduce a correlation-permutation approach that aligns a music encoder's internal layers with a speech encoder. We extend previous work to the case of merging transformer layers. The method computes a permutation matrix that maximizes the model's features-wise cross-correlations layer by layer, enabling effective fusion of these otherwise disjoint models. The merged model retains speech capabilities through this method while significantly enhancing music performance, achieving an improvement of 14.83 points in average score compared to linear interpolation model merging. This work allows the creation of unified audio models from independently trained encoders.

📄 PDF Abstract BibTeX arXiv:2506.11403

Code (0)

등록된 구현이 없습니다.

Tasks

Re-basin

Similar Papers 제목 키워드 기반

The challenge of realistic music generation: modelling raw audio at scale

2018-06-26 · NeurIPS 2018 12 · Sander Dieleman, Aäron van den Oord, Karen Simonyan

Realistic music generation is a challenging task. When building generative models of music that are learnt from data, typically high-level representations such as scores or MIDI are used that abstract away the idiosyncra…

Music Generation

Symmetry-Aware Graph Metanetwork Autoencoders: Model Merging through Parameter Canonicalization

2025-11-16 · Odysseas Boufalis, Jorge Carrasco-Pollo, Joshua Rosenthal, Eduardo Terres-Caballero 외 arxiv

Neural network parameterizations exhibit inherent symmetries that yield multiple equivalent minima within the loss landscape. Scale Graph Metanetworks (ScaleGMNs) explicitly leverage these symmetries by proposing an arch…

MMMOS: Multi-domain Multi-axis Audio Quality Assessment

2025-07-05 · Yi-Cheng Lin, Jia-Hung Chen, Hung-yi Lee arxiv

Accurate audio quality estimation is essential for developing and evaluating audio generation, retrieval, and enhancement systems. Existing non-intrusive assessment models predict a single Mean Opinion Score (MOS) for sp…

Audio Quality AssessmentAudio Generation

Signed-Permutation Coordinate Transport for RMSNorm Transformers

2026-06-30 · John Sweeney arxiv

Modern LLM workflows move coordinate-indexed objects across checkpoints: steering vectors, sparse autoencoders, top-$k$ neuron sets, attribution lists, and merge alignments. This is only well posed after fixing the model…

SLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing

2026-01-14 · Ziyang Ma, Guanrou Yang, Wenxi Chen, Zhifu Gao 외 arxiv

The recent surge in open-source Multimodal Large Language Models (MLLM) frameworks, such as LLaVA, provides a convenient kickoff for artificial intelligence developers and researchers. However, most of the MLLM framework…

parameter-efficient fine-tuningSpeech RecognitionAudio captioning