paper-with-me

Papers

Language Model Mapping in Multimodal Music Learning: A Grand Challenge Proposal

2025-03-01 · Daniel Chin, Gus Xia

We have seen remarkable success in representation learning and language models (LMs) using deep neural networks. Many studies aim to build the underlying connections among different modalities via the alignment and mappings at the token or embedding level, but so far, most methods are very data-hungry, limiting their performance in domains such as music where paired data are less abundant. We argue that the embedding alignment is only at the surface level of multimodal alignment. In this paper, we propose a grand challenge of \textit{language model mapping} (LMM), i.e., how to map the essence implied in the LM of one domain to the LM of another domain under the assumption that LMs of different modalities are tracking the same underlying phenomena. We first introduce a basic setup of LMM, highlighting the goal to unveil a deeper aspect of cross-modal alignment as well as to achieve more sample-efficiency learning. We then discuss why music is an ideal domain in which to conduct LMM research. After that, we connect LMM in music with a more general and challenging scientific problem of \textit{learning to take actions based on both sensory input and abstract symbols}, and in the end, present an advanced version of the challenge problem setup.

📄 PDF Abstract BibTeX arXiv:2503.00427

Code (0)

등록된 구현이 없습니다.

Tasks

cross-modal alignmentLanguage ModelingLanguage ModellingRepresentation Learning

Similar Papers 제목 키워드 기반

POLY-SIM: Polyglot Speaker Identification with Missing Modality Grand Challenge 2026 Evaluation Plan

2026-03-25 · Marta Moscati, Muhammad Saad Saeed, Marina Zanoni, Mubashir Noman 외 arxiv

Multimodal speaker identification systems typically assume the availability of complete and homogeneous audio-visual modalities during both training and testing. However, in real-world applications, such assumptions ofte…

Speaker Identification

Vis2Mus: Exploring Multimodal Representation Mapping for Controllable Music Generation

2022-11-10 · Runbang Zhang, Yixiao Zhang, Kai Shao, Ying Shan 외

In this study, we explore the representation mapping from the domain of visual arts to the domain of music, with which we can use visual arts as an effective handle to control music generation. Unlike most studies in mul…

Music GenerationRepresentation LearningStyle Transfer

Second Grand-Challenge and Workshop on Multimodal Language (Challenge-HML)

2020-07-01 · WS 2020 7 ·

Proceedings of Grand Challenge and Workshop on Human Multimodal Language (Challenge-HML)

2018-07-01 · WS 2018 7 ·

Remixing Music for Hearing Aids Using Ensemble of Fine-Tuned Source Separators

2024-01-11 · Matthew Daly

This paper introduces our system submission for the Cadenza ICASSP 2024 Grand Challenge, which presents the problem of remixing and enhancing music for hearing aid users. Our system placed first in the challenge, achievi…