Language Model Mapping in Multimodal Music Learning: A Grand Challenge Proposal
We have seen remarkable success in representation learning and language models (LMs) using deep neural networks. Many studies aim to build the underlying connections among different modalities via the alignment and mappings at the token or embedding level, but so far, most methods are very data-hungry, limiting their performance in domains such as music where paired data are less abundant. We argue that the embedding alignment is only at the surface level of multimodal alignment. In this paper, we propose a grand challenge of \textit{language model mapping} (LMM), i.e., how to map the essence implied in the LM of one domain to the LM of another domain under the assumption that LMs of different modalities are tracking the same underlying phenomena. We first introduce a basic setup of LMM, highlighting the goal to unveil a deeper aspect of cross-modal alignment as well as to achieve more sample-efficiency learning. We then discuss why music is an ideal domain in which to conduct LMM research. After that, we connect LMM in music with a more general and challenging scientific problem of \textit{learning to take actions based on both sensory input and abstract symbols}, and in the end, present an advanced version of the challenge problem setup.
Code (0)
등록된 구현이 없습니다.
Tasks
cross-modal alignmentLanguage ModelingLanguage ModellingRepresentation LearningSimilar Papers 제목 키워드 기반
POLY-SIM: Polyglot Speaker Identification with Missing Modality Grand Challenge 2026 Evaluation Plan
Multimodal speaker identification systems typically assume the availability of complete and homogeneous audio-visual modalities during both training and testing. However, in real-world applications, such assumptions ofte…
Speaker IdentificationVis2Mus: Exploring Multimodal Representation Mapping for Controllable Music Generation
In this study, we explore the representation mapping from the domain of visual arts to the domain of music, with which we can use visual arts as an effective handle to control music generation. Unlike most studies in mul…
Music GenerationRepresentation LearningStyle TransferSecond Grand-Challenge and Workshop on Multimodal Language (Challenge-HML)
Proceedings of Grand Challenge and Workshop on Human Multimodal Language (Challenge-HML)
Remixing Music for Hearing Aids Using Ensemble of Fine-Tuned Source Separators
This paper introduces our system submission for the Cadenza ICASSP 2024 Grand Challenge, which presents the problem of remixing and enhancing music for hearing aid users. Our system placed first in the challenge, achievi…