paper-with-me

홈 › Papers

MuChin: A Chinese Colloquial Description Benchmark for Evaluating Language Models in the Field of Music

2024-02-15 · ZiHao Wang, Shuyu Li, Tao Zhang, Qi Wang, Pengfei Yu, Jinyang Luo, Yan Liu, Ming Xi, Kejun Zhang

The rapidly evolving multimodal Large Language Models (LLMs) urgently require new benchmarks to uniformly evaluate their performance on understanding and textually describing music. However, due to semantic gaps between Music Information Retrieval (MIR) algorithms and human understanding, discrepancies between professionals and the public, and low precision of annotations, existing music description datasets cannot serve as benchmarks. To this end, we present MuChin, the first open-source music description benchmark in Chinese colloquial language, designed to evaluate the performance of multimodal LLMs in understanding and describing music. We established the Caichong Music Annotation Platform (CaiMAP) that employs an innovative multi-person, multi-stage assurance method, and recruited both amateurs and professionals to ensure the precision of annotations and alignment with popular semantics. Utilizing this method, we built a dataset with multi-dimensional, high-precision music annotations, the Caichong Music Dataset (CaiMD), and carefully selected 1,000 high-quality entries to serve as the test set for MuChin. Based on MuChin, we analyzed the discrepancies between professionals and amateurs in terms of music description, and empirically demonstrated the effectiveness of annotated data for fine-tuning LLMs. Ultimately, we employed MuChin to evaluate existing music understanding models on their ability to provide colloquial descriptions of music. All data related to the benchmark, along with the scoring code and detailed appendices, have been open-sourced (https://github.com/CarlWangChina/MuChin/).

📄 PDF Abstract BibTeX arXiv:2402.09871

Code (1)

CarlWangChina/MuChin 공식 구현 pytorch

Tasks

Information RetrievalMusic Information Retrieval

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

ACE: Automatic Colloquialism, Typographical and Orthographic Errors Detection for Chinese Language

2016-12-01 · COLING 2016 12 · Shichao Dong, Gabriel Pui Cheong Fung, Binyang Li, Baolin Peng 외

We present a system called ACE for Automatic Colloquialism and Errors detection for written Chinese. ACE is based on the combination of N-gram model and rule-base model. Although it focuses on detecting colloquial Canton…

Language ModelingLanguage Modelling

MuDiT & MuSiT: Alignment with Colloquial Expression in Description-to-Song Generation

2024-07-03 · ZiHao Wang, Haoxuan Liu, Jiaxing Yu, Tao Zhang 외

Amid the rising intersection of generative AI and human artistic processes, this study probes the critical yet less-explored terrain of alignment in human-centric automatic song composition. We propose a novel task of Co…

DescriptiveRhythm

MIXCD: System Description for Evaluating Chinese Word Similarity at SemEval-2012

2012-07-01 · SEMEVAL 2012 7 · Yingjie Zhang, Bin Li, Xin-yu Dai, Jia-Jun Chen
Information RetrievalSemantic Textual SimilarityWord Sense DisambiguationWord Similarity

Transformer-based Automatic Speech Recognition of Formal and Colloquial Czech in MALACH Project

2022-06-15 · Jan Lehečka, Josef V. Psutka, Josef Psutka

Czech is a very specific language due to its large differences between the formal and the colloquial form of speech. While the formal (written) form is used mainly in official documents, literature, and public speeches, …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Formspeech-recognition+1

WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models

2026-02-12 · Yangzhuo Li, Shengpeng Ji, Yifu Chen, Tianle Liang 외 arxiv

With the rapid integration of advanced reasoning capabilities into spoken dialogue models, the field urgently demands benchmarks that transcend simple interactions to address real-world complexity. However, current evalu…