paper-with-me

홈 › Papers

Codebook based Audio Feature Representation for Music Information Retrieval

2013-12-19 · Yonatan Vaizman, Brian McFee, Gert Lanckriet

Digital music has become prolific in the web in recent decades. Automated recommendation systems are essential for users to discover music they love and for artists to reach appropriate audience. When manual annotations and user preference data is lacking (e.g. for new artists) these systems must rely on \emph{content based} methods. Besides powerful machine learning tools for classification and retrieval, a key component for successful recommendation is the \emph{audio content representation}. Good representations should capture informative musical patterns in the audio signal of songs. These representations should be concise, to enable efficient (low storage, easy indexing, fast search) management of huge music repositories, and should also be easy and fast to compute, to enable real-time interaction with a user supplying new songs to the system. Before designing new audio features, we explore the usage of traditional local features, while adding a stage of encoding with a pre-computed \emph{codebook} and a stage of pooling to get compact vectorial representations. We experiment with different encoding methods, namely \emph{the LASSO}, \emph{vector quantization (VQ)} and \emph{cosine similarity (CS)}. We evaluate the representations' quality in two music information retrieval applications: query-by-tag and query-by-example. Our results show that concise representations can be used for successful performance in both applications. We recommend using top-$\tau$ VQ encoding, which consistently performs well in both applications, and requires much less computation time than the LASSO.

📄 PDF Abstract BibTeX arXiv:1312.5457

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalManagementMusic Information RetrievalQuantizationRecommendation SystemsRetrievalTAG

Similar Papers 제목 키워드 기반

StepAudio 3 Gen Technical Report

2026-09-11 · Bin Lin, Bo Zhao, Boyang Wang, Boyang Zhang 외 hf

We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types…

Audio Generation

StepAudio 3 Music Technical Report

2026-09-11 · Chengli Feng, Zhiyue Wu, Jiahao Song, Zheqi Dai 외 hf

We introduce StepAudio 3 Music, a large-scale, long-form music generation model that supports explicit musical planning and open-domain text-controlled generation. The StepAudio Music Tokenizer represents audio as a 50-H…

Reinforcement LearningMusic Generation

Exploring State-Space-Model based Language Model in Music Generation

2025-07-09 · Wei-Jaw Lee, Fang-Chih Hsieh, Xuanjun Chen, Fang-Duo Tsai 외 arxiv

The recent surge in State Space Models (SSMs), particularly the emergence of Mamba, has established them as strong alternatives or complementary modules to Transformers across diverse domains. In this work, we aim to exp…

Text-to-Music Generation

An Independence-promoting Loss for Music Generation with Language Models

2024-06-04 · Jean-Marie Lemercier, Simon Rouard, Jade Copet, Yossi Adi 외

Music generation schemes using language modeling rely on a vocabulary of audio tokens, generally provided as codes in a discrete latent space learnt by an auto-encoder. Multi-stage quantizers are often employed to produc…

Language ModelingLanguage ModellingMusic Generation

Two-Dimensional Quantization for Geometry-Aware Audio Coding

2025-12-01 · Tal Shuster, Eliya Nachmani arxiv

Recent neural audio codecs have achieved impressive reconstruction quality, typically relying on quantization methods such as Residual Vector Quantization (RVQ), Vector Quantization (VQ) and Finite Scalar Quantization (F…

Representation Learning