paper-with-me

Papers

An efficient supervised dictionary learning method for audio signal recognition

2018-12-12 · Imad Rida, Romain Hérault, Gilles Gasso

Machine hearing or listening represents an emerging area. Conventional approaches rely on the design of handcrafted features specialized to a specific audio task and that can hardly generalized to other audio fields. For example, Mel-Frequency Cepstral Coefficients (MFCCs) and its variants were successfully applied to computational auditory scene recognition while Chroma vectors are good at music chord recognition. Unfortunately, these predefined features may be of variable discrimination power while extended to other tasks or even within the same task due to different nature of clips. Motivated by this need of a principled framework across domain applications for machine listening, we propose a generic and data-driven representation learning approach. For this sake, a novel and efficient supervised dictionary learning method is presented. The method learns dissimilar dictionaries, one per each class, in order to extract heterogeneous information for classification. In other words, we are seeking to minimize the intra-class homogeneity and maximize class separability. This is made possible by promoting pairwise orthogonality between class specific dictionaries and controlling the sparsity structure of the audio clip's decomposition over these dictionaries. The resulting optimization problem is non-convex and solved using a proximal gradient descent method. Experiments are performed on both computational auditory scene (East Anglia and Rouen) and synthetic music chord recognition datasets. Obtained results show that our method is capable to reach state-of-the-art hand-crafted features for both applications.

📄 PDF Abstract BibTeX arXiv:1812.04748

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Signal RecognitionChord RecognitionDictionary LearningRepresentation LearningScene Recognition

Similar Papers 제목 키워드 기반

A Framework for Generative and Contrastive Learning of Audio Representations

2020-10-22 · Prateek Verma, Julius Smith

In this paper, we present a framework for contrastive learning for audio representations, in a self supervised frame work without access to any ground truth labels. The core idea in self supervised contrastive learning i…

Contrastive Learning

Weakly-supervised Dictionary Learning

2018-02-05 · Zeyu You, Raviv Raich, Xiaoli Z. Fern, Jinsub Kim

We present a probabilistic modeling and inference framework for discriminative analysis dictionary learning under a weak supervision setting. Dictionary learning approaches have been widely used for tasks such as low-lev…

DenoisingDictionary LearningGeneral ClassificationTime Series+1

Supervised Dictionary Learning

2008-12-01 · NeurIPS 2008 12 · Julien Mairal, Jean Ponce, Guillermo Sapiro, Andrew Zisserman 외

It is now well established that sparse signal models are well suited to restoration tasks and can effectively be learned from audio, image, and video data. Recent research has been aimed at learning discriminative sparse…

Dictionary LearningGeneral ClassificationTexture Classification

Probabilistic Modelling of Signal Mixtures with Differentiable Dictionaries

2022-11-28 · Lukáš Samuel Marták, Rainer Kelz, Gerhard Widmer

We introduce a novel way to incorporate prior information into (semi-) supervised non-negative matrix factorization, which we call differentiable dictionary search. It enables general, highly flexible and principled mode…

Completely Unsupervised Phoneme Recognition by Adversarially Learning Mapping Relationships from Audio Embeddings

2018-04-01 · Da-Rong Liu, Kuan-Yu Chen, Hung-Yi Lee, Lin-shan Lee

Unsupervised discovery of acoustic tokens from audio corpora without annotation and learning vector representations for these tokens have been widely studied. Although these techniques have been shown successful in some …

Generative Adversarial NetworkPhoneme Recognition