paper-with-me

Papers

Fast MVAE: Joint separation and classification of mixed sources based on multichannel variational autoencoder with auxiliary classifier

2018-12-16 · Li Li, Hirokazu Kameoka, Shoji Makino

This paper proposes an alternative algorithm for multichannel variational autoencoder (MVAE), a recently proposed multichannel source separation approach. While MVAE is notable in its impressive source separation performance, the convergence-guaranteed optimization algorithm and that it allows us to estimate source-class labels simultaneously with source separation, there are still two major drawbacks, i.e., the high computational complexity and unsatisfactory source classification accuracy. To overcome these drawbacks, the proposed method employs an auxiliary classifier VAE, an information-theoretic extension of the conditional VAE, for learning the generative model of the source spectrograms. Furthermore, with the trained auxiliary classifier, we introduce a novel algorithm for the optimization that is able to not only reduce the computational time but also improve the source classification performance. We call the proposed method "fast MVAE (fMVAE)". Experimental evaluations revealed that fMVAE achieved comparative source separation performance to MVAE and about 80% source classification accuracy rate while it reduced about 93% computational time.

📄 PDF Abstract BibTeX arXiv:1812.06391

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGeneral Classification

Methods 이 논문이 사용한 방법론

USD Coin Customer Service Number +1-833-534-1729 설명 없음
Auxiliary Classifier Auxiliary Classifiers are type of architectural component that seek to improve the convergence of very deep networks. They are classifier heads we attach to layers before the…

Similar Papers 제목 키워드 기반

Generalized Multichannel Variational Autoencoder for Underdetermined Source Separation

2018-09-29 · Shogo Seki, Hirokazu Kameoka, Li Li, Tomoki Toda 외

This paper deals with a multichannel audio source separation problem under underdetermined conditions. Multichannel Non-negative Matrix Factorization (MNMF) is one of powerful approaches, which adopts the NMF concept for…

Audio Source Separation

Semi-blind source separation with multichannel variational autoencoder

2018-08-02 · Hirokazu Kameoka, Li Li, Shota Inoue, Shoji Makino

This paper proposes a multichannel source separation technique called the multichannel variational autoencoder (MVAE) method, which uses a conditional VAE (CVAE) to model and estimate the power spectrograms of the source…

blind source separationDecoder

Target Speech Extraction Based on Blind Source Separation and X-vector-based Speaker Selection Trained with Data Augmentation

2020-05-16 · Zhaoyi Gu, Lele Liao, Kai Chen, Jing Lu

Extracting the desired speech from a mixture is a meaningful and challenging task. The end-to-end DNN-based methods, though attractive, face the problem of generalization. In this paper, we explore a sequential approach …

blind source separationData AugmentationSpeaker RecognitionSpeech Extraction

Joint Multimodal Learning with Deep Generative Models

2016-11-07 · Masahiro Suzuki, Kotaro Nakayama, Yutaka Matsuo

We investigate deep generative models that can exchange multiple modalities bi-directionally, e.g., generating images from corresponding texts and vice versa. Recently, some studies handle multiple modalities on deep gen…

Integrating Random Effects in Variational Autoencoders for Dimensionality Reduction of Correlated Data

2024-12-22 · Giora Simchoni, Saharon Rosset

Variational Autoencoders (VAE) are widely used for dimensionality reduction of large-scale tabular and image datasets, under the assumption of independence between data observations. In practice, however, datasets are of…

Dimensionality Reduction