Neural Fast Full-Rank Spatial Covariance Analysis for Blind Source Separation
This paper describes an efficient unsupervised learning method for a neural source separation model that utilizes a probabilistic generative model of observed multichannel mixtures proposed for blind source separation (BSS). For this purpose, amortized variational inference (AVI) has been used for directly solving the inverse problem of BSS with full-rank spatial covariance analysis (FCA). Although this unsupervised technique called neural FCA is in principle free from the domain mismatch problem, it is computationally demanding due to the full rankness of the spatial model in exchange for robustness against relatively short reverberations. To reduce the model complexity without sacrificing performance, we propose neural FastFCA based on the jointly-diagonalizable yet full-rank spatial model. Our neural separation model introduced for AVI alternately performs neural network blocks and single steps of an efficient iterative algorithm called iterative source steering. This alternating architecture enables the separation model to quickly separate the mixture spectrogram by leveraging both the deep neural network and the multichannel optimization algorithm. The training objective with AVI is derived to maximize the marginalized likelihood of the observed mixtures. The experiment using mixture signals of two to four sound sources shows that neural FastFCA outperforms conventional BSS methods and reduces the computational time to about 2% of that for the neural FCA.
Code (0)
등록된 구현이 없습니다.
Tasks
blind source separationVariational InferenceMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Fast Multichannel Source Separation Based on Jointly Diagonalizable Spatial Covariance Matrices
This paper describes a versatile method that accelerates multichannel source separation methods based on full-rank spatial modeling. A popular approach to multichannel source separation is to integrate a spatial model wi…
Speech EnhancementGrassmannian Splatting I: Moving rank-2 Spacetime Surfels for Dynamic Scene Rendering
We introduce Grassmannian splatting, a dynamic scene representation whose primitives are Gaussians supported on 3-planes in spacetime $\R^4$: generically, spatial 2-planes in uniform translation along their normals. Each…
Spatial Channel Covariance Estimation for Hybrid Architectures Based on Tensor Decompositions
Spatial channel covariance information can replace full instantaneous channel state information for the analog precoder design in hybrid analog/digital architectures. Obtaining spatial channel covariance estimation, howe…
Compressive SensingTensor DecompositionSemi-Supervised Multichannel Speech Enhancement With a Deep Speech Prior
This paper describes a semi-supervised multichannel speech enhancement method that uses clean speech data for prior training. Although multichannel nonnegative matrix factorization (MNMF) and its constrained variant call…
Speech EnhancementComputation of the Maximum Likelihood estimator in low-rank Factor Analysis
Factor analysis, a classical multivariate statistical technique is popularly used as a fundamental tool for dimensionality reduction in statistics, econometrics and data science. Estimation is often carried out via the M…
Dimensionality ReductionEconometrics