paper-with-me

Papers

MIMO Self-attentive RNN Beamformer for Multi-speaker Speech Separation

2021-04-17 · Xiyun Li, Yong Xu, Meng Yu, Shi-Xiong Zhang, Jiaming Xu, Bo Xu, Dong Yu

Recently, our proposed recurrent neural network (RNN) based all deep learning minimum variance distortionless response (ADL-MVDR) beamformer method yielded superior performance over the conventional MVDR by replacing the matrix inversion and eigenvalue decomposition with two recurrent neural networks. In this work, we present a self-attentive RNN beamformer to further improve our previous RNN-based beamformer by leveraging on the powerful modeling capability of self-attention. Temporal-spatial self-attention module is proposed to better learn the beamforming weights from the speech and noise spatial covariance matrices. The temporal self-attention module could help RNN to learn global statistics of covariance matrices. The spatial self-attention module is designed to attend on the cross-channel correlation in the covariance matrices. Furthermore, a multi-channel input with multi-speaker directional features and multi-speaker speech separation outputs (MIMO) model is developed to improve the inference efficiency. The evaluations demonstrate that our proposed MIMO self-attentive RNN beamformer improves both the automatic speech recognition (ASR) accuracy and the perceptual estimation of speech quality (PESQ) against prior arts.

📄 PDF Abstract BibTeX arXiv:2104.08450

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionSpeech Separation

Similar Papers 제목 키워드 기반

ASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in Meetings

2024-06-05 · Theo Mariotte, Anthony Larcher, Silvio Montresor, Jean-Hugh Thomas

Speaker Diarization (SD) aims at grouping speech segments that belong to the same speaker. This task is required in many speech-processing applications, such as rich meeting transcription. In this context, distant microp…

speaker-diarizationSpeaker Diarization

MIMO-DBnet: Multi-channel Input and Multiple Outputs DOA-aware Beamforming Network for Speech Separation

2022-12-07 · Yanjie Fu, Haoran Yin, Meng Ge, Longbiao Wang 외

Recently, many deep learning based beamformers have been proposed for multi-channel speech separation. Nevertheless, most of them rely on extra cues known in advance, such as speaker feature, face image or directional in…

Speech Separation

Speaker diarisation using 2D self-attentive combination of embeddings

2019-02-08 · Guangzhi Sun, Chao Zhang, Phil Woodland

Speaker diarisation systems often cluster audio segments using speaker embeddings such as i-vectors and d-vectors. Since different types of embeddings are often complementary, this paper proposes a generic framework to i…

Diversity

MIMO-SPEECH: End-to-End Multi-Channel Multi-Speaker Speech Recognition

2019-10-15 · Xuankai Chang, Wangyou Zhang, Yanmin Qian, Jonathan Le Roux 외

Recently, the end-to-end approach has proven its efficacy in monaural multi-speaker speech recognition. However, high word error rates (WERs) still prevent these systems from being used in practical applications. On the …

speech-recognitionSpeech RecognitionSpeech Separation

Terahertz-Band Joint Ultra-Massive MIMO Radar-Communications: Model-Based and Model-Free Hybrid Beamforming

2021-02-27 · Ahmet M. Elbir, Kumar Vijay Mishra, Symeon Chatzinotas

Wireless communications and sensing at terahertz (THz) band are increasingly investigated as promising short-range technologies because of the availability of high operational bandwidth at THz. In order to address the ex…

model