paper-with-me

홈 › Papers

Speaker Adapted Beamforming for Multi-Channel Automatic Speech Recognition

2018-06-19 · Tobias Menne, Ralf Schlüter, Hermann Ney

This paper presents, in the context of multi-channel ASR, a method to adapt a mask based, statistically optimal beamforming approach to a speaker of interest. The beamforming vector of the statistically optimal beamformer is computed by utilizing speech and noise masks, which are estimated by a neural network. The proposed adaptation approach is based on the integration of the beamformer, which includes the mask estimation network, and the acoustic model of the ASR system. This allows for the propagation of the training error, from the acoustic modeling cost function, all the way through the beamforming operation and through the mask estimation network. By using the results of a first pass recognition and by keeping all other parameters fixed, the mask estimation network can therefore be fine tuned by retraining. Utterances of a speaker of interest can thus be used in a two pass approach, to optimize the beamforming for the speech characteristics of that specific speaker. It is shown that this approach improves the ASR performance of a state-of-the-art multi-channel ASR system on the CHiME-4 data. Furthermore the effect of the adaptation on the estimated speech masks is discussed.

📄 PDF Abstract BibTeX arXiv:1806.07407

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Spatial-aware Speaker Diarization for Multi-channel Multi-party Meeting

2022-09-24 · Jie Wang, Yuji Liu, Binling Wang, Yiming Zhi 외

This paper describes a spatial-aware speaker diarization system for the multi-channel multi-party meeting. The diarization system obtains direction information of speaker by microphone array. Speaker spatial embedding is…

speaker-diarizationSpeaker Diarization

A Comparative Study on Multichannel Speaker-Attributed Automatic Speech Recognition in Multi-party Meetings

2022-11-01 · Mohan Shi, Jie Zhang, Zhihao Du, Fan Yu 외

Speaker-attributed automatic speech recognition (SA-ASR) in multi-party meeting scenarios is one of the most valuable and challenging ASR task. It was shown that single-channel frame-level diarization with serialized out…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+3

Joint Beamforming and Speaker-Attributed ASR for Real Distant-Microphone Meeting Transcription

2024-10-29 · Can Cui, Imran Ahamad Sheikh, Mostafa Sadeghi, Emmanuel Vincent

Distant-microphone meeting transcription is a challenging task. State-of-the-art end-to-end speaker-attributed automatic speech recognition (SA-ASR) architectures lack a multichannel noise and reverberation reduction fro…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

End-to-End Multi-speaker ASR with Independent Vector Analysis

2022-04-01 · Robin Scheibler, Wangyou Zhang, Xuankai Chang, Shinji Watanabe 외

We develop an end-to-end system for multi-channel, multi-speaker automatic speech recognition. We propose a frontend for joint source separation and dereverberation based on the independent vector analysis (IVA) paradigm…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

End-to-End Single-Channel Speaker-Turn Aware Conversational Speech Translation

2023-11-01 · Juan Zuluaga-Gomez, Zhaocheng Huang, Xing Niu, Rohit Paturi 외

Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversations by multiple speakers. In this paper…

Automatic Speech Recognitionspeech-recognitionSpeech RecognitionSpeech-to-Text+2