Target Speaker Selection for Neural Network Beamforming in Multi-Speaker Scenarios
We propose a speaker selection mechanism (SSM) for the training of an end-to-end beamforming neural network, based on recent findings that a listener usually looks to the target speaker with a certain undershot angle. The mechanism allows the neural network model to learn toward which speaker to focus, during training, in a multi-speaker scenario, based on the position of listener and speakers. However, only audio information is necessary during inference. We perform acoustic simulations demonstrating the feasibility and performance when the SSM is employed in training. The results show significant increase in speech intelligibility, quality, and distortion metrics when compared to the minimum variance distortionless filter and the same neural network model trained without SSM. The success of the proposed method is a significant step forward toward the solution of the cocktail party problem.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Deep Ad-hoc Beamforming Based on Speaker Extraction for Target-Dependent Speech Separation
Recently, the research on ad-hoc microphone arrays with deep learning has drawn much attention, especially in speech enhancement and separation. Because an ad-hoc microphone array may cover such a large area that multipl…
channel selectionDeep LearningSpeech EnhancementSpeech SeparationL-SpEx: Localized Target Speaker Extraction
Speaker extraction aims to extract the target speaker's voice from a multi-talker speech mixture given an auxiliary reference utterance. Recent studies show that speaker extraction benefits from the location or direction…
Target Speaker ExtractionVisual-Informed Speech Enhancement Using Attention-Based Beamforming
Recent studies have demonstrated that incorporating auxiliary information, such as speaker voiceprint or visual cues, can substantially improve Speech Enhancement (SE) performance. However, single-channel methods often y…
Visual Speech RecognitionSpeaker IdentificationSpeech EnhancementActivity DetectionMicrophone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge
In this paper, we introduce a multi-talker distant automatic speech recognition (DASR) system we designed for the DASR task 1 of the CHiME-8 challenge. Our system performs speaker counting, diarization, and ASR. It handl…
Action DetectionActivity DetectionAutomatic Speech Recognitionspeaker-diarization+4TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings
Since diarization and source separation of meeting data are closely related tasks, we here propose an approach to perform the two objectives jointly. It builds upon the target-speaker voice activity detection (TS-VAD) di…
Action DetectionActivity DetectionSpeech Recognition