paper-with-me

Papers

ASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in Meetings

2024-06-05 · Theo Mariotte, Anthony Larcher, Silvio Montresor, Jean-Hugh Thomas

Speaker Diarization (SD) aims at grouping speech segments that belong to the same speaker. This task is required in many speech-processing applications, such as rich meeting transcription. In this context, distant microphone arrays usually capture the audio signal. Beamforming, i.e., spatial filtering, is a common practice to process multi-microphone audio data. However, it often requires an explicit localization of the active source to steer the filter. This paper proposes a self-attention-based algorithm to select the output of a bank of fixed spatial filters. This method serves as a feature extractor for joint Voice Activity (VAD) and Overlapped Speech Detection (OSD). The speaker diarization is then inferred from the detected segments. The approach shows convincing distant VAD, OSD, and SD performance, e.g. 14.5% DER on the AISHELL-4 dataset. The analysis of the self-attention weights demonstrates their explainability, as they correlate with the speaker's angular locations.

📄 PDF Abstract BibTeX arXiv:2406.03251

Code (0)

등록된 구현이 없습니다.

Tasks

speaker-diarizationSpeaker Diarization

Similar Papers 제목 키워드 기반

MIMO Self-attentive RNN Beamformer for Multi-speaker Speech Separation

2021-04-17 · Xiyun Li, Yong Xu, Meng Yu, Shi-Xiong Zhang 외

Recently, our proposed recurrent neural network (RNN) based all deep learning minimum variance distortionless response (ADL-MVDR) beamformer method yielded superior performance over the conventional MVDR by replacing the…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Joint Beamforming and Speaker-Attributed ASR for Real Distant-Microphone Meeting Transcription

2024-10-29 · Can Cui, Imran Ahamad Sheikh, Mostafa Sadeghi, Emmanuel Vincent

Distant-microphone meeting transcription is a challenging task. State-of-the-art end-to-end speaker-attributed automatic speech recognition (SA-ASR) architectures lack a multichannel noise and reverberation reduction fro…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Beamformer-Guided Target Speaker Extraction

2023-03-15 · Mohamed Elminshawi, Srikanth Raj Chetupalli, Emanuël A. P. Habets

We propose a Beamformer-guided Target Speaker Extraction (BG-TSE) method to extract a target speaker's voice from a multi-channel recording informed by the direction of arrival of the target. The proposed method employs …

Target Speaker Extraction

An End-to-end Architecture of Online Multi-channel Speech Separation

2020-09-07 · Jian Wu, Zhuo Chen, Jinyu Li, Takuya Yoshioka 외

Multi-speaker speech recognition has been one of the keychallenges in conversation transcription as it breaks the singleactive speaker assumption employed by most state-of-the-artspeech recognition systems. Speech separa…

speech-recognitionSpeech RecognitionSpeech Separation

Adaptive Dereverberation, Noise and Interferer Reduction Using Sparse Weighted Linearly Constrained Minimum Power Beamforming

2023-03-13 · Henri Gode, Simon Doclo

Interfering sources, background noise and reverberation degrade speech quality and intelligibility in hearing aid applications. In this paper, we present an adaptive algorithm aiming at dereverberation, noise and interfe…

Speech Enhancement