paper-with-me

홈 › Papers

Spatial Diarization for Meeting Transcription with Ad-Hoc Acoustic Sensor Networks

2023-11-27 · Tobias Gburrek, Joerg Schmalenstroeer, Reinhold Haeb-Umbach

We propose a diarization system, that estimates "who spoke when" based on spatial information, to be used as a front-end of a meeting transcription system running on the signals gathered from an acoustic sensor network (ASN). Although the spatial distribution of the microphones is advantageous, exploiting the spatial diversity for diarization and signal enhancement is challenging, because the microphones' positions are typically unknown, and the recorded signals are initially unsynchronized in general. Here, we approach these issues by first blindly synchronizing the signals and then estimating time differences of arrival (TDOAs). The TDOA information is exploited to estimate the speakers' activity, even in the presence of multiple speakers being simultaneously active. This speaker activity information serves as a guide for a spatial mixture model, on which basis the individual speaker's signals are extracted via beamforming. Finally, the extracted signals are forwarded to a speech recognizer. Additionally, a novel initialization scheme for spatial mixture models based on the TDOA estimates is proposed. Experiments conducted on real recordings from the LibriWASN data set have shown that our proposed system is advantageous compared to a system using a spatial mixture model, which does not make use of external diarization information.

📄 PDF Abstract BibTeX arXiv:2311.15597

Code (1)

fgnt/spatiospectral_diarization

Tasks

Diversity

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

LibriWASN: A Data Set for Meeting Separation, Diarization, and Recognition with Asynchronous Recording Devices

2023-08-21 · Joerg Schmalenstroeer, Tobias Gburrek, Reinhold Haeb-Umbach

We present LibriWASN, a data set whose design follows closely the LibriCSS meeting recognition data set, with the marked difference that the data is recorded with devices that are randomly positioned on a meeting table a…

The CUHK-TENCENT speaker diarization system for the ICASSP 2022 multi-channel multi-party meeting transcription challenge

2022-02-04 · Naijun Zheng, Na Li, Xixin Wu, Lingwei Meng 외

This paper describes our speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription (M2MeT) challenge, where Mandarin meeting data were recorded in multi-channel format for diarization an…

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)+5

A Meeting Transcription System for an Ad-Hoc Acoustic Sensor Network

2022-05-02 · Tobias Gburrek, Christoph Boeddeker, Thilo von Neumann, Tobias Cord-Landwehr 외

We propose a system that transcribes the conversation of a typical meeting scenario that is captured by a set of initially unsynchronized microphone arrays at unknown positions. It consists of subsystems for signal synch…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)PositionSpeech Enhancement+2

The xmuspeech system for multi-channel multi-party meeting transcription challenge

2022-02-11 · Jie Wang, Yuji Liu, Binling Wang, Yiming Zhi 외

This paper describes the system developed by the XMUSPEECH team for the Multi-channel Multi-party Meeting Transcription Challenge (M2MeT). For the speaker diarization task, we propose a multi-channel speaker diarization …

speaker-diarizationSpeaker Diarization

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition

2025-05-20 · Ming Gao, Shilong Wu, Hang Chen, Jun Du 외

Meetings are a valuable yet challenging scenario for speech applications due to complex acoustic conditions. This paper summarizes the outcomes of the MISP 2025 Challenge, hosted at Interspeech 2025, which focuses on mul…

Audio-Visual Speech Recognitionspeaker-diarizationSpeaker Diarizationspeech-recognition+2