paper-with-me

Papers

Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization

2023-09-28 · Thilo von Neumann, Christoph Boeddeker, Tobias Cord-Landwehr, Marc Delcroix, Reinhold Haeb-Umbach

We propose a modular pipeline for the single-channel separation, recognition, and diarization of meeting-style recordings and evaluate it on the Libri-CSS dataset. Using a Continuous Speech Separation (CSS) system with a TF-GridNet separation architecture, followed by a speaker-agnostic speech recognizer, we achieve state-of-the-art recognition performance in terms of Optimal Reference Combination Word Error Rate (ORC WER). Then, a d-vector-based diarization module is employed to extract speaker embeddings from the enhanced signals and to assign the CSS outputs to the correct speaker. Here, we propose a syntactically informed diarization using sentence- and word-level boundaries of the ASR module to support speaker turn detection. This results in a state-of-the-art Concatenated minimum-Permutation Word Error Rate (cpWER) for the full meeting recognition pipeline.

📄 PDF Abstract BibTeX arXiv:2309.16482

Code (0)

등록된 구현이 없습니다.

Tasks

SentenceSpeech Separation

Similar Papers 제목 키워드 기반

Recognizing Overlapped Speech in Meetings: A Multichannel Separation Approach Using Neural Networks

2018-10-08 · Takuya Yoshioka, Hakan Erdogan, Zhuo Chen, Xiong Xiao 외

The goal of this work is to develop a meeting transcription system that can recognize speech even when utterances of different speakers are overlapped. While speech overlaps have been regarded as a major obstacle in accu…

speech-recognitionSpeech RecognitionSpeech Separation

Low-Latency Speaker-Independent Continuous Speech Separation

2019-04-13 · Takuya Yoshioka, Zhuo Chen, Changliang Liu, Xiong Xiao 외

Speaker independent continuous speech separation (SI-CSS) is a task of converting a continuous audio stream, which may contain overlapping voices of unknown speakers, into a fixed number of continuous signals each of whi…

speech-recognitionSpeech RecognitionSpeech Separation

Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription

2023-09-15 · Peter Vieting, Simon Berger, Thilo von Neumann, Christoph Boeddeker 외

Many real-life applications of automatic speech recognition (ASR) require processing of overlapped speech. A common method involves first separating the speech into overlap-free streams on which ASR is performed. Recentl…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis

2020-11-03 · Desh Raj, Pavel Denisov, Zhuo Chen, Hakan Erdogan 외

Multi-speaker speech recognition of unsegmented recordings has diverse applications such as meeting transcription and automatic subtitle generation. With technical advances in systems dealing with speech separation, spea…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+3

LibriWASN: A Data Set for Meeting Separation, Diarization, and Recognition with Asynchronous Recording Devices

2023-08-21 · Joerg Schmalenstroeer, Tobias Gburrek, Reinhold Haeb-Umbach

We present LibriWASN, a data set whose design follows closely the LibriCSS meeting recognition data set, with the marked difference that the data is recorded with devices that are randomly positioned on a meeting table a…