paper-with-me

Papers

Frame-wise and overlap-robust speaker embeddings for meeting diarization

2023-06-01 · Tobias Cord-Landwehr, Christoph Boeddeker, Cătălin Zorilă, Rama Doddipatla, Reinhold Haeb-Umbach

Using a Teacher-Student training approach we developed a speaker embedding extraction system that outputs embeddings at frame rate. Given this high temporal resolution and the fact that the student produces sensible speaker embeddings even for segments with speech overlap, the frame-wise embeddings serve as an appropriate representation of the input speech signal for an end-to-end neural meeting diarization (EEND) system. We show in experiments that this representation helps mitigate a well-known problem of EEND systems: when increasing the number of speakers the diarization performance drop is significantly reduced. We also introduce block-wise processing to be able to diarize arbitrarily long meetings.

📄 PDF Abstract BibTeX arXiv:2306.00625

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

EEND End-to-End Neural Diarization is a neural network for speaker diarization in which a neural network directly outputs speaker diarization results given a multi-speaker…

Similar Papers 제목 키워드 기반

Geodesic interpolation of frame-wise speaker embeddings for the diarization of meeting scenarios

2024-01-08 · Tobias Cord-Landwehr, Christoph Boeddeker, Cătălin Zorilă, Rama Doddipatla 외

We propose a modified teacher-student training for the extraction of frame-wise speaker embeddings that allows for an effective diarization of meeting scenarios containing partially overlapping speech. To this end, a geo…

Clustering

Speaker Embedding-aware Neural Diarization: an Efficient Framework for Overlapping Speech Diarization in Meeting Scenarios

2022-03-18 · Zhihao Du, Shiliang Zhang, Siqi Zheng, Zhijie Yan

Overlapping speech diarization has been traditionally treated as a multi-label classification problem. In this paper, we reformulate this task as a single-label prediction problem by encoding multiple binary labels into …

Action DetectionActivity DetectionMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATION+1

Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models

2025-06-06 · Yuke Lin, Ming Cheng, Ze Li, Beilong Tang 외

Multi-speaker automatic speech recognition (MS-ASR) faces significant challenges in transcribing overlapped speech, a task critical for applications like meeting transcription and conversational analysis. While serialize…

Automatic Speech Recognitionspeaker-diarizationSpeaker Diarizationspeech-recognition+1

Utterance-Wise Meeting Transcription System Using Asynchronous Distributed Microphones

2020-07-31 · Shota Horiguchi, Yusuke Fujita, Kenji Nagamatsu

A novel framework for meeting transcription using asynchronous microphones is proposed in this paper. It consists of audio synchronization, speaker diarization, utterance-wise speech enhancement using guided source separ…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+3

G-STAR: End-to-End Global Speaker-Tracking Attributed Recognition

2026-03-11 · Jing Peng, Ziyi Chen, Haoyu Li, Yucheng Wang 외 arxiv

We study timestamped speaker-attributed automatic speech recognition (SA-ASR) for long-form, multi-party speech with overlap. In this setting, chunk-wise inference must preserve meeting-level speaker identity consistency…

Speech Recognition