SCDiar: a streaming diarization system based on speaker change detection and speech recognition
In hours-long meeting scenarios, real-time speech stream often struggles with achieving accurate speaker diarization, commonly leading to speaker identification and speaker count errors. To address this challenge, we propose SCDiar, a system that operates on speech segments, split at the token level by a speaker change detection (SCD) module. Building on these segments, we introduce several enhancements to efficiently select the best available segment for each speaker. These improvements lead to significant gains across various benchmarks. Notably, on real-world meeting data involving more than ten participants, SCDiar outperforms previous systems by up to 53.6\% in accuracy, substantially narrowing the performance gap between online and offline systems.
Code (0)
등록된 구현이 없습니다.
Tasks
Change Detectionspeaker-diarizationSpeaker DiarizationSpeaker Identificationspeech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Highly Efficient Real-Time Streaming and Fully On-Device Speaker Diarization with Multi-Stage Clustering
While recent research advances in speaker diarization mostly focus on improving the quality of diarization results, there is also an increasing interest in improving the efficiency of diarization systems. In this paper, …
ClusteringCPUspeaker-diarizationSpeaker DiarizationTurn-to-Diarize: Online Speaker Diarization Constrained by Transformer Transducer Speaker Turn Detection
In this paper, we present a novel speaker diarization system for streaming on-device applications. In this system, we use a transformer transducer to detect the speaker turns, represent each speaker turn by a speaker emb…
Clusteringspeaker-diarizationSpeaker DiarizationLS-EEND: Long-Form Streaming End-to-End Neural Diarization with Online Attractor Extraction
This work proposes a frame-wise online/streaming end-to-end neural diarization (EEND) method, which detects speaker activities in a frame-in-frame-out fashion. The proposed model mainly consists of a causal embedding enc…
DecoderFormDiariST: Streaming Speech Translation with Speaker Diarization
End-to-end speech translation (ST) for conversation recordings involves several under-explored challenges such as speaker diarization (SD) without accurate word time stamps and handling of overlapping speech in a streami…
speaker-diarizationSpeaker Diarizationspeech-recognitionSpeech Recognition+1VibeVoice-ASR-Streaming Technical Report
Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing …
Speaker Diarization