paper-with-me

홈 › Papers

A Comprehensive Evaluation of Incremental Speech Recognition and Diarization for Conversational AI

2020-12-01 · COLING 2020 8 · Angus Addlesee, Yanchao Yu, Arash Eshghi

Automatic Speech Recognition (ASR) systems are increasingly powerful and more accurate, but also more numerous with several options existing currently as a service (e.g. Google, IBM, and Microsoft). Currently the most stringent standards for such systems are set within the context of their use in, and for, Conversational AI technology. These systems are expected to operate incrementally in real-time, be responsive, stable, and robust to the pervasive yet peculiar characteristics of conversational speech such as disfluencies and overlaps. In this paper we evaluate the most popular of such systems with metrics and experiments designed with these standards in mind. We also evaluate the speaker diarization (SD) capabilities of the same systems which will be particularly important for dialogue systems designed to handle multi-party interaction. We found that Microsoft has the leading incremental ASR system which preserves disfluent materials and IBM has the leading incremental SD system in addition to the ASR that is most robust to speech overlaps. Google strikes a balance between the two but none of these systems are yet suitable to reliably handle natural spontaneous conversations in real-time.

📄 PDF Abstract BibTeX

Code (2)

wallscope-research/incremental-asr-evaluation 공식 구현
wallscope-research/incremental-asr-processing 공식 구현

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarizationspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Speaker Diarization with Lexical Information

2020-04-13 · Tae Jin Park, Kyu J. Han, Jing Huang, Xiaodong He 외

This work presents a novel approach for speaker diarization to leverage lexical information provided by automatic speech recognition. We propose a speaker diarization system that can incorporate word-level speaker turn p…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Clusteringspeaker-diarization+3

Speaker Mask Transformer for Multi-talker Overlapped Speech Recognition

2023-12-18 · Peng Shen, Xugang Lu, Hisashi Kawai

Multi-talker overlapped speech recognition remains a significant challenge, requiring not only speech recognition but also speaker diarization tasks to be addressed. In this paper, to better address these tasks, we first…

speaker-diarizationSpeaker Diarizationspeech-recognitionSpeech Recognition

Online speaker diarization of meetings guided by speech separation

2024-01-30 · Elio Gruttadauria, Mathieu Fontaine, Slim Essid

Overlapped speech is notoriously problematic for speaker diarization systems. Consequently, the use of speech separation has recently been proposed to improve their performance. Although promising, speech separation mode…

Action DetectionActivity Detectionspeaker-diarizationSpeaker Diarization+1

North America Bixby Speaker Diarization System for the VoxCeleb Speaker Recognition Challenge 2021

2021-09-28 · Myungjong Kim, Taeyeon Ki, Aviral Anshu, Vijendra Raj Apsingekar

This paper describes the submission to the speaker diarization track of VoxCeleb Speaker Recognition Challenge 2021 done by North America Bixby Lab of Samsung Research America. Our speaker diarization system consists of …

Clusteringspeaker-diarizationSpeaker DiarizationSpeaker Recognition+1

GIST-AiTeR System for the Diarization Task of the 2022 VoxCeleb Speaker Recognition Challenge

2022-09-21 · Dongkeon Park, Yechan Yu, Kyeong Wan Park, Ji Won Kim 외

This report describes the submission system of the GIST-AiTeR team at the 2022 VoxCeleb Speaker Recognition Challenge (VoxSRC) Track 4. Our system mainly includes speech enhancement, voice activity detection , multi-scal…

Action DetectionActivity DetectionClusteringSpeaker Recognition+1