paper-with-me

Papers

Evaluating Automatic Speech Recognition in an Incremental Setting

2023-02-23 · Ryan Whetten, Mir Tahsin Imtiaz, Casey Kennington

The increasing reliability of automatic speech recognition has proliferated its everyday use. However, for research purposes, it is often unclear which model one should choose for a task, particularly if there is a requirement for speed as well as accuracy. In this paper, we systematically evaluate six speech recognizers using metrics including word error rate, latency, and the number of updates to already recognized words on English test data, as well as propose and compare two methods for streaming audio into recognizers for incremental recognition. We further propose Revokes per Second as a new metric for evaluating incremental recognition and demonstrate that it provides insights into overall model performance. We find that, generally, local recognizers are faster and require fewer updates than cloud-based recognizers. Finally, we find Meta's Wav2Vec model to be the fastest, and find Mozilla's DeepSpeech model to be the most stable in its predictions.

📄 PDF Abstract BibTeX arXiv:2302.12049

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Test 설명 없음
SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Online Continual Learning of End-to-End Speech Recognition Models

2022-07-11 · Muqiao Yang, Ian Lane, Shinji Watanabe

Continual Learning, also known as Lifelong Learning, aims to continually learn from new data as it becomes available. While prior research on continual learning in automatic speech recognition has focused on the adaptati…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Continual LearningLifelong learning+3

ILASR: Privacy-Preserving Incremental Learning for Automatic Speech Recognition at Production Scale

2022-07-19 · Gopinath Chennupati, Milind Rao, Gurpreet Chadha, Aaron Eakin 외

Incremental learning is one paradigm to enable model building and updating at scale with streaming data. For end-to-end automatic speech recognition (ASR) tasks, the absence of human annotated labels along with the need …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Continual LearningIncremental Learning+3

Simultaneous Speech-to-Speech Translation System with Neural Incremental ASR, MT, and TTS

2020-11-10 · Katsuhito Sudoh, Takatomo Kano, Sashi Novitasari, Tomoya Yanagita 외

This paper presents a newly developed, simultaneous neural speech-to-speech translation system and its evaluation. The system consists of three fully-incremental neural processing modules for automatic speech recognition…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationSimultaneous Speech-to-Speech Translation+8

PI-Whisper: Designing an Adaptive and Incremental Automatic Speech Recognition System for Edge Devices

2024-06-21 · Amir Nassereldine, Dancheng Liu, Chenhui Xu, Ruiyang Qin 외

Edge-based automatic speech recognition (ASR) technologies are increasingly prevalent in the development of intelligent and personalized assistants. However, resource-constrained ASR models face significant challenges in…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Fairnessspeech-recognition+1

Incremental Machine Speech Chain Towards Enabling Listening while Speaking in Real-time

2020-11-04 · Sashi Novitasari, Andros Tjandra, Tomoya Yanagita, Sakriani Sakti 외

Inspired by a human speech chain mechanism, a machine speech chain framework based on deep learning was recently proposed for the semi-supervised development of automatic speech recognition (ASR) and text-to-speech synth…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+4