paper-with-me

홈 › Papers

Evaluation of real-time transcriptions using end-to-end ASR models

2024-09-09 · Carlos Arriaga, Alejandro Pozo, Javier Conde, Alvaro Alonso

Automatic Speech Recognition (ASR) or Speech-to-text (STT) has greatly evolved in the last few years. Traditional architectures based on pipelines have been replaced by joint end-to-end (E2E) architectures that simplify and streamline the model training process. In addition, new AI training methods, such as weak-supervised learning have reduced the need for high-quality audio datasets for model training. However, despite all these advancements, little to no research has been done on real-time transcription. In real-time scenarios, the audio is not pre-recorded, and the input audio must be fragmented to be processed by the ASR systems. To achieve real-time requirements, these fragments must be as short as possible to reduce latency. However, audio cannot be split at any point as dividing an utterance into two separate fragments will generate an incorrect transcription. Also, shorter fragments provide less context for the ASR model. For this reason, it is necessary to design and test different splitting algorithms to optimize the quality and delay of the resulting transcription. In this paper, three audio splitting algorithms are evaluated with different ASR models to determine their impact on both the quality of the transcription and the end-to-end delay. The algorithms are fragmentation at fixed intervals, voice activity detection (VAD), and fragmentation with feedback. The results are compared to the performance of the same model, without audio fragmentation, to determine the effects of this division. The results show that VAD fragmentation provides the best quality with the highest delay, whereas fragmentation at fixed intervals provides the lowest quality and the lowest delay. The newly proposed feedback algorithm exchanges a 2-4% increase in WER for a reduction of 1.5-2s delay, respectively, to the VAD splitting.

📄 PDF Abstract BibTeX arXiv:2409.05674

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionSpeech-to-Text

Methods 이 논문이 사용한 방법론

Fragmentation Given a pattern $P,$ that is more complicated than the patterns, we fragment $P$ into simpler patterns such that their exact count is known. In the subgraph GNN proposed earlier,…

Similar Papers 제목 키워드 기반

Étude de l'informativité des transcriptions : une approche basée sur le résumé automatique

2018-09-04 · Carlos-Emiliano González-Gallardo, Malek Hajjem, Eric SanJuan, Juan-Manuel Torres-Moreno

In this paper we propose a new approach to evaluate the informativeness of transcriptions coming from Automatic Speech Recognition systems. This approach, based in the notion of informativeness, is focused on the framewo…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Informativenessspeech-recognition+2

The SI TEDx-UM speech database: a new Slovenian Spoken Language Resource

2016-05-01 · LREC 2016 5 · Andrej {\v{Z}}gank, Mirjam Sepesy Mau{\v{c}}ec, Darinka Verdonik

This paper presents a new Slovenian spoken language resource built from TEDx Talks. The speech database contains 242 talks in total duration of 54 hours. The annotation and transcription of acquired spoken material was g…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

Improved Visually Prompted Keyword Localisation in Real Low-Resource Settings

2024-09-09 · Leanne Nortje, Dan Oneata, Herman Kamper

Given an image query, visually prompted keyword localisation (VPKL) aims to find occurrences of the depicted word in a speech collection. This can be useful when transcriptions are not available for a low-resource langua…

Few-Shot Learning

Textual Supervision for Visually Grounded Spoken Language Understanding

2020-10-06 · Findings of the Association for Computational Linguistics 2020 · Bertrand Higy, Desmond Elliott, Grzegorz Chrupała

Visually-grounded models of spoken language understanding extract semantic information directly from speech, without relying on transcriptions. This is useful for low-resource languages, where transcriptions can be expen…

Spoken Language Understanding

Measuring Annotation Efficiency for Handwritten Devanagari Recognition: Sample-Complexity Curves for Four Pretraining Regimes

2026-09-15 · Manglesh Kumar Pandey, Sumit Kumar Banshal arxiv

To train handwritten text recognition systems we need word images and their corresponding transcriptions, and these transcriptions are produced manually. For a script that can be read by only a small number of specialist…

Handwritten Text Recognition