paper-with-me

홈 › Papers

Audio Segmentation for Robust Real-Time Speech Recognition Based on Neural Networks

2016-12-01 · IWSLT 2016 12 · Micha Wetzel, Matthias Sperber, Alexander Waibel

Speech that contains multimedia content can pose a serious challenge for real-time automatic speech recognition (ASR) for two reasons: (1) The ASR produces meaningless output, hurting the readability of the transcript. (2) The search space of the ASR is blown up when multimedia content is encountered, resulting in large delays that compromise real-time requirements. This paper introduces a segmenter that aims to remove these problems by detecting music and noise segments in real-time and replacing them with silence. We propose a two step approach, consisting of frame classification and smoothing. First, a classifier detects speech and multimedia on the frame level. In the second step the smoothing algorithm considers the temporal context to prevent rapid class fluctuations. We investigate in frame classification and smoothing settings to obtain an appealing accuracy-latency-tradeoff. The proposed segmenter yields increases the transcript quality of an ASR system by removing on average 39 % of the errors caused by non-speech in the audio stream, while maintaining a real-time applicable delay of 270 milliseconds.

📄 PDF Abstract BibTeX

Code (1)

ches-001/audio-segmenter pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Cocktail-Party Audio-Visual Speech Recognition

2025-06-02 · Thai-Binh Nguyen, Ngoc-Quan Pham, Alexander Waibel

Audio-Visual Speech Recognition (AVSR) offers a robust solution for speech recognition in challenging environments, such as cocktail-party scenarios, where relying solely on audio proves insufficient. However, current AV…

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

AV Taris: Online Audio-Visual Speech Recognition

2020-12-14 · George Sterpu, Naomi Harte

In recent years, Automatic Speech Recognition (ASR) technology has approached human-level performance on conversational speech under relatively clean listening conditions. In more demanding situations involving distant m…

Action DetectionActivity DetectionAudio-Visual Speech RecognitionAutomatic Speech Recognition+4

Toward Streaming ASR with Non-Autoregressive Insertion-based Model

2020-12-18 · Yuya Fujita, Tianzi Wang, Shinji Watanabe, Motoi Omachi

Neural end-to-end (E2E) models have become a promising technique to realize practical automatic speech recognition (ASR) systems. When realizing such a system, one important issue is the segmentation of audio to deal wit…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Segmentationspeech-recognition+1

Towards Audio Token Compression in Large Audio Language Models

2025-11-26 · Saurabhchand Bhati, Samuel Thomas, Hilde Kuehne, Rogerio Feris 외 arxiv

Large Audio Language Models (LALMs) deliver strong performance across speech and audio tasks, but their audio encoders generate high-rate token sequences (e.g., 25 tokens/s), making attention computation costly and limit…

Speech-to-Speech TranslationSpeech Recognition

Moshi: a speech-text foundation model for real-time dialogue

2024-09-17 · Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer 외

We introduce Moshi, a speech-text foundation model and full-duplex spoken dialogue framework. Current systems for spoken dialogue rely on pipelines of independent components, namely voice activity detection, speech recog…

Action DetectionActivity DetectionLanguage ModelingLanguage Modelling+5