paper-with-me

홈 › Papers

Super-Human Performance in Online Low-latency Recognition of Conversational Speech

2020-10-07 · Thai-Son Nguyen, Sebastian Stueker, Alex Waibel

Achieving super-human performance in recognizing human speech has been a goal for several decades, as researchers have worked on increasingly challenging tasks. In the 1990's it was discovered, that conversational speech between two humans turns out to be considerably more difficult than read speech as hesitations, disfluencies, false starts and sloppy articulation complicate acoustic processing and require robust handling of acoustic, lexical and language context, jointly. Early attempts with statistical models could only reach error rates over 50% and far from human performance (WER of around 5.5%). Neural hybrid models and recent attention-based encoder-decoder models have considerably improved performance as such contexts can now be learned in an integral fashion. However, processing such contexts requires an entire utterance presentation and thus introduces unwanted delays before a recognition result can be output. In this paper, we address performance as well as latency. We present results for a system that can achieve super-human performance (at a WER of 5.0%, over the Switchboard conversational benchmark) at a word based latency of only 1 second behind a speaker's speech. The system uses multiple attention-based encoder-decoder networks integrated within a novel low latency incremental inference approach.

📄 PDF Abstract BibTeX arXiv:2010.03449

Code (1)

yoojungsun0/Psych239 pytorch

Tasks

Decoder

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Multi-Head Attention 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Self-regularised Minimum Latency Training for Streaming Transformer-based Speech Recognition

2023-04-24 · Mohan Li, Rama Doddipatla, Catalin Zorila

This paper proposes a self-regularised minimum latency training (SR-MLT) method for streaming Transformer-based automatic speech recognition (ASR) systems. In previous works, latency was optimised by truncating the onlin…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Online Automatic Speech Recognition with Listen, Attend and Spell Model

2020-08-12 · Roger Hsiao, Dogan Can, Tim Ng, Ruchir Travadi 외

The Listen, Attend and Spell (LAS) model and other attention-based automatic speech recognition (ASR) models have known limitations when operated in a fully online mode. In this paper, we analyze the online operation of …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Scaling Up Online Speech Recognition Using ConvNets

2020-01-27 · Vineel Pratap, Qiantong Xu, Jacob Kahn, Gilad Avidov 외

We design an online end-to-end speech recognition system based on Time-Depth Separable (TDS) convolutions and Connectionist Temporal Classification (CTC). We improve the core TDS architecture in order to limit the future…

Decoderspeech-recognitionSpeech Recognition

Low-Latency Human Action Recognition with Weighted Multi-Region Convolutional Neural Network

2018-05-08 · Yunfeng Wang, Wengang Zhou, Qilin Zhang, Xiaotian Zhu 외

Spatio-temporal contexts are crucial in understanding human actions in videos. Recent state-of-the-art Convolutional Neural Network (ConvNet) based action recognition systems frequently involve 3D spatio-temporal ConvNet…

Action RecognitionChunkingOptical Flow EstimationTemporal Action Localization

Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans

2026-07-22 · Wentao Jiang, Youchen Xie, Haidi Fan, Yajing Chen 외 hf

Existing co-speech gesture generation methods are predominantly studied in offline settings, where gestures are synthesized from complete speech segments. However, interactive digital humans in real-world scenarios are r…