paper-with-me

홈 › Papers

Interactive decoding of words from visual speech recognition models

2021-07-01 · Brendan Shillingford, Yannis Assael, Misha Denil

This work describes an interactive decoding method to improve the performance of visual speech recognition systems using user input to compensate for the inherent ambiguity of the task. Unlike most phoneme-to-word decoding pipelines, which produce phonemes and feed these through a finite state transducer, our method instead expands words in lockstep, facilitating the insertion of interaction points at each word position. Interaction points enable us to solicit input during decoding, allowing users to interactively direct the decoding process. We simulate the behavior of user input using an oracle to give an automated evaluation, and show promise for the use of this method for text input.

📄 PDF Abstract BibTeX arXiv:2107.00692

Code (0)

등록된 구현이 없습니다.

Tasks

Positionspeech-recognitionSpeech RecognitionVisual Speech Recognition

Similar Papers 제목 키워드 기반

Who Needs Words? Lexicon-Free Speech Recognition

2019-04-09 · Tatiana Likhomanenko, Gabriel Synnaeve, Ronan Collobert

Lexicon-free speech recognition naturally deals with the problem of out-of-vocabulary (OOV) words. In this paper, we show that character-based language models (LM) can perform as well as word-based LMs for speech recogni…

speech-recognitionSpeech Recognition

Tree-constrained Pointer Generator with Graph Neural Network Encodings for Contextual Speech Recognition

2022-07-02 · Guangzhi Sun, Chao Zhang, Philip C. Woodland

Incorporating biasing words obtained as contextual knowledge is critical for many automatic speech recognition (ASR) applications. This paper proposes the use of graph neural network (GNN) encodings in a tree-constrained…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Graph Neural Networkspeech-recognition+1

An investigation of phone-based subword units for end-to-end speech recognition

2020-04-08 · Weiran Wang, Guangsen Wang, Aadyot Bhatnagar, Yingbo Zhou 외

Phones and their context-dependent variants have been the standard modeling units for conventional speech recognition systems, while characters and subwords have demonstrated their effectiveness for end-to-end recognitio…

DecoderLanguage ModelingLanguage Modellingspeech-recognition+1

Streaming Speech-to-Confusion Network Speech Recognition

2023-06-02 · Denis Filimonov, Prabhat Pandey, Ariya Rastrow, Ankur Gandhe 외

In interactive automatic speech recognition (ASR) systems, low-latency requirements limit the amount of search space that can be explored during decoding, particularly in end-to-end neural ASR. In this paper, we present …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Recognition of Isolated Words using Zernike and MFCC features for Audio Visual Speech Recognition

2014-07-04 · Prashant Bordea, Amarsinh Varpeb, Ramesh Manzac, Pravin Yannawara

Automatic Speech Recognition (ASR) by machine is an attractive research topic in signal processing domain and has attracted many researchers to contribute in this area. In recent year, there have been many advances in au…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+2