Improving the Gap in Visual Speech Recognition Between Normal and Silent Speech Based on Metric Learning
This paper presents a novel metric learning approach to address the performance gap between normal and silent speech in visual speech recognition (VSR). The difference in lip movements between the two poses a challenge for existing VSR models, which exhibit degraded accuracy when applied to silent speech. To solve this issue and tackle the scarcity of training data for silent speech, we propose to leverage the shared literal content between normal and silent speech and present a metric learning approach based on visemes. Specifically, we aim to map the input of two speech types close to each other in a latent space if they have similar viseme representations. By minimizing the Kullback-Leibler divergence of the predicted viseme probability distributions between and within the two speech types, our model effectively learns and predicts viseme identities. Our evaluation demonstrates that our method improves the accuracy of silent VSR, even when limited training data is available.
Code (0)
등록된 구현이 없습니다.
Tasks
Metric Learningspeech-recognitionSpeech RecognitionVisual Speech RecognitionSimilar Papers 제목 키워드 기반
Visual-Only Recognition of Normal, Whispered and Silent Speech
Silent speech interfaces have been recently proposed as a way to enable communication when the acoustic signal is not available. This introduces the need to build visual speech recognition systems for silent and whispere…
Silent Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech RecognitionSilent versus modal multi-speaker speech recognition from ultrasound and video
We investigate multi-speaker speech recognition from ultrasound images of the tongue and video images of the lips. We train our systems on imaging data from modal speech, and evaluate on matched test sets of two speaking…
Silent Speech Recognitionspeech-recognitionSpeech RecognitionA Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
Silent Speech Interfaces (SSIs) offer a noninvasive alternative to brain-computer interfaces for soundless verbal communication. We introduce Multimodal Orofacial Neural Audio (MONA), a system that leverages cross-modal …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)cross-modal alignmentLanguage Modelling+4Cocktail-Party Audio-Visual Speech Recognition
Audio-Visual Speech Recognition (AVSR) offers a robust solution for speech recognition in challenging environments, such as cocktail-party scenarios, where relying solely on audio proves insufficient. However, current AV…
Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech RecognitionContinuous Silent Speech Recognition using EEG
In this paper we explore continuous silent speech recognition using electroencephalography (EEG) signals. We implemented a connectionist temporal classification (CTC) automatic speech recognition (ASR) model to translate…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)EEGElectroencephalogram (EEG)+4