paper-with-me

홈 › Papers

Speaker-Adapted End-to-End Visual Speech Recognition for Continuous Spanish

2023-11-21 · David Gimeno-Gómez, Carlos-D. Martínez-Hinarejos

Different studies have shown the importance of visual cues throughout the speech perception process. In fact, the development of audiovisual approaches has led to advances in the field of speech technologies. However, although noticeable results have recently been achieved, visual speech recognition remains an open research problem. It is a task in which, by dispensing with the auditory sense, challenges such as visual ambiguities and the complexity of modeling silence must be faced. Nonetheless, some of these challenges can be alleviated when the problem is approached from a speaker-dependent perspective. Thus, this paper studies, using the Spanish LIP-RTVE database, how the estimation of specialized end-to-end systems for a specific person could affect the quality of speech recognition. First, different adaptation strategies based on the fine-tuning technique were proposed. Then, a pre-trained CTC/Attention architecture was used as a baseline throughout our experiments. Our findings showed that a two-step fine-tuning process, where the VSR system is first adapted to the task domain, provided significant improvements when the speaker adaptation was addressed. Furthermore, results comparable to the current state of the art were reached even when only a limited amount of data was available.

📄 PDF Abstract BibTeX arXiv:2311.12480

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech RecognitionVisual Speech Recognition

Similar Papers 제목 키워드 기반

Personalized Speech Recognition for Children with Test-Time Adaptation

2024-09-19 · Zhonghao Shi, Harshvardhan Srivastava, Xuan Shi, Shrikanth Narayanan 외

Accurate automatic speech recognition (ASR) for children is crucial for effective real-time child-AI interaction, especially in educational applications. However, off-the-shelf ASR models primarily pre-trained on adult d…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

CNVSRC 2023: The First Chinese Continuous Visual Speech Recognition Challenge

2024-06-14 · Chen Chen, Zehua Liu, Xiaolou Li, Lantian Li 외

The first Chinese Continuous Visual Speech Recognition Challenge aimed to probe the performance of Large Vocabulary Continuous Visual Speech Recognition (LVC-VSR) on two tasks: (1) Single-speaker VSR for a particular spe…

speech-recognitionSpeech RecognitionVisual Speech Recognition

An analysis of degenerating speech due to progressive dysarthria on ASR performance

2022-10-31 · Katrin Tomanek, Katie Seaver, Pan-Pan Jiang, Richard Cave 외

Although personalized automatic speech recognition (ASR) models have recently been designed to recognize even severely impaired speech, model performance may degrade over time for persons with degenerating speech. The ai…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Speaker-Independent Speech-Driven Visual Speech Synthesis using Domain-Adapted Acoustic Models

2019-05-15 · Ahmed Hussen Abdelaziz, Barry-John Theobald, Justin Binder, Gabriele Fanelli 외

Speech-driven visual speech synthesis involves mapping features extracted from acoustic speech to the corresponding lip animation controls for a face model. This mapping can take many forms, but a powerful approach is to…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Face Modelspeech-recognition+2

Visual gesture variability between talkers in continuous visual speech

2017-10-03 · Helen L. Bear

Recent adoption of deep learning methods to the field of machine lipreading research gives us two options to pursue to improve system performance. Either, we develop end-to-end systems holistically or, we experiment to f…

Lipreading