paper-with-me

홈 › Papers

Automatic Viseme Vocabulary Construction to Enhance Continuous Lip-reading

2017-04-26 · Adriana Fernandez-Lopez, Federico M. Sukno

Speech is the most common communication method between humans and involves the perception of both auditory and visual channels. Automatic speech recognition focuses on interpreting the audio signals, but it has been demonstrated that video can provide information that is complementary to the audio. Thus, the study of automatic lip-reading is important and is still an open problem. One of the key challenges is the definition of the visual elementary units (the visemes) and their vocabulary. Many researchers have analyzed the importance of the phoneme to viseme mapping and have proposed viseme vocabularies with lengths between 11 and 15 visemes. These viseme vocabularies have usually been manually defined by their linguistic properties and in some cases using decision trees or clustering techniques. In this work, we focus on the automatic construction of an optimal viseme vocabulary based on the association of phonemes with similar appearance. To this end, we construct an automatic system that uses local appearance descriptors to extract the main characteristics of the mouth region and HMMs to model the statistic relations of both viseme and phoneme sequences. To compare the performance of the system different descriptors (PCA, DCT and SIFT) are analyzed. We test our system in a Spanish corpus of continuous speech. Our results indicate that we are able to recognize approximately 58% of the visemes, 47% of the phonemes and 23% of the words in a continuous speech scenario and that the optimal viseme vocabulary for Spanish is composed by 20 visemes.

📄 PDF Abstract BibTeX arXiv:1704.08035

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClusteringLip Readingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Towards Lipreading Sentences with Active Appearance Models

2018-05-29

Automatic lipreading has major potential impact for speech recognition, supplementing and complementing the acoustic modality. Most attempts at lipreading have been performed on small vocabulary tasks, due to a shortfall…

Audio-Visual Speech RecognitionLipreadingspeech-recognitionSpeech Recognition+1

Comparing phonemes and visemes with DNN-based lipreading

2018-05-08 · Kwanchiva Thangthai, Helen L. Bear, Richard Harvey

There is debate if phoneme or viseme units are the most effective for a lipreading system. Some studies use phoneme units even though phonemes describe unique short sounds; other studies tried to improve lipreading accur…

DecoderLipreading

Towards Dynamic Neural Communication and Speech Neuroprosthesis Based on Viseme Decoding

2025-01-09 · Ji-Ha Park, Seo-Hyun Lee, Soowon Kim, Seong-Whan Lee

Decoding text, speech, or images from human neural signals holds promising potential both as neuroprosthesis for patients and as innovative communication tools for general users. Although neural signals contain various i…

Face Reconstruction

Learning Phonetic Context-Dependent Viseme for Enhancing Speech-Driven 3D Facial Animation

2025-07-28 · Hyung Kyu Kim, Hak Gu Kim arxiv

Speech-driven 3D facial animation aims to generate realistic facial movements synchronized with audio. Traditional methods primarily minimize reconstruction loss by aligning each frame with ground-truth. However, this fr…

Visual gesture variability between talkers in continuous visual speech

2017-10-03 · Helen L. Bear

Recent adoption of deep learning methods to the field of machine lipreading research gives us two options to pursue to improve system performance. Either, we develop end-to-end systems holistically or, we experiment to f…

Lipreading