Visual gesture variability between talkers in continuous visual speech
Recent adoption of deep learning methods to the field of machine lipreading research gives us two options to pursue to improve system performance. Either, we develop end-to-end systems holistically or, we experiment to further our understanding of the visual speech signal. The latter option is more difficult but this knowledge would enable researchers to both improve systems and apply the new knowledge to other domains such as speech therapy. One challenge in lipreading systems is the correct labeling of the classifiers. These labels map an estimated function between visemes on the lips and the phonemes uttered. Here we ask if such maps are speaker-dependent? Prior work investigated isolated word recognition from speaker-dependent (SD) visemes, we extend this to continuous speech. Benchmarked against SD results, and the isolated words performance, we test with RMAV dataset speakers and observe that with continuous speech, the trajectory between visemes has a greater negative effect on the speaker differentiation.
Code (0)
등록된 구현이 없습니다.
Tasks
LipreadingSimilar Papers 제목 키워드 기반
Leveraging Speech for Gesture Detection in Multimodal Communication
Gestures are inherent to human interaction and often complement speech in face-to-face communication, forming a multimodal communication system. An important task in gesture analysis is detecting a gesture's beginning an…
Which phoneme-to-viseme maps best improve visual-only computer lip-reading?
A critical assumption of all current visual speech recognition systems is that there are visual speech units called visemes which can be mapped to units of acoustic speech, the phonemes. Despite there being a number of p…
Lip Readingspeech-recognitionSpeech RecognitionVisual Speech RecognitionContinuous ErrP detections during multimodal human-robot interaction
Human-in-the-loop approaches are of great importance for robot applications. In the presented study, we implemented a multimodal human-robot interaction (HRI) scenario, in which a simulated robot communicates with its hu…
EEGElectroencephalogram (EEG)feature selectionA Gesture Recognition System for Detecting Behavioral Patterns of ADHD
We present an application of gesture recognition using an extension of Dynamic Time Warping (DTW) to recognize behavioural patterns of Attention Deficit Hyperactivity Disorder (ADHD). We propose an extension of DTW using…
Dynamic Time WarpingGesture RecognitionPrompt-to-Gesture: Measuring the Capabilities of Image-to-Video Deictic Gesture Generation
Gesture recognition research, unlike NLP, continues to face acute data scarcity, with progress constrained by the need for costly human recordings or image processing approaches that cannot generate authentic variability…
Gesture RecognitionGesture GenerationVideo Generation