Investigating the dynamics of hand and lips in French Cued Speech using attention mechanisms and CTC-based decoding
Hard of hearing or profoundly deaf people make use of cued speech (CS) as a communication tool to understand spoken language. By delivering cues that are relevant to the phonetic information, CS offers a way to enhance lipreading. In literature, there have been several studies on the dynamics between the hand and the lips in the context of human production. This article proposes a way to investigate how a neural network learns this relation for a single speaker while performing a recognition task using attention mechanisms. Further, an analysis of the learnt dynamics is utilized to establish the relationship between the two modalities and extract automatic segments. For the purpose of this study, a new dataset has been recorded for French CS. Along with the release of this dataset, a benchmark will be reported for word-level recognition, a novelty in the automatic recognition of French CS.
Code (0)
등록된 구현이 없습니다.
Tasks
LipreadingSimilar Papers 제목 키워드 기반
Re-synchronization using the Hand Preceding Model for Multi-modal Fusion in Automatic Continuous Cued Speech Recognition
Cued Speech (CS) is an augmented lip reading complemented by hand coding, and it is very helpful to the deaf people. Automatic CS recognition can help communications between the deaf people and others. Due to the asynchr…
Lip ReadingPhoneme RecognitionPositionspeech-recognition+1Multistream neural architectures for cued-speech recognition using a pre-trained visual feature extractor and constrained CTC decoding
This paper proposes a simple and effective approach for automatic recognition of Cued Speech (CS), a visual communication tool that helps people with hearing impairment to understand spoken language with the help of hand…
DecoderLipreadingspeech-recognitionSpeech RecognitionCLeLfPC: a Large Open Multi-Speaker Corpus of French Cued Speech
Cued Speech is a communication system developed for deaf people to complement speechreading at the phonetic level with hands. This visual communication mode uses handshapes in different placements near the face in combin…
TransliterationMapping de l'espace spectral vers l'espace visuel de la parole : les voyelles du fran\ccais en langue fran\ccaise parl\'ee compl\'et\'ee (Mapping of the spectral space to the visual speech space for French vowels cued in Cued Speech) [in French]
A Novel Interpretable and Generalizable Re-synchronization Model for Cued Speech based on a Multi-Cuer Corpus
Cued Speech (CS) is a multi-modal visual coding system combining lip reading with several hand cues at the phonetic level to make the spoken language visible to the hearing impaired. Previous studies solved asynchronous …
Lip Reading