paper-with-me

홈 › Papers

Cross-Modal Mutual Learning for Cued Speech Recognition

2022-12-02 · Lei Liu, Li Liu

Automatic Cued Speech Recognition (ACSR) provides an intelligent human-machine interface for visual communications, where the Cued Speech (CS) system utilizes lip movements and hand gestures to code spoken language for hearing-impaired people. Previous ACSR approaches often utilize direct feature concatenation as the main fusion paradigm. However, the asynchronous modalities i.e., lip, hand shape and hand position) in CS may cause interference for feature concatenation. To address this challenge, we propose a transformer based cross-modal mutual learning framework to prompt multi-modal interaction. Compared with the vanilla self-attention, our model forces modality-specific information of different modalities to pass through a modality-invariant codebook, collating linguistic representations for tokens of each modality. Then the shared linguistic knowledge is used to re-synchronize multi-modal sequences. Moreover, we establish a novel large-scale multi-speaker CS dataset for Mandarin Chinese. To our knowledge, this is the first work on ACSR for Mandarin Chinese. Extensive experiments are conducted for different languages i.e., Chinese, French, and British English). Results demonstrate that our model exhibits superior recognition performance to the state-of-the-art by a large margin.

📄 PDF Abstract BibTeX arXiv:2212.01083

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Computation and Parameter Efficient Multi-Modal Fusion Transformer for Cued Speech Recognition

2024-01-31 · Lei Liu, Li Liu, Haizhou Li

Cued Speech (CS) is a pure visual coding method used by hearing-impaired people that combines lip reading with several specific hand shapes to make the spoken language visible. Automatic CS recognition (ACSR) seeks to tr…

Lip Readingspeech-recognitionSpeech Recognition

Cued-Agent: A Collaborative Multi-Agent System for Automatic Cued Speech Recognition

2025-08-01 · Guanjie Huang, Danny H. K. Tsang, Shan Yang, Guangzhi Lei 외 arxiv

Cued Speech (CS) is a visual communication system that combines lip-reading with hand coding to facilitate communication for individuals with hearing impairments. Automatic CS Recognition (ACSR) aims to convert CS hand g…

Speech Recognition

Re-synchronization using the Hand Preceding Model for Multi-modal Fusion in Automatic Continuous Cued Speech Recognition

2020-02-23

Cued Speech (CS) is an augmented lip reading complemented by hand coding, and it is very helpful to the deaf people. Automatic CS recognition can help communications between the deaf people and others. Due to the asynchr…

Lip ReadingPhoneme RecognitionPositionspeech-recognition+1

Investigating the dynamics of hand and lips in French Cued Speech using attention mechanisms and CTC-based decoding

2023-06-14 · Sanjana Sankar, Denis Beautemps, Frédéric Elisei, Olivier Perrotin 외

Hard of hearing or profoundly deaf people make use of cued speech (CS) as a communication tool to understand spoken language. By delivering cues that are relevant to the phonetic information, CS offers a way to enhance l…

Lipreading

Cued Speech Generation Leveraging a Pre-trained Audiovisual Text-to-Speech Model

2025-01-08 · Sanjana Sankar, Martin Lenglet, Gerard Bailly, Denis Beautemps 외

This paper presents a novel approach for the automatic generation of Cued Speech (ACSG), a visual communication system used by people with hearing impairment to better elicit the spoken language. We explore transfer lear…

text-to-speechText to SpeechTransfer Learning