paper-with-me

Papers

Re-synchronization using the Hand Preceding Model for Multi-modal Fusion in Automatic Continuous Cued Speech Recognition

2020-02-23

Cued Speech (CS) is an augmented lip reading complemented by hand coding, and it is very helpful to the deaf people. Automatic CS recognition can help communications between the deaf people and others. Due to the asynchronous nature of lips and hand movements, fusion of them in automatic CS recognition is a challenging problem. In this work, we propose a novel re-synchronization procedure for multi-modal fusion, which aligns the hand features with lips feature. It is realized by delaying hand position and hand shape with their optimal hand preceding time which is derived by investigating the temporal organizations of hand position and hand shape movements in CS. This re-synchronization procedure is incorporated into a practical continuous CS recognition system that combines convolutional neural network (CNN) with multi-stream hidden markov model (MSHMM). A significant improvement of about 4.6\% has been achieved retaining 76.6\% CS phoneme recognition correctness compared with the state-of-the-art architecture (72.04\%), which did not take into account the asynchrony of multi-modal fusion in CS. To our knowledge, this is the first work to tackle the asynchronous multi-modal fusion in the automatic continuous CS recognition.

📄 PDF Abstract BibTeX arXiv:2001.00854

Code (0)

등록된 구현이 없습니다.

Tasks

Lip ReadingPhoneme RecognitionPositionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

A Novel Interpretable and Generalizable Re-synchronization Model for Cued Speech based on a Multi-Cuer Corpus

2023-06-05 · Lufei Gao, Shan Huang, Li Liu

Cued Speech (CS) is a multi-modal visual coding system combining lip reading with several hand cues at the phonetic level to make the spoken language visible to the hearing impaired. Previous studies solved asynchronous …

Lip Reading

A-MESS: Anchor based Multimodal Embedding with Semantic Synchronization for Multimodal Intent Recognition

2025-03-25 · Yaomin Shen, Xiaojian Lin, Wei Fan

In the domain of multimodal intent recognition (MIR), the objective is to recognize human intent by integrating a variety of modalities, such as language text, body gestures, and tones. However, existing approaches face …

Contrastive LearningIntent RecognitionLanguage ModelingLanguage Modelling+3

Variational Test-time Optimization for Diffusion Synchronization

2026-06-14 · Hyunsoo Lee, Farrin Marouf Sofian, Kushagra Pandey, Stephan Mandt arxiv

Collaborative generation, which coordinates multiple diffusion trajectories to extend the capabilities of pretrained priors, has emerged as a powerful paradigm for extending the applicability of diffusion models. Among e…

Multimodal Emotion Recognition using Audio-Video Transformer Fusion with Cross Attention

2024-07-26 · Joe Dhanith P R, Shravan Venkatraman, Vigya Sharma, Santhosh Malarvannan 외

Understanding emotions is a fundamental aspect of human communication. Integrating audio and video signals offers a more comprehensive understanding of emotional states compared to traditional methods that rely on a sing…

Emotion RecognitionMultimodal Emotion Recognition

OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation

2026-05-12 · Guohui Zhang, XiaoXiao Ma, Jie Huang, Hang Xu 외 arxiv

Recent advances in joint audio-video generation have been remarkable, yet real-world applications demand strong per-modality fidelity, cross-modal alignment, and fine-grained synchronization. Reinforcement Learning (RL) …

Reinforcement LearningVideo Generation