paper-with-me

Papers

Audio-Visual Speech and Gesture Recognition by Sensors of Mobile Devices

2023-02-17 · Sensors 2023 2 · Dmitry Ryumin, Denis Ivanko, Elena Ryumina

Audio-visual speech recognition (AVSR) is one of the most promising solutions for reliable speech recognition, particularly when audio is corrupted by noise. Additional visual information can be used for both automatic lip-reading and gesture recognition. Hand gestures are a form of non-verbal communication and can be used as a very important part of modern human–computer interaction systems. Currently, audio and video modalities are easily accessible by sensors of mobile devices. However, there is no out-of-the-box solution for automatic audio-visual speech and gesture recognition. This study introduces two deep neural network-based model architectures: one for AVSR and one for gesture recognition. The main novelty regarding audio-visual speech recognition lies in fine-tuning strategies for both visual and acoustic features and in the proposed end-to-end model, which considers three modality fusion approaches: prediction-level, feature-level, and model-level. The main novelty in gesture recognition lies in a unique set of spatio-temporal features, including those that consider lip articulation information. As there are no available datasets for the combined task, we evaluated our methods on two different large-scale corpora—LRW and AUTSL—and outperformed existing methods on both audio-visual speech recognition and gesture recognition tasks. We achieved AVSR accuracy for the LRW dataset equal to 98.76% and gesture recognition rate for the AUTSL dataset equal to 98.56%. The results obtained demonstrate not only the high performance of the proposed methodology, but also the fundamental possibility of recognizing audio-visual speech and gestures by sensors of mobile devices.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Audio-Visual Speech RecognitionGesture RecognitionLip ReadingSign Language Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

Similar Papers 제목 키워드 기반

SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization

2024-06-18 · Young Jin Ahn, Jungwoo Park, Sangha Park, Jonghyun Choi 외

Visual Speech Recognition (VSR) stands at the intersection of computer vision and speech recognition, aiming to interpret spoken content from visual cues. A prominent challenge in VSR is the presence of homophenes-visual…

Landmark-based LipreadingLipreadingspeech-recognitionSpeech Recognition+1

Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers

2025-03-13 · Yasheng Sun, Zhiliang Xu, Hang Zhou, Jiazhi Guan 외

Co-speech gesture video synthesis is a challenging task that requires both probabilistic modeling of human gestures and the synthesis of realistic images that align with the rhythmic nuances of speech. To address these c…

EmotionGesture: Audio-Driven Diverse Emotional Co-Speech 3D Gesture Generation

2023-05-30 · Xingqun Qi, Chen Liu, Lincheng Li, Jie Hou 외

Generating vivid and diverse 3D co-speech gestures is crucial for various applications in animating virtual avatars. While most existing methods can generate gestures from audio directly, they usually overlook that emoti…

Gesture GenerationRhythm

Co-Speech Gesture Video Generation with Implicit Motion-Audio Entanglement

2025-01-01 · CVPR 2025 1 · Xinjie Li, Ziyi Chen, Xinlu Yu, Iek-Heng Chu 외

Co-speech gestures are essential to non-verbal communication, enhancing both the naturalness and effectiveness of human interaction. Although recent methods have made progress in generating co-speech gesture videos, …

Gesture GenerationMotion GenerationVideo Generation

Generating coherent spontaneous speech and gesture from text

2021-01-14 · Simon Alexanderson, Éva Székely, Gustav Eje Henter, Taras Kucherenko 외

Embodied human communication encompasses both verbal (speech) and non-verbal information (e.g., gesture and head movements). Recent advances in machine learning have substantially improved the technologies for generating…

Gesture GenerationMotion GenerationSpeech Synthesistext-to-speech+1