paper-with-me

홈 › Papers

Gesture2Speech: How Far Can Hand Movements Shape Expressive Speech?

2026-03-20 · Lokesh Kumar, Nirmesh Shah, Ashishkumar P. Gudmalwar, Pankaj Wasnik arxiv

Human communication seamlessly integrates speech and bodily motion, where hand gestures naturally complement vocal prosody to express intent, emotion, and emphasis. While recent text-to-speech (TTS) systems have begun incorporating multimodal cues such as facial expressions or lip movements, the role of hand gestures in shaping prosody remains largely underexplored. We propose a novel multimodal TTS framework, Gesture2Speech, that leverages visual gesture cues to modulate prosody in synthesized speech. Motivated by the observation that confident and expressive speakers coordinate gestures with vocal prosody, we introduce a multimodal Mixture-of-Experts (MoE) architecture that dynamically fuses linguistic content and gesture features within a dedicated style extraction module. The fused representation conditions an LLM-based speech decoder, enabling prosodic modulation that is temporally aligned with hand movements. We further design a gesture-speech alignment loss that explicitly models their temporal correspondence to ensure fine-grained synchrony between gestures and prosodic contours. Evaluations on the PATS dataset show that Gesture2Speech outperforms state-of-the-art baselines in both speech naturalness and gesture-speech synchrony. To the best of our knowledge, this is the first work to utilize hand gesture cues for prosody control in neural speech synthesis. Demo samples are available at https://research.sri-media-analysis.com/aaai26-beeu-gesture2speech/

📄 PDF Abstract BibTeX arXiv:2603.19831

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling

2023-12-31 · CVPR 2024 1 · Haiyang Liu, Zihao Zhu, Giorgio Becherini, Yichen Peng 외

We propose EMAGE, a framework to generate full-body human gestures from audio and masked gestures, encompassing facial, local body, hands, and global movements. To achieve this, we first introduce BEAT2 (BEAT-SMPLX-FLAME…

3D Face AnimationDiversityGesture GenerationRhythm

EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

2025-01-18 · Linrui Tian, Siqi Hu, Qi Wang, Bang Zhang 외

In this paper, we propose a novel audio-driven talking head method capable of simultaneously generating highly expressive facial expressions and hand gestures. Unlike existing methods that focus on generating full-body o…

Gesture GenerationVideo Generation

Are Gestures Worth a Thousand Words? An Analysis of Interviews in the Political Domain

2021-06-01 · ACL (mmsr, IWCS) 2021 6 · Daniela Trotta, Sara Tonelli

Speaker gestures are semantically co-expressive with speech and serve different pragmatic functions to accompany oral modality. Therefore, gestures are an inseparable part of the language system: they may add clarity to …

Retrieval

Generating Natural and Expressive Robot Gestures through Iterative Reinforcement Learning with Human Feedback using LLMs

2026-06-17 · Chris Lee, Flora Salim, Benjamin Tag, Francisco Cruz arxiv

Expressive gestures are essential for natural and effective communication, complementing speech when verbal cues alone are insufficient (e.g., pointing). For social robots such as the humanoid Pepper, producing natural a…

Reinforcement LearningGesture GenerationCode Generation

FineHand: Learning Hand Shapes for American Sign Language Recognition

2020-03-04 · Al Amin Hosain, Panneer Selvam Santhalingam, Parth Pathak, Huzefa Rangwala 외

American Sign Language recognition is a difficult gesture recognition problem, characterized by fast, highly articulate gestures. These are comprised of arm movements with different hand shapes, facial expression and hea…

Gesture RecognitionSign Language Recognition