paper-with-me

홈 › Papers

A Comprehensive Multi-scale Approach for Speech and Dynamics Synchrony in Talking Head Generation

2023-07-04 · Louis Airale, Dominique Vaufreydaz, Xavier Alameda-Pineda

Animating still face images with deep generative models using a speech input signal is an active research topic and has seen important recent progress.However, much of the effort has been put into lip syncing and rendering quality while the generation of natural head motion, let alone the audio-visual correlation between head motion and speech, has often been neglected.In this work, we propose a multi-scale audio-visual synchrony loss and a multi-scale autoregressive GAN to better handle short and long-term correlation between speech and the dynamics of the head and lips.In particular, we train a stack of syncer models on multimodal input pyramids and use these models as guidance in a multi-scale generator network to produce audio-aligned motion unfolding over diverse time scales.Both the pyramid of audio-visual syncers and the generative models are trained in a low-dimensional space that fully preserves dynamics cues.The experiments show significant improvements over the state-of-the-art in head motion dynamics quality and especially in multi-scale audio-visual synchrony on a collection of benchmark datasets.

📄 PDF Abstract BibTeX arXiv:2307.03270

Code (1)

louisbearing/hmo-audio 공식 구현 pytorch

Tasks

Talking Head Generation

Similar Papers 제목 키워드 기반

Distinct Theta Synchrony across Speech Modes: Perceived, Spoken, Whispered, and Imagined

2025-11-11 · Jung-Sun Lee, Ha-Na Jo, Eunyeong Ko arxiv

Human speech production encompasses multiple modes such as perceived, overt, whispered, and imagined, each reflecting distinct neural mechanisms. Among these, theta-band synchrony has been closely associated with languag…

Multimodal Emotion Coupling via Speech-to-Facial and Bodily Gestures in Dyadic Interaction

2025-05-08 · Von Ralph Dane Marquez Herbuela, Yukie Nagai

Human emotional expression emerges through coordinated vocal, facial, and gestural signals. While speech face alignment is well established, the broader dynamics linking emotionally expressive speech to regional facial a…

Face Alignment

Spatiotemporal Emotional Synchrony in Dyadic Interactions: The Role of Speech Conditions in Facial and Vocal Affective Alignment

2025-04-29 · Von Ralph Dane Marquez Herbuela, Yukie Nagai

Understanding how humans express and synchronize emotions across multiple communication channels particularly facial expressions and speech has significant implications for emotion recognition systems and human computer …

Dynamic Time WarpingEmotion Recognition

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation

2024-12-21 · Lucas Goncalves, Prashant Mathur, Xing Niu, Brady Houston 외

Audio-Visual Speech-to-Speech Translation typically prioritizes improving translation quality and naturalness. However, an equally critical aspect in audio-visual content is lip-synchrony-ensuring that the movements of t…

Speech-to-Speech TranslationTranslation

On the Audio-visual Synchronization for Lip-to-Speech Synthesis

2023-03-01 · ICCV 2023 1 · Zhe Niu, Brian Mak

Most lip-to-speech (LTS) synthesis models are trained and evaluated under the assumption that the audio-video pairs in the dataset are perfectly synchronized. In this work, we show that the commonly used audio-visual dat…

Audio-Visual SynchronizationLip to Speech SynthesisSpeech Synthesis