paper-with-me

Papers

Rhythmic Gesticulator: Rhythm-Aware Co-Speech Gesture Synthesis with Hierarchical Neural Embeddings

2022-10-04 · Tenglong Ao, Qingzhe Gao, Yuke Lou, Baoquan Chen, Libin Liu

Automatic synthesis of realistic co-speech gestures is an increasingly important yet challenging task in artificial embodied agent creation. Previous systems mainly focus on generating gestures in an end-to-end manner, which leads to difficulties in mining the clear rhythm and semantics due to the complex yet subtle harmony between speech and gestures. We present a novel co-speech gesture synthesis method that achieves convincing results both on the rhythm and semantics. For the rhythm, our system contains a robust rhythm-based segmentation pipeline to ensure the temporal coherence between the vocalization and gestures explicitly. For the gesture semantics, we devise a mechanism to effectively disentangle both low- and high-level neural embeddings of speech and motion based on linguistic theory. The high-level embedding corresponds to semantics, while the low-level embedding relates to subtle variations. Lastly, we build correspondence between the hierarchical embeddings of the speech and the motion, resulting in rhythm- and semantics-aware gesture synthesis. Evaluations with existing objective metrics, a newly proposed rhythmic metric, and human feedback show that our method outperforms state-of-the-art systems by a clear margin.

📄 PDF Abstract BibTeX arXiv:2210.01448

Code (1)

aubrey-ao/humanbehavioranimation 공식 구현 pytorch

Tasks

Gesture GenerationRhythm

Similar Papers 제목 키워드 기반

LLM Gesticulator: Leveraging Large Language Models for Scalable and Controllable Co-Speech Gesture Synthesis

2024-10-06 · Haozhou Pang, Tianwei Ding, Lanshan He, Ming Tao 외

In this work, we present LLM Gesticulator, an LLM-based audio-driven co-speech gesture generation framework that synthesizes full-body animations that are rhythmically aligned with the input audio while exhibiting natura…

Gesture Generation

Semantic Gesticulator: Semantics-Aware Co-Speech Gesture Synthesis

2024-05-16 · Zeyi Zhang, Tenglong Ao, Yuyao Zhang, Qingzhe Gao 외

In this work, we present Semantic Gesticulator, a novel framework designed to synthesize realistic gestures accompanying speech with strong semantic correspondence. Semantically meaningful gestures are crucial for effect…

Language ModellingLarge Language ModelRhythmSemantic correspondence

SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis

2024-12-21 · Xiangyue Zhang, Jianfang Li, Jiaxu Zhang, Ziqiang Dang 외

A good co-speech motion generation cannot be achieved without a careful integration of common rhythmic motion and rare yet essential semantic motion. In this work, we propose SemTalk for holistic co-speech motion generat…

Gesture GenerationMotion GenerationRhythm

DuoGesture: Neuro-Inspired and Biomechanically Informed Dual-Stream Co-Speech Gesture Generation

2026-05-25 · Ferdinand Paar, Lanmiao Liu, Aslı Özyürek, Serge Thill 외 arxiv

Co-speech gesture generation requires both semantic expressivity and biomechanically plausible rhythmic motion. Existing holistic gesture models mix lexically grounded semantic gestures with frequent prosody-aligned beat…

Gesture Generation

Gesticulator: A framework for semantically-aware speech-driven gesture generation

2020-01-25 · Taras Kucherenko, Patrik Jonell, Sanne van Waveren, Gustav Eje Henter 외

During speech, people spontaneously gesticulate, which plays a key role in conveying information. Similarly, realistic co-speech gestures are crucial to enable natural and smooth interactions with social agents. Current …

Gesture Generation