paper-with-me

홈 › Papers

Audio to Body Dynamics

2017-12-19 · CVPR 2018 6 · Eli Shlizerman, Lucio M. Dery, Hayden Schoen, Ira Kemelmacher-Shlizerman

We present a method that gets as input an audio of violin or piano playing, and outputs a video of skeleton predictions which are further used to animate an avatar. The key idea is to create an animation of an avatar that moves their hands similarly to how a pianist or violinist would do, just from audio. Aiming for a fully detailed correct arms and fingers motion is a goal, however, it's not clear if body movement can be predicted from music at all. In this paper, we present the first result that shows that natural body dynamics can be predicted at all. We built an LSTM network that is trained on violin and piano recital videos uploaded to the Internet. The predicted points are applied onto a rigged avatar to create the animation.

📄 PDF Abstract BibTeX arXiv:1712.09382

Code (1)

facebookresearch/Audio2BodyDynamics pytorch

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

To React or not to React: End-to-End Visual Pose Forecasting for Personalized Avatar during Dyadic Conversations

2019-10-05 · Chaitanya Ahuja, Shugao Ma, Louis-Philippe Morency, Yaser Sheikh

Non verbal behaviours such as gestures, facial expressions, body posture, and para-linguistic cues have been shown to complement or clarify verbal messages. Hence to improve telepresence, in form of an avatar, it is impo…

AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers

2025-03-25 · CVPR 2025 1 · Jiazhi Guan, Kaisiyuan Wang, Zhiliang Xu, Quanwei Yang 외

Despite the recent progress of audio-driven video generation, existing methods mostly focus on driving facial movements, leading to non-coherent head and body dynamics. Moving forward, it is desirable yet challenging to …

Video Generation

Music Gesture for Visual Sound Separation

2020-04-20 · CVPR 2020 6 · Chuang Gan, Deng Huang, Hang Zhao, Joshua B. Tenenbaum 외

Recent deep learning approaches have achieved impressive performance on visual sound separation tasks. However, these approaches are mostly built on appearance and optical flow like motion feature representations, which …

Optical Flow Estimation

LiveGesture Streamable Co-Speech Gesture Generation Model

2026-04-13 · Muhammad Usama Saleem, Mayur Jagdishbhai Patel, Ekkasit Pinyoanuntapong, Zhongxing Qin 외 arxiv

We propose LiveGesture, the first fully streamable, speech-driven full-body gesture generation framework that operates with zero look-ahead and supports arbitrary sequence length. Unlike existing co-speech gesture method…

Gesture Generation

VividAnimator: An End-to-End Audio and Pose-driven Half-Body Human Animation Framework

2025-10-11 · Donglin Huang, Yongyuan Li, Tianhang Liu, Junming Huang 외 arxiv

Existing for audio- and pose-driven human animation methods often struggle with stiff head movements and blurry hands, primarily due to the weak correlation between audio and head movements and the structural complexity …