paper-with-me

Papers

Audio-Synchronized Visual Animation

2024-03-08 · Lin Zhang, Shentong Mo, Yijing Zhang, Pedro Morgado

Current visual generation methods can produce high quality videos guided by texts. However, effectively controlling object dynamics remains a challenge. This work explores audio as a cue to generate temporally synchronized image animations. We introduce Audio Synchronized Visual Animation (ASVA), a task animating a static image to demonstrate motion dynamics, temporally guided by audio clips across multiple classes. To this end, we present AVSync15, a dataset curated from VGGSound with videos featuring synchronized audio visual events across 15 categories. We also present a diffusion model, AVSyncD, capable of generating dynamic animations guided by audios. Extensive evaluations validate AVSync15 as a reliable benchmark for synchronized generation and demonstrate our models superior performance. We further explore AVSyncDs potential in a variety of audio synchronized generation tasks, from generating full videos without a base image to controlling object motions with various sounds. We hope our established benchmark can open new avenues for controllable visual generation. More videos on project webpage https://lzhangbj.github.io/projects/asva/asva.html.

📄 PDF Abstract BibTeX arXiv:2403.05659

Code (1)

lzhangbj/ASVA 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
BASE 설명 없음

Similar Papers 제목 키워드 기반

Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm

2025-08-05 · Lin Zhang, Zefan Cai, Yufan Zhou, Shentong Mo 외 arxiv

Recent advances in audio-synchronized visual animation enable control of video content using audios from specific classes. However, existing methods rely heavily on expensive manual curation of high-quality, class-specif…

Speaker-Independent Speech-Driven Visual Speech Synthesis using Domain-Adapted Acoustic Models

2019-05-15 · Ahmed Hussen Abdelaziz, Barry-John Theobald, Justin Binder, Gabriele Fanelli 외

Speech-driven visual speech synthesis involves mapping features extracted from acoustic speech to the corresponding lip animation controls for a face model. This mapping can take many forms, but a powerful approach is to…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Face Modelspeech-recognition+2

SyncAnimation: A Real-Time End-to-End Framework for Audio-Driven Human Pose and Talking Head Animation

2025-01-24 · Yujian Liu, Shidang Xu, Jing Guo, Dingbin Wang 외

Generating talking avatar driven by audio remains a significant challenge. Existing methods typically require high computational costs and often lack sufficient facial detail and realism, making them unsuitable for appli…

NeRF

LinguaLinker: Audio-Driven Portraits Animation with Implicit Facial Control Enhancement

2024-07-26 · Rui Zhang, Yixiao Fang, Zhengnan Lu, Pei Cheng 외

This study delves into the intricacies of synchronizing facial dynamics with multilingual audio inputs, focusing on the creation of visually compelling, time-synchronized animations through diffusion-based techniques. Di…

MMFace4D: A Large-Scale Multi-Modal 4D Face Dataset for Audio-Driven 3D Face Animation

2023-03-17 · Haozhe Wu, Jia Jia, Junliang Xing, Hongwei Xu 외

Audio-Driven Face Animation is an eagerly anticipated technique for applications such as VR/AR, games, and movie making. With the rapid development of 3D engines, there is an increasing demand for driving 3D faces with a…

3D Face Animation