paper-with-me

Papers

Semantic-Aware Implicit Neural Audio-Driven Video Portrait Generation

2022-01-19 · Xian Liu, Yinghao Xu, Qianyi Wu, Hang Zhou, Wayne Wu, Bolei Zhou

Animating high-fidelity video portrait with speech audio is crucial for virtual reality and digital entertainment. While most previous studies rely on accurate explicit structural information, recent works explore the implicit scene representation of Neural Radiance Fields (NeRF) for realistic generation. In order to capture the inconsistent motions as well as the semantic difference between human head and torso, some work models them via two individual sets of NeRF, leading to unnatural results. In this work, we propose Semantic-aware Speaking Portrait NeRF (SSP-NeRF), which creates delicate audio-driven portraits using one unified set of NeRF. The proposed model can handle the detailed local facial semantics and the global head-torso relationship through two semantic-aware modules. Specifically, we first propose a Semantic-Aware Dynamic Ray Sampling module with an additional parsing branch that facilitates audio-driven volume rendering. Moreover, to enable portrait rendering in one unified neural radiance field, a Torso Deformation module is designed to stabilize the large-scale non-rigid torso motions. Extensive evaluations demonstrate that our proposed approach renders more realistic video portraits compared to previous methods. Project page: https://alvinliu0.github.io/projects/SSP-NeRF

📄 PDF Abstract BibTeX arXiv:2201.07786

Code (0)

등록된 구현이 없습니다.

Tasks

NeRF

Similar Papers 제목 키워드 기반

AudioScenic: Audio-Driven Video Scene Editing

2024-04-25 · Kaixin Shen, Ruijie Quan, Linchao Zhu, Jun Xiao 외

Audio-driven visual scene editing endeavors to manipulate the visual background while leaving the foreground content unchanged, according to the given audio signals. Unlike current efforts focusing primarily on image edi…

SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation

2025-08-01 · Kien T. Pham, Yingqing He, Yazhou Xing, Qifeng Chen 외 arxiv

Audio-driven video generation aims to synthesize realistic videos that align with input audio recordings, akin to the human ability to visualize scenes from auditory input. However, existing approaches predominantly focu…

Video Generation

YingVideo-MV: Music-Driven Multi-Stage Video Generation

2025-12-02 · Jiahui Chen, Weida Wang, Runhua Shi, Huan Yang 외 arxiv

While diffusion model for audio-driven avatar video generation have achieved notable process in synthesizing long sequences with natural audio-visual synchronization and identity consistency, the generation of music-perf…

Video Generation

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation

2026-08-01 · Chenggong Hu, Shaoyin Ma, Yi Wang, Li Sun 외 arxiv

Audio-driven emotional talking face generation aims to synthesize realistic videos with expressive facial dynamics. However, existing methods struggle to balance controllability and visual fidelity. Although implicit rep…

Talking Face GenerationContinuous Control

Mamba-Enhanced Implicit Motion Learning for Audio-Driven Portrait Animation

2026-06-02 · Xuan Wei, Jiahui Chen, Kaiheng Li, Mingyu Shao 외 arxiv

Audio-driven human motion video generation aims to synthesize realistic and temporally coherent human animations from a single static image, with applications in talking-head synthesis, co-speech gesture generation, and …

Gesture GenerationVideo Generation