paper-with-me

Papers

Instruct-NeuralTalker: Editing Audio-Driven Talking Radiance Fields with Instructions

2023-06-19 · Yuqi Sun, Ruian He, Weimin Tan, Bo Yan

Recent neural talking radiance field methods have shown great success in photorealistic audio-driven talking face synthesis. In this paper, we propose a novel interactive framework that utilizes human instructions to edit such implicit neural representations to achieve real-time personalized talking face generation. Given a short speech video, we first build an efficient talking radiance field, and then apply the latest conditional diffusion model for image editing based on the given instructions and guiding implicit representation optimization towards the editing target. To ensure audio-lip synchronization during the editing process, we propose an iterative dataset updating strategy and utilize a lip-edge loss to constrain changes in the lip region. We also introduce a lightweight refinement network for complementing image details and achieving controllable detail generation in the final rendered image. Our method also enables real-time rendering at up to 30FPS on consumer hardware. Multiple metrics and user verification show that our approach provides a significant improvement in rendering quality compared to state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2306.10813

Code (0)

등록된 구현이 없습니다.

Tasks

Face GenerationTalking Face Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Audio-driven High-resolution Seamless Talking Head Video Editing via StyleGAN

2024-07-08 · Jiacheng Su, KunHong Liu, Liyan Chen, Junfeng Yao 외

The existing methods for audio-driven talking head video editing have the limitations of poor visual effects. This paper tries to tackle this problem through editing talking face images seamless with different emotions b…

DisentanglementVideo Editing

JoyGen: Audio-Driven 3D Depth-Aware Talking-Face Video Editing

2025-01-03 · Qili Wang, Dajiang Wu, Zihang Xu, Junshi Huang 외

Significant progress has been made in talking-face video generation research; however, precise lip-audio synchronization and high visual quality remain challenging in editing lip shapes based on input audio. This paper i…

3D ReconstructionFace GenerationMotion GenerationTalking Face Generation+2

AVI-Talking: Learning Audio-Visual Instructions for Expressive 3D Talking Face Generation

2024-02-25 · Yasheng Sun, Wenqing Chu, Hang Zhou, Kaisiyuan Wang 외

While considerable progress has been made in achieving accurate lip synchronization for 3D speech-driven talking face generation, the task of incorporating expressive facial detail synthesis aligned with the speaker's sp…

Face GenerationHallucinationTalking Face Generation

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

2025-05-28 · Zhe Kong, Feng Gao, Yong Zhang, Zhuoliang Kang 외

Audio-driven human animation methods, such as talking head and talking body generation, have made remarkable progress in generating synchronized facial movements and appealing visual quality videos. However, existing met…

Human AnimationInstruction FollowingVideo Generation

VideoReTalking: Audio-based Lip Synchronization for Talking Head Video Editing In the Wild

2022-11-27 · Kun Cheng, Xiaodong Cun, Yong Zhang, Menghan Xia 외

We present VideoReTalking, a new system to edit the faces of a real-world talking head video according to input audio, producing a high-quality and lip-syncing output video even with a different emotion. Our system disen…

Video EditingVideo Generation