paper-with-me

Papers

Audio-Driven Talking Face Video Generation with Dynamic Convolution Kernels

2022-01-16 · Zipeng Ye, Mengfei Xia, Ran Yi, Juyong Zhang, Yu-Kun Lai, Xuwei Huang, Guoxin Zhang, Yong-Jin Liu

In this paper, we present a dynamic convolution kernel (DCK) strategy for convolutional neural networks. Using a fully convolutional network with the proposed DCKs, high-quality talking-face video can be generated from multi-modal sources (i.e., unmatched audio and video) in real time, and our trained model is robust to different identities, head postures, and input audios. Our proposed DCKs are specially designed for audio-driven talking face video generation, leading to a simple yet effective end-to-end system. We also provide a theoretical analysis to interpret why DCKs work. Experimental results show that our method can generate high-quality talking-face video with background at 60 fps. Comparison and evaluation between our method and the state-of-the-art methods demonstrate the superiority of our method.

📄 PDF Abstract BibTeX arXiv:2201.05986

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

JoyGen: Audio-Driven 3D Depth-Aware Talking-Face Video Editing

2025-01-03 · Qili Wang, Dajiang Wu, Zihang Xu, Junshi Huang 외

Significant progress has been made in talking-face video generation research; however, precise lip-audio synchronization and high visual quality remain challenging in editing lip shapes based on input audio. This paper i…

3D ReconstructionFace GenerationMotion GenerationTalking Face Generation+2

Speech Driven Talking Face Generation from a Single Image and an Emotion Condition

2020-08-08 · Sefik Emre Eskimez, You Zhang, Zhiyao Duan

Visual emotion expression plays an important role in audiovisual speech communication. In this work, we propose a novel approach to rendering visual emotion expression in speech-driven talking face generation. Specifical…

Emotion RecognitionFace GenerationTalking Face Generation

One-shot Talking Face Generation from Single-speaker Audio-Visual Correlation Learning

2021-12-06 · Suzhen Wang, Lincheng Li, Yu Ding, Xin Yu

Audio-driven one-shot talking face generation methods are usually trained on video resources of various persons. However, their created videos often suffer unnatural mouth shapes and asynchronous lips because those metho…

Face GenerationTalking Face Generation

PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face Generation

2024-12-10 · Fatemeh Nazarieh, ZhenHua Feng, Diptesh Kanojia, Muhammad Awais 외

Audio-driven talking face generation is a challenging task in digital communication. Despite significant progress in the area, most existing methods concentrate on audio-lip synchronization, often overlooking aspects suc…

Face GenerationTalking Face Generation

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

2025-06-03 · Chetwin Low, Weimin WANG

In this paper, we present TalkingMachines -- an efficient framework that transforms pretrained video generation models into real-time, audio-driven character animators. TalkingMachines enables natural conversational expe…

DecoderKnowledge DistillationLanguage ModelingLanguage Modelling+2