paper-with-me

홈 › Papers

OpFlowTalker: Realistic and Natural Talking Face Generation via Optical Flow Guidance

2024-05-23 · Shuheng Ge, Haoyu Xing, Li Zhang, Xiangqian Wu

Creating realistic, natural, and lip-readable talking face videos remains a formidable challenge. Previous research primarily concentrated on generating and aligning single-frame images while overlooking the smoothness of frame-to-frame transitions and temporal dependencies. This often compromised visual quality and effects in practical settings, particularly when handling complex facial data and audio content, which frequently led to semantically incongruent visual illusions. Specifically, synthesized videos commonly featured disorganized lip movements, making them difficult to understand and recognize. To overcome these limitations, this paper introduces the application of optical flow to guide facial image generation, enhancing inter-frame continuity and semantic consistency. We propose "OpFlowTalker", a novel approach that utilizes predicted optical flow changes from audio inputs rather than direct image predictions. This method smooths image transitions and aligns changes with semantic content. Moreover, it employs a sequence fusion technique to replace the independent generation of single frames, thus preserving contextual information and maintaining temporal coherence. We also developed an optical flow synchronization module that regulates both full-face and lip movements, optimizing visual synthesis by balancing regional dynamics. Furthermore, we introduce a Visual Text Consistency Score (VTCS) that accurately measures lip-readability in synthesized videos. Extensive empirical evidence validates the effectiveness of our approach.

📄 PDF Abstract BibTeX arXiv:2405.14709

Code (0)

등록된 구현이 없습니다.

Tasks

Face GenerationImage GenerationOptical Flow EstimationTalking Face Generation

Similar Papers 제목 키워드 기반

Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

2024-01-16 · Zhenhui Ye, Tianyun Zhong, Yi Ren, Jiaqi Yang 외

One-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simulta…

3D ReconstructionFace GenerationSuper-ResolutionTalking Face Generation

Audio-Driven Talking Face Generation with Diverse yet Realistic Facial Animations

2023-04-18 · Rongliang Wu, Yingchen Yu, Fangneng Zhan, Jiahui Zhang 외

Audio-driven talking face generation, which aims to synthesize talking faces with realistic facial animations (including accurate lip movements, vivid facial expression details and natural head poses) corresponding to th…

Face GenerationTalking Face Generation

Audio-driven Talking Face Video Generation with Learning-based Personalized Head Pose

2020-02-24 · Ran Yi, Zipeng Ye, Juyong Zhang, Hujun Bao 외

Real-world talking faces often accompany with natural head movement. However, most existing talking face video generation methods only consider facial animation with fixed head pose. In this paper, we address this proble…

3D Face AnimationVideo Generation

CP-EB: Talking Face Generation with Controllable Pose and Eye Blinking Embedding

2023-11-15 · Jianzong Wang, Yimin Deng, ZiQi Liang, xulong Zhang 외

This paper proposes a talking face generation method named "CP-EB" that takes an audio signal as input and a person image as reference, to synthesize a photo-realistic people talking video with head poses controlled by a…

Face GenerationTalking Face Generation

Text-Driven Emotionally Continuous Talking Face Generation

2026-03-06 · Hao Yang, Yanyan Zhao, Tian Zheng, Hongbo Zhang 외 arxiv

Talking Face Generation (TFG) strives to create realistic and emotionally expressive digital faces. While previous TFG works have mastered the creation of naturalistic facial movements, they typically express a fixed tar…

Talking Face Generation