paper-with-me

Papers

EfficientSync: Real-Time Lip Synchronization via Deformation-Based Reference Texture Mixing

2026-08-19 · Fa-Ting Hong, Runzhen Liu, Luchuan Song, Hongmin Cai, Chuhua Xian arxiv

Audio-driven lip synchronization manipulates the mouth region of a talking-face video to match the driving audio while preserving head pose, identity, and background. Although the task is inherently local editing, prevailing approaches reconstruct the entire lower face with heavy GAN- or diffusion-based decoders, incurring substantial latency and, more critically, hallucinating intra-oral details such as teeth and lip wrinkles instead of preserving authentic textures. We contend that the bottleneck in identity preservation is not the scarcity of reference frames, but the lack of a mechanism that faithfully transfers the genuine textures they already contain. We therefore present EfficientSync, a real-time deformation-based framework that retains reference textures rather than resynthesizing them. First, the Dynamic Texture Mixer reformulates multi-reference fusion as channel-wise selection, evaluating each spatially aligned reference in a global context and aggregating them by channel-wise weighted summation, preserving textural integrity at low cost. Second, Spatio-Temporal Shifted Adaptive Masking decomposes the source frame into lip-generation conditions and an independent background prior, suppressing lower-face leakage while blending the synthesized mouth seamlessly into the background. Third, STAR Sampling, a zero-overhead pre-processing step, retrieves the sharpest and most topologically diverse reference frames. Experiments on HDTF and VFHQ show state-of-the-art visual quality and identity preservation at 166 FPS on a single GPU. Video demos: https://alunaticat.github.io/EfficientSync/index.html.

📄 PDF Abstract BibTeX arXiv:2608.18832

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization

2026-06-09 · Paul Hyunbin Cho, Jinhyuk Jang, SeokYoung Lee, Joungbin Lee 외 arxiv

Diffusion-based lip synchronization models achieve strong visual quality and audio-visual alignment, but full-sequence bidirectional attention and many denoising steps make them impractical for real-time inference. We pr…

GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian Splatting

2026-07-01 · Haijie Yang, Zhenyu Zhang, Yixuan Dong, Jianjun Qian 외 arxiv

Audio-driven talking head synthesis has achieved impressive progress in lip synchronization and visual quality, yet generating expressive emotional avatars with controllable intensity remains challenging, especially unde…

Audio-driven Talking Face Generation with Stabilized Synchronization Loss

2023-07-18 · Dogucan Yaman, Fevziye Irem Eyiokur, Leonard Bärmann, Hazim Kemal Ekenel 외

Talking face generation aims to create realistic videos with accurate lip synchronization and high visual quality, using given audio and reference video while preserving identity and visual characteristics. In this paper…

Audio-Visual SynchronizationFace GenerationTalking Face Generation

D^3-Talker: Dual-Branch Decoupled Deformation Fields for Few-Shot 3D Talking Head Synthesis

2025-08-20 · Yuhang Guo, Kaijun Deng, Siyang Song, Jindong Xie 외 arxiv

A key challenge in 3D talking head synthesis lies in the reliance on a long-duration talking head video to train a new model for each target identity from scratch. Recent methods have attempted to address this issue by e…

An Application of Model Reference Adaptive Control for Multi-Agent Synchronization in Drone Networks

2024-06-30 · Miguel F. Arevalo-Castiblanco, Yejin Wi, Marzia Cescon and, Cesar A. Uribe

This paper presents the application of a Distributed Model Reference Adaptive Control (DMRAC) strategy for robust multi-agent synchronization of a network of drones. The proposed approach enables the development of contr…