paper-with-me

Papers

A Neural Lip-Sync Framework for Synthesizing Photorealistic Virtual News Anchors

2020-02-20 · Ruobing Zheng, Zhou Zhu, Bo Song, Changjiang Ji

Lip sync has emerged as a promising technique for generating mouth movements from audio signals. However, synthesizing a high-resolution and photorealistic virtual news anchor is still challenging. Lack of natural appearance, visual consistency, and processing efficiency are the main problems with existing methods. This paper presents a novel lip-sync framework specially designed for producing high-fidelity virtual news anchors. A pair of Temporal Convolutional Networks are used to learn the cross-modal sequential mapping from audio signals to mouth movements, followed by a neural rendering network that translates the synthetic facial map into a high-resolution and photorealistic appearance. This fully trainable framework provides end-to-end processing that outperforms traditional graphics-based methods in many low-delay applications. Experiments also show the framework has advantages over modern neural-based methods in both visual appearance and efficiency.

📄 PDF Abstract BibTeX arXiv:2002.08700

Code (0)

등록된 구현이 없습니다.

Tasks

Constrained Lip-synchronizationImage-to-Image TranslationNeural Rendering

Similar Papers 제목 키워드 기반

Synthesizing Photorealistic Virtual Humans Through Cross-modal Disentanglement

2022-09-03 · CVPR 2023 1 · Siddarth Ravichandran, Ondřej Texler, Dimitar Dinev, Hyun Jae Kang

Over the last few decades, many aspects of human life have been enhanced with virtual domains, from the advent of digital assistants such as Amazon's Alexa and Apple's Siri to the latest metaverse efforts of the rebrande…

Data AugmentationDisentanglementFace SwappingImage Generation+1

TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy

2026-06-24 · Hao Sun, Hao Yan, Mengting Chen, Quanjian Song 외 arxiv

While Video Virtual Try-on (VVT) has achieved remarkable progress in synthesizing realistic garment overlays on dynamic subjects, existing paradigms remains fundamentally constrained by a passive dependency on source cam…

Virtual Try-on

Pose Guided Fashion Image Synthesis Using Deep Generative Model

2019-06-17 · Wei Sun, Jawadul H. Bappy, Shanglin Yang, Yi Xu 외

Generating a photorealistic image with intended human pose is a promising yet challenging research topic for many applications such as smart photo editing, movie making, virtual try-on, and fashion display. In this paper…

DecoderImage GenerationPose-Guided Image GenerationVirtual Try-on

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation

2026-04-26 · Chunyu Li, Jiaye Li, Ruiqiao Mei, Haoyuan Xia 외 arxiv

Real-time text-driven joint audio-video avatar generation requires jointly synthesizing portrait video and speech with high fidelity and precise synchronization, yet existing audio-visual diffusion models remain too slow…

ICo3D: An Interactive Conversational 3D Virtual Human

2026-01-19 · Richard Shaw, Youngkyoon Jang, Athanasios Papaioannou, Arthur Moreau 외 arxiv

This work presents Interactive Conversational 3D Virtual Human (ICo3D), a method for generating an interactive, conversational, and photorealistic 3D human avatar. Based on multi-view captures of a subject, we create an …