paper-with-me

Papers

MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset

2024-06-20 · Kim Sung-Bin, Lee Chae-Yeon, Gihun Son, Oh Hyun-Bin, Janghoon Ju, Suekyeong Nam, Tae-Hyun Oh

Recent studies in speech-driven 3D talking head generation have achieved convincing results in verbal articulations. However, generating accurate lip-syncs degrades when applied to input speech in other languages, possibly due to the lack of datasets covering a broad spectrum of facial movements across languages. In this work, we introduce a novel task to generate 3D talking heads from speeches of diverse languages. We collect a new multilingual 2D video dataset comprising over 420 hours of talking videos in 20 languages. With our proposed dataset, we present a multilingually enhanced model that incorporates language-specific style embeddings, enabling it to capture the unique mouth movements associated with each language. Additionally, we present a metric for assessing lip-sync accuracy in multilingual settings. We demonstrate that training a 3D talking head model with our proposed dataset significantly enhances its multilingual performance. Codes and datasets are available at https://multi-talk.github.io/.

📄 PDF Abstract BibTeX arXiv:2406.14272

Code (0)

등록된 구현이 없습니다.

Tasks

Talking Head Generation

Similar Papers 제목 키워드 기반

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

2025-05-28 · Zhe Kong, Feng Gao, Yong Zhang, Zhuoliang Kang 외

Audio-driven human animation methods, such as talking head and talking body generation, have made remarkable progress in generating synchronized facial movements and appealing visual quality videos. However, existing met…

Human AnimationInstruction FollowingVideo Generation

Dual Audio-Centric Modality Coupling for Talking Head Generation

2025-03-26 · Ao Fu, Ziqi Ni, Yi Zhou

The generation of audio-driven talking head videos is a key challenge in computer vision and graphics, with applications in virtual avatars and digital media. Traditional approaches often struggle with capturing the comp…

NeRFTalking Head Generationtext-to-speechText to Speech

One-Shot Pose-Driving Face Animation Platform

2024-07-12 · He Feng, Donglin Di, Yongjia Ma, Wei Chen 외

The objective of face animation is to generate dynamic and expressive talking head videos from a single reference face, utilizing driving conditions derived from either video or audio inputs. Current approaches often req…

Learning Frame-Wise Emotion Intensity for Audio-Driven Talking-Head Generation

2024-09-29 · Jingyi Xu, Hieu Le, Zhixin Shu, Yang Wang 외

Human emotional expression is inherently dynamic, complex, and fluid, characterized by smooth transitions in intensity throughout verbal communication. However, the modeling of such intensity fluctuations has been largel…

Talking Head Generation

What comprises a good talking-head video generation?: A Survey and Benchmark

2020-05-07 · Lele Chen, Guofeng Cui, Ziyi Kou, Haitian Zheng 외

Over the years, performance evaluation has become essential in computer vision, enabling tangible progress in many sub-fields. While talking-head video generation has become an emerging research topic, existing evaluatio…

Talking Head GenerationVideo Generation