A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages
Speech-driven talking face synthesis (TFS) focuses on generating lifelike facial animations from speech input. Current TFS models perform well in English but struggle with non-English languages, producing inaccurate mouth shapes and rigid facial expressions. These limitations are mainly caused by English-dominated training datasets and the lack of cross-language generalization ability.To address these challenges, we propose Multilingual Experts (MuEx), a novel framework featuring a Phoneme-Guided Mixture-of-Experts (PG-MoE) architecture that employs phonemes and visemes as universal intermediaries to bridge the gap between audio and visual modalities, enabling lifelike multilingual TFS. We extract speech and visual features as phonemes and visemes, respectively, which represent the basic units of speech sounds and mouth movements, to alleviate linguistic differences and dataset bias.Furthermore, we introduce the Phoneme-Viseme Alignment Mechanism (PV-Align), which establishes robust cross-modal correspondences between phonemes and visemes to improve audiovisual synchronization. In addition, we construct a Multilingual Talking Face Dataset (MTFD) comprising 12 diverse languages with 95.04 hours of high-quality videos for training and evaluating multilingual TFS performance.Extensive experiments demonstrate that MuEx achieves superior performance across all languages in MTFD and exhibits effective zero-shot generalization to unseen languages without additional training.
Code (0)
등록된 구현이 없습니다.
Tasks
Zero-shot GeneralizationSimilar Papers 제목 키워드 기반
NPVForensics: Jointing Non-critical Phonemes and Visemes for Deepfake Detection
Deepfake technologies empowered by deep learning are rapidly evolving, creating new security concerns for society. Existing multimodal detection methods usually capture audio-visual inconsistencies to expose Deepfake vid…
DeepFake DetectionFace SwappingLearning Audio-Driven Viseme Dynamics for 3D Face Animation
We present a novel audio-driven facial animation approach that can generate realistic lip-synchronized 3D facial animations from the input audio. Our approach learns viseme dynamics from speech videos, produces animator-…
3D Face AnimationAutomatic Viseme Vocabulary Construction to Enhance Continuous Lip-reading
Speech is the most common communication method between humans and involves the perception of both auditory and visual channels. Automatic speech recognition focuses on interpreting the audio signals, but it has been demo…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClusteringLip Reading+2Phoneme-to-viseme mappings: the good, the bad, and the ugly
Visemes are the visual equivalent of phonemes. Although not precisely defined, a working definition of a viseme is "a set of phonemes which have identical appearance on the lips". Therefore a phoneme falls into one visem…
Estimating speech from lip dynamics
The goal of this project is to develop a limited lip reading algorithm for a subset of the English language. We consider a scenario in which no audio information is available. The raw video is processed and the position …
Lip ReadingPositionSentence