paper-with-me

홈 › Papers

A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple Languages

2025-10-08 · Zibo Su, Kun Wei, Jiahua Li, Xu Yang, Cheng Deng arxiv

Speech-driven talking face synthesis (TFS) focuses on generating lifelike facial animations from speech input. Current TFS models perform well in English but struggle with non-English languages, producing inaccurate mouth shapes and rigid facial expressions. These limitations are mainly caused by English-dominated training datasets and the lack of cross-language generalization ability.To address these challenges, we propose Multilingual Experts (MuEx), a novel framework featuring a Phoneme-Guided Mixture-of-Experts (PG-MoE) architecture that employs phonemes and visemes as universal intermediaries to bridge the gap between audio and visual modalities, enabling lifelike multilingual TFS. We extract speech and visual features as phonemes and visemes, respectively, which represent the basic units of speech sounds and mouth movements, to alleviate linguistic differences and dataset bias.Furthermore, we introduce the Phoneme-Viseme Alignment Mechanism (PV-Align), which establishes robust cross-modal correspondences between phonemes and visemes to improve audiovisual synchronization. In addition, we construct a Multilingual Talking Face Dataset (MTFD) comprising 12 diverse languages with 95.04 hours of high-quality videos for training and evaluating multilingual TFS performance.Extensive experiments demonstrate that MuEx achieves superior performance across all languages in MTFD and exhibits effective zero-shot generalization to unseen languages without additional training.

📄 PDF Abstract BibTeX arXiv:2510.06612

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot Generalization

Similar Papers 제목 키워드 기반

NPVForensics: Jointing Non-critical Phonemes and Visemes for Deepfake Detection

2023-06-12 · Yu Chen, Yang Yu, Rongrong Ni, Yao Zhao 외

Deepfake technologies empowered by deep learning are rapidly evolving, creating new security concerns for society. Existing multimodal detection methods usually capture audio-visual inconsistencies to expose Deepfake vid…

DeepFake DetectionFace Swapping

Learning Audio-Driven Viseme Dynamics for 3D Face Animation

2023-01-15 · Linchao Bao, Haoxian Zhang, Yue Qian, Tangli Xue 외

We present a novel audio-driven facial animation approach that can generate realistic lip-synchronized 3D facial animations from the input audio. Our approach learns viseme dynamics from speech videos, produces animator-…

3D Face Animation

Automatic Viseme Vocabulary Construction to Enhance Continuous Lip-reading

2017-04-26 · Adriana Fernandez-Lopez, Federico M. Sukno

Speech is the most common communication method between humans and involves the perception of both auditory and visual channels. Automatic speech recognition focuses on interpreting the audio signals, but it has been demo…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)ClusteringLip Reading+2

Phoneme-to-viseme mappings: the good, the bad, and the ugly

2018-05-08 · Helen L. Bear, Richard Harvey

Visemes are the visual equivalent of phonemes. Although not precisely defined, a working definition of a viseme is "a set of phonemes which have identical appearance on the lips". Therefore a phoneme falls into one visem…

Estimating speech from lip dynamics

2017-08-03 · Jithin Donny George, Ronan Keane, Conor Zellmer

The goal of this project is to develop a limited lip reading algorithm for a subset of the English language. We consider a scenario in which no audio information is available. The raw video is processed and the position …

Lip ReadingPositionSentence