paper-with-me

홈 › Papers

MuteSwap: Visual-informed Silent Video Identity Conversion

2025-07-01 · Yifan Liu, Yu Fang, Zhouhan Lin arxiv

Conventional voice conversion modifies voice characteristics from a source speaker to a target speaker, relying on audio input from both sides. However, this process becomes infeasible when clean audio is unavailable, such as in silent videos or noisy environments. In this work, we focus on the task of Silent Face-based Voice Conversion (SFVC), which does voice conversion entirely from visual inputs. i.e., given images of a target speaker and a silent video of a source speaker containing lip motion, SFVC generates speech aligning the identity of the target speaker while preserving the speech content in the source silent video. As this task requires generating intelligible speech and converting identity using only visual cues, it is particularly challenging. To address this, we introduce MuteSwap, a novel framework that employs contrastive learning to align cross-modality identities and minimize mutual information to separate shared visual features. Experimental results show that MuteSwap achieves impressive performance in both speech synthesis and identity conversion, especially under noisy conditions where methods dependent on audio input fail to produce intelligible results, demonstrating both the effectiveness of our training approach and the feasibility of SFVC.

📄 PDF Abstract BibTeX arXiv:2507.00498

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningVoice ConversionSpeech Synthesis

Similar Papers 제목 키워드 기반

VisageSynTalk: Unseen Speaker Video-to-Speech Synthesis via Speech-Visage Feature Selection

2022-06-15 · Joanna Hong, Minsu Kim, Yong Man Ro

The goal of this work is to reconstruct speech from a silent talking face video. Recent studies have shown impressive performance on synthesizing speech from silent talking face videos. However, they have not explicitly …

feature selectionSpeech Synthesis

From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech

2025-03-21 · CVPR 2025 1 · Ji-Hoon Kim, Jeongsoo Choi, Jaehun Kim, Chaeyoung Jung 외

The objective of this study is to generate high-quality speech from silent talking face videos, a task also known as video-to-speech synthesis. A significant challenge in video-to-speech synthesis lies in the substantial…

Speech Synthesis

Audio-driven Talking Face Generation with Stabilized Synchronization Loss

2023-07-18 · Dogucan Yaman, Fevziye Irem Eyiokur, Leonard Bärmann, Hazim Kemal Ekenel 외

Talking face generation aims to create realistic videos with accurate lip synchronization and high visual quality, using given audio and reference video while preserving identity and visual characteristics. In this paper…

Audio-Visual SynchronizationFace GenerationTalking Face Generation

Speaker disentanglement in video-to-speech conversion

2021-05-20 · Dan Oneata, Adriana Stan, Horia Cucu

The task of video-to-speech aims to translate silent video of lip movement to its corresponding audio signal. Previous approaches to this task are generally limited to the case of a single speaker, but a method that acco…

DisentanglementSpeech Synthesis

Assessing Identity Leakage in Talking Face Generation: Metrics and Evaluation Framework

2025-11-05 · Dogucan Yaman, Fevziye Irem Eyiokur, Hazım Kemal Ekenel, Alexander Waibel arxiv

Video editing-based talking face generation aims to preserve video details such as pose, lighting, and gestures while modifying only lip motion, often using an identity reference image to maintain speaker consistency. Ho…

Talking Face Generation