paper-with-me

홈 › Papers

From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech

2025-03-21 · CVPR 2025 1 · Ji-Hoon Kim, Jeongsoo Choi, Jaehun Kim, Chaeyoung Jung, Joon Son Chung

The objective of this study is to generate high-quality speech from silent talking face videos, a task also known as video-to-speech synthesis. A significant challenge in video-to-speech synthesis lies in the substantial modality gap between silent video and multi-faceted speech. In this paper, we propose a novel video-to-speech system that effectively bridges this modality gap, significantly enhancing the quality of synthesized speech. This is achieved by learning of hierarchical representations from video to speech. Specifically, we gradually transform silent video into acoustic feature spaces through three sequential stages -- content, timbre, and prosody modeling. In each stage, we align visual factors -- lip movements, face identity, and facial expressions -- with corresponding acoustic counterparts to ensure the seamless transformation. Additionally, to generate realistic and coherent speech from the visual representations, we employ a flow matching model that estimates direct trajectories from a simple prior distribution to the target speech distribution. Extensive experiments demonstrate that our method achieves exceptional generation quality comparable to real utterances, outperforming existing methods by a significant margin.

📄 PDF Abstract BibTeX arXiv:2503.16956

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Who is Authentic Speaker

2024-04-30 · Qiang Huang

Voice conversion (VC) using deep learning technologies can now generate high quality one-to-many voices and thus has been used in some practical application fields, such as entertainment and healthcare. However, voice co…

Speaker RecognitionVoice Conversion

Disjoint Mapping Network for Cross-modal Matching of Voices and Faces

2018-07-12 · ICLR 2019 5 · Yandong Wen, Mahmoud Al Ismail, Weiyang Liu, Bhiksha Raj 외

We propose a novel framework, called Disjoint Mapping Network (DIMNet), for cross-modal biometric matching, in particular of voices and faces. Different from the existing methods, DIMNet does not explicitly learn the joi…

On Learning Associations of Faces and Voices

2018-05-15 · Changil Kim, Hijung Valentina Shin, Tae-Hyun Oh, Alexandre Kaspar 외

In this paper, we study the associations between human faces and voices. Audiovisual integration, specifically the integration of facial and vocal information is a well-researched area in neuroscience. It is shown that t…

Speaker Identification

Speaker Recognition in Realistic Scenario Using Multimodal Data

2023-02-25 · Saqlain Hussain Shah, Muhammad Saad Saeed, Shah Nawaz, Muhammad Haroon Yousaf

In recent years, an association is established between faces and voices of celebrities leveraging large scale audio-visual information from YouTube. The availability of large scale audio-visual datasets is instrumental i…

Speaker Recognition

Audio-Visual Speaker Verification via Joint Cross-Attention

2023-09-28 · R. Gnana Praveen, Jahangir Alam

Speaker verification has been widely explored using speech signals, which has shown significant improvement using deep models. Recently, there has been a surge in exploring faces and voices as they can offer more complem…

Speaker Verification