paper-with-me

홈 › Papers

RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling

2025-05-28 · Long-Khanh Pham, Thanh V. T. Tran, Minh-Tan Pham, Van Nguyen

Lip-to-speech (L2S) synthesis, which reconstructs speech from visual cues, faces challenges in accuracy and naturalness due to limited supervision in capturing linguistic content, accents, and prosody. In this paper, we propose RESOUND, a novel L2S system that generates intelligible and expressive speech from silent talking face videos. Leveraging source-filter theory, our method involves two components: an acoustic path to predict prosody and a semantic path to extract linguistic features. This separation simplifies learning, allowing independent optimization of each representation. Additionally, we enhance performance by integrating speech units, a proven unsupervised speech representation technique, into waveform generation alongside mel-spectrograms. This allows RESOUND to synthesize prosodic speech while preserving content and speaker identity. Experiments conducted on two standard L2S benchmarks confirm the effectiveness of the proposed method across various metrics.

📄 PDF Abstract BibTeX arXiv:2505.22024

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Vocoder-Based Speech Synthesis from Silent Videos

2020-04-06 · Daniel Michelsanti, Olga Slizovskaia, Gloria Haro, Emilia Gómez 외

Both acoustic and visual information influence human perception of speech. For this reason, the lack of audio in a video sequence determines an extremely low speech intelligibility for untrained lip readers. In this pape…

Multi-Task LearningSpeech Synthesis

Speech Reconstruction from Silent Tongue and Lip Articulation By Pseudo Target Generation and Domain Adversarial Training

2023-04-12 · Rui-Chen Zheng, Yang Ai, Zhen-Hua Ling

This paper studies the task of speech reconstruction from ultrasound tongue images and optical lip videos recorded in a silent speaking mode, where people only activate their intra-oral and extra-oral articulators withou…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech

2025-03-21 · CVPR 2025 1 · Ji-Hoon Kim, Jeongsoo Choi, Jaehun Kim, Chaeyoung Jung 외

The objective of this study is to generate high-quality speech from silent talking face videos, a task also known as video-to-speech synthesis. A significant challenge in video-to-speech synthesis lies in the substantial…

Speech Synthesis

Vid2speech: Speech Reconstruction from Silent Video

2017-01-02 · Ariel Ephrat, Shmuel Peleg

Speechreading is a notoriously difficult task for humans to perform. In this paper we present an end-to-end model based on a convolutional neural network (CNN) for generating an intelligible acoustic speech signal from s…

SVoice: Enabling Voice Communication in Silence via Acoustic Sensing on Commodity Devices

2022-11-09 · SenSys 2022 11 · Yongjian Fu, Shuning Wang, Linghui Zhong, Lili Chen 외

Silent Speech Interface (SSI) has been proposed as a means of reconstructing audible speech from silent articulatory gestures for covert voice communication in public and voice assistance for the aphasic. Prior arts o…