paper-with-me

홈 › Papers

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation

2024-12-21 · Lucas Goncalves, Prashant Mathur, Xing Niu, Brady Houston, Chandrashekhar Lavania, Srikanth Vishnubhotla, Lijia Sun, Anthony Ferritto

Audio-Visual Speech-to-Speech Translation typically prioritizes improving translation quality and naturalness. However, an equally critical aspect in audio-visual content is lip-synchrony-ensuring that the movements of the lips match the spoken content-essential for maintaining realism in dubbed videos. Despite its importance, the inclusion of lip-synchrony constraints in AVS2S models has been largely overlooked. This study addresses this gap by integrating a lip-synchrony loss into the training process of AVS2S models. Our proposed method significantly enhances lip-synchrony in direct audio-visual speech-to-speech translation, achieving an average LSE-D score of 10.67, representing a 9.2% reduction in LSE-D over a strong baseline across four language pairs. Additionally, it maintains the naturalness and high quality of the translated speech when overlaid onto the original video, without any degradation in translation quality.

📄 PDF Abstract BibTeX arXiv:2412.16530

Code (0)

등록된 구현이 없습니다.

Tasks

Speech-to-Speech TranslationTranslation

Similar Papers 제목 키워드 기반

On the Audio-visual Synchronization for Lip-to-Speech Synthesis

2023-03-01 · ICCV 2023 1 · Zhe Niu, Brian Mak

Most lip-to-speech (LTS) synthesis models are trained and evaluated under the assumption that the audio-video pairs in the dataset are perfectly synchronized. In this work, we show that the commonly used audio-visual dat…

Audio-Visual SynchronizationLip to Speech SynthesisSpeech Synthesis

A Comprehensive Multi-scale Approach for Speech and Dynamics Synchrony in Talking Head Generation

2023-07-04 · Louis Airale, Dominique Vaufreydaz, Xavier Alameda-Pineda

Animating still face images with deep generative models using a speech input signal is an active research topic and has seen important recent progress.However, much of the effort has been put into lip syncing and renderi…

Talking Head Generation

Lip Graph Assisted Audio-Visual Speech Recognition Using Bidirectional Synchronous Fusion

2020-10-25 · Interspeech 2020 10 · Hong Liu, Zhan Chen, Bing Yang

Current studies have shown that extracting representative visual features and efficiently fusing audio and visual modalities are vital for audio-visual speech recognition (AVSR), but these are still challenging. To this …

Audio-Visual Speech RecognitionLandmark-based Lipreadingspeech-recognitionSpeech Recognition+1

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization

2025-05-06 · Detao Bai, Zhiheng Ma, Xihan Wei, Liefeng Bo

The inherent synchronization between a speaker's lip movements, voice, and the underlying linguistic content offers a rich source of information for improving speech processing tasks, especially in challenging conditions…

Active Speaker DetectionAudio-Visual Speech RecognitionAudio-Visual SynchronizationRepresentation Learning+4

An Alternative to Low-level-Sychrony-Based Methods for Speech Detection

2010-12-01 · NeurIPS 2010 12 · Javier R. Movellan, Paul L. Ruvolo

Determining whether someone is talking has applications in many areas such as speech recognition, speaker diarization, social robotics, facial expression recognition, and human computer interaction. One popular approach…

Facial Expression RecognitionFacial Expression Recognition (FER)speaker-diarizationSpeaker Diarization+2