paper-with-me

홈 › Papers

AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation

2023-12-05 · CVPR 2024 1 · Jeongsoo Choi, Se Jin Park, Minsu Kim, Yong Man Ro

This paper proposes a novel direct Audio-Visual Speech to Audio-Visual Speech Translation (AV2AV) framework, where the input and output of the system are multimodal (i.e., audio and visual speech). With the proposed AV2AV, two key advantages can be brought: 1) We can perform real-like conversations with individuals worldwide in a virtual meeting by utilizing our own primary languages. In contrast to Speech-to-Speech Translation (A2A), which solely translates between audio modalities, the proposed AV2AV directly translates between audio-visual speech. This capability enhances the dialogue experience by presenting synchronized lip movements along with the translated speech. 2) We can improve the robustness of the spoken language translation system. By employing the complementary information of audio-visual speech, the system can effectively translate spoken language even in the presence of acoustic noise, showcasing robust performance. To mitigate the problem of the absence of a parallel AV2AV translation dataset, we propose to train our spoken language translation system with the audio-only dataset of A2A. This is done by learning unified audio-visual speech representations through self-supervised learning in advance to train the translation system. Moreover, we propose an AV-Renderer that can generate raw audio and video in parallel. It is designed with zero-shot speaker modeling, thus the speaker in source audio-visual speech can be maintained at the target translated audio-visual speech. The effectiveness of AV2AV is evaluated with extensive experiments in a many-to-many language translation setting. Demo page is available on https://choijeongsoo.github.io/av2av.

📄 PDF Abstract BibTeX arXiv:2312.02512

Code (1)

choijeongsoo/av2av 공식 구현 pytorch

Tasks

Self-Supervised LearningSpeech-to-Speech TranslationTranslation

Similar Papers 제목 키워드 기반

MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation

2023-03-01 · Mohamed Anwar, Bowen Shi, Vedanuj Goswami, Wei-Ning Hsu 외

We introduce MuAViC, a multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation providing 1200 hours of audio-visual speech in 9 languages. It is fully transcribed and covers 6…

Audio-Visual Speech RecognitionRobust Speech Recognitionspeech-recognitionSpeech Recognition+4

AV-TranSpeech: Audio-Visual Robust Speech-to-Speech Translation

2023-05-24 · Rongjie Huang, Huadai Liu, Xize Cheng, Yi Ren 외

Direct speech-to-speech translation (S2ST) aims to convert speech from one language into another, and has demonstrated significant progress to date. Despite the recent success, current S2ST models still suffer from disti…

Speech-to-Speech TranslationTranslation

TransFace: Unit-Based Audio-Visual Speech Synthesizer for Talking Head Translation

2023-12-23 · Xize Cheng, Rongjie Huang, Linjun Li, Tao Jin 외

Direct speech-to-speech translation achieves high-quality results through the introduction of discrete units obtained from self-supervised learning. This approach circumvents delays and cascading errors associated with m…

es-enfr-enSelf-Supervised LearningSpeech-to-Speech Translation+1

AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition

2025-08-11 · Junxiao Xue, Xiaozhen Liu, Xuecheng Wu, Xinyi Yin 외 arxiv

Audio-visual speech recognition (AVSR) combines audio-visual modalities to improve speech recognition, especially in noisy environments. However, most existing methods deploy the unidirectional enhancement or symmetric f…

Audio-Visual Speech Recognition

End-to-end Audiovisual Speech Recognition

2018-02-18 · IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2018 9 · Stavros Petridis, Themos Stafylakis, Pingchuan Ma, Feipeng Cai 외

Several end-to-end deep learning approaches have been recently presented which extract either audio or visual features from the input images or audio signals and perform speech recognition. However, research on end-to-en…

Lipreadingspeech-recognitionSpeech Recognition