paper-with-me

Papers

Audio-Visual Speech Enhancement In Complex Scenarios With Separation And Dereverberation Joint Modeling

2025-10-29 · Jiarong Du, Zhan Jin, Peijun Yang, Juan Liu, Zhuo Li, Xin Liu, Ming Li arxiv

Audio-visual speech enhancement (AVSE) is a task that uses visual auxiliary information to extract a target speaker's speech from mixed audio. In real-world scenarios, there often exist complex acoustic environments, accompanied by various interfering sounds and reverberation. Most previous methods struggle to cope with such complex conditions, resulting in poor perceptual quality of the extracted speech. In this paper, we propose an effective AVSE system that performs well in complex acoustic environments. Specifically, we design a "separation before dereverberation" pipeline that can be extended to other AVSE networks. The 4th COGMHEAR Audio-Visual Speech Enhancement Challenge (AVSEC) aims to explore new approaches to speech processing in multimodal complex environments. We validated the performance of our system in AVSEC-4: we achieved excellent results in the three objective metrics on the competition leaderboard, and ultimately secured first place in the human subjective listening test.

📄 PDF Abstract BibTeX arXiv:2510.26825

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

LA-VocE: Low-SNR Audio-visual Speech Enhancement using Neural Vocoders

2022-11-20 · Rodrigo Mira, Buye Xu, Jacob Donley, Anurag Kumar 외

Audio-visual speech enhancement aims to extract clean speech from a noisy environment by leveraging not only the audio itself but also the target speaker's lip movements. This approach has been shown to yield improvement…

Speech EnhancementSpeech Synthesis

Audio-visual Speech Enhancement Using Conditional Variational Auto-Encoders

2019-08-07 · Mostafa Sadeghi, Simon Leglaive, Xavier Alameda-Pineda, Laurent Girin 외

Variational auto-encoders (VAEs) are deep generative latent variable models that can be used for learning the distribution of complex data. VAEs have been successfully used to learn a probabilistic prior over speech sign…

Speech Enhancement

Improved Lite Audio-Visual Speech Enhancement

2020-08-30 · Shang-Yi Chuang, Hsin-Min Wang, Yu Tsao

Numerous studies have investigated the effectiveness of audio-visual multimodal learning for speech enhancement (AVSE) tasks, seeking a solution that uses visual data as auxiliary and complementary input to reduce the no…

Speech Enhancement

Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations

2025-07-29 · T. Aleksandra Ma, Sile Yin, Li-Chia Yang, Shuo Zhang arxiv

Speech enhancement in audio-only settings remains challenging, particularly in the presence of interfering speakers. This paper presents a simple yet effective real-time audio-visual speech enhancement (AVSE) system, RAV…

Audio-Visual Speech RecognitionActive Speaker DetectionSpeech Enhancement

Contextual Audio-Visual Switching For Speech Enhancement in Real-World Environments

2018-08-28 · Ahsan Adeel, Mandar Gogate, Amir Hussain

Human speech processing is inherently multimodal, where visual cues (lip movements) help to better understand the speech in noise. Lip-reading driven speech enhancement significantly outperforms benchmark audio-only appr…

Lip ReadingSpeech Enhancement