paper-with-me

Papers

Audio-Visual Speech Separation in Noisy Environments with a Lightweight Iterative Model

2023-05-31 · Héctor Martel, Julius Richter, Kai Li, Xiaolin Hu, Timo Gerkmann

We propose Audio-Visual Lightweight ITerative model (AVLIT), an effective and lightweight neural network that uses Progressive Learning (PL) to perform audio-visual speech separation in noisy environments. To this end, we adopt the Asynchronous Fully Recurrent Convolutional Neural Network (A-FRCNN), which has shown successful results in audio-only speech separation. Our architecture consists of an audio branch and a video branch, with iterative A-FRCNN blocks sharing weights for each modality. We evaluated our model in a controlled environment using the NTCD-TIMIT dataset and in-the-wild using a synthetic dataset that combines LRS3 and WHAM!. The experiments demonstrate the superiority of our model in both settings with respect to various audio-only and audio-visual baselines. Furthermore, the reduced footprint of our model makes it suitable for low resource applications.

📄 PDF Abstract BibTeX arXiv:2306.00160

Code (1)

hmartelb/avlit 공식 구현 pytorch

Tasks

Speech Separation

Similar Papers 제목 키워드 기반

Diffusion-Based Unsupervised Audio-Visual Speech Separation in Noisy Environments with Noise Prior

2025-09-17 · Yochai Yemini, Rami Ben-Ari, Sharon Gannot, Ethan Fetaya arxiv

In this paper, we address the problem of single-microphone speech separation in the presence of ambient noise. We propose a generative unsupervised technique that directly models both clean speech and structured noise co…

Speech Separation

A Multi-Stage Triple-Path Method for Speech Separation in Noisy and Reverberant Environments

2023-03-07 · Zhaoxi Mu, Xinyu Yang, Xiangyuan Yang, Wenjing Zhu

In noisy and reverberant environments, the performance of deep learning-based speech separation methods drops dramatically because previous methods are not designed and optimized for such situations. To address this issu…

DenoisingSpeech DenoisingSpeech Separation

Audio-visual speech separation based on joint feature representation with cross-modal attention

2022-03-05 · Junwen Xiong, Peng Zhang, Lei Xie, Wei Huang 외

Multi-modal based speech separation has exhibited a specific advantage on isolating the target character in multi-talker noisy environments. Unfortunately, most of current separation strategies prefer a straightforward f…

Optical Flow EstimationSpeech Separation

Robust Active Speaker Detection in Noisy Environments

2024-03-27 · Siva Sai Nagender Vasireddy, Chenxu Zhang, Xiaohu Guo, Yapeng Tian

This paper addresses the issue of active speaker detection (ASD) in noisy environments and formulates a robust active speaker detection (rASD) problem. Existing ASD approaches leverage both audio and visual modalities, b…

Active Speaker DetectionSpeech Separation

Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention

2025-09-28 · Kai Li, Kejun Gao, Xiaolin Hu arxiv

Audio-visual speech separation (AVSS) methods leverage visual cues to extract target speech and have demonstrated strong separation quality in noisy acoustic environments. However, these methods usually involve a large n…

Speech Separation