paper-with-me

Papers

TDFNet: An Efficient Audio-Visual Speech Separation Model with Top-down Fusion

2024-01-25 · Samuel Pegg, Kai Li, Xiaolin Hu

Audio-visual speech separation has gained significant traction in recent years due to its potential applications in various fields such as speech recognition, diarization, scene analysis and assistive technologies. Designing a lightweight audio-visual speech separation network is important for low-latency applications, but existing methods often require higher computational costs and more parameters to achieve better separation performance. In this paper, we present an audio-visual speech separation model called Top-Down-Fusion Net (TDFNet), a state-of-the-art (SOTA) model for audio-visual speech separation, which builds upon the architecture of TDANet, an audio-only speech separation method. TDANet serves as the architectural foundation for the auditory and visual networks within TDFNet, offering an efficient model with fewer parameters. On the LRS2-2Mix dataset, TDFNet achieves a performance increase of up to 10\% across all performance metrics compared with the previous SOTA method CTCNet. Remarkably, these results are achieved using fewer parameters and only 28\% of the multiply-accumulate operations (MACs) of CTCNet. In essence, our method presents a highly effective and efficient solution to the challenges of speech separation within the audio-visual domain, making significant strides in harnessing visual information optimally.

📄 PDF Abstract BibTeX arXiv:2401.14185

Code (1)

spkgyk/TDFNet 공식 구현 pytorch

Tasks

speech-recognitionSpeech RecognitionSpeech Separation

Similar Papers 제목 키워드 기반

RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation

2023-09-29 · Samuel Pegg, Kai Li, Xiaolin Hu

Audio-visual speech separation methods aim to integrate different modalities to generate high-quality separated speech, thereby enhancing the performance of downstream tasks such as speech recognition. Most existing stat…

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionSpeech Separation+1

Audio-Visual Speech Separation Using Cross-Modal Correspondence Loss

2021-03-02 · Naoki Makishima, Mana Ihori, Akihiko Takashima, Tomohiro Tanaka 외

We present an audio-visual speech separation learning method that considers the correspondence between the separated signals and the visual signals to reflect the speech characteristics during training. Audio-visual spee…

Speech Separation

CSLNSpeech: solving extended speech separation problem with the help of Chinese sign language

2020-07-21 · Jiasong Wu, Xuan Li, Taotao Li, Fanman Meng 외

Previous audio-visual speech separation methods use the synchronization of the speaker's facial movement and speech in the video to supervise the speech separation in a self-supervised way. In this paper, we propose a mo…

Self-Supervised LearningSpeech Separation

Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation

2018-04-10 · Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel 외

We present a joint audio-visual model for isolating a single speech signal from a mixture of sounds such as other speakers and background noise. Solving this task using only audio as input is extremely challenging and do…

Speech Separation

An Audio-Visual Speech Separation Model Inspired by Cortico-Thalamo-Cortical Circuits

2022-12-21 · Kai Li, Fenghua Xie, Hang Chen, Kexin Yuan 외

Audio-visual approaches involving visual inputs have laid the foundation for recent progress in speech separation. However, the optimization of the concurrent usage of auditory and visual inputs is still an active resear…

Speech Separation