paper-with-me

Papers

Scenario-Aware Audio-Visual TF-GridNet for Target Speech Extraction

2023-10-30 · Zexu Pan, Gordon Wichern, Yoshiki Masuyama, Francois G. Germain, Sameer Khurana, Chiori Hori, Jonathan Le Roux

Target speech extraction aims to extract, based on a given conditioning cue, a target speech signal that is corrupted by interfering sources, such as noise or competing speakers. Building upon the achievements of the state-of-the-art (SOTA) time-frequency speaker separation model TF-GridNet, we propose AV-GridNet, a visual-grounded variant that incorporates the face recording of a target speaker as a conditioning factor during the extraction process. Recognizing the inherent dissimilarities between speech and noise signals as interfering sources, we also propose SAV-GridNet, a scenario-aware model that identifies the type of interfering scenario first and then applies a dedicated expert model trained specifically for that scenario. Our proposed model achieves SOTA results on the second COG-MHEAR Audio-Visual Speech Enhancement Challenge, outperforming other models by a significant margin, objectively and in a listening test. We also perform an extensive analysis of the results under the two scenarios.

📄 PDF Abstract BibTeX arXiv:2310.19644

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker SeparationSpeech EnhancementSpeech Extraction

Similar Papers 제목 키워드 기반

Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction

2025-05-27 · Zexu Pan, Shengkui Zhao, Tingting Wang, Kun Zhou 외

Audio-visual speaker extraction isolates a target speaker's speech from a mixture speech signal conditioned on a visual cue, typically using the target speaker's face recording. However, in real-world scenarios, other co…

InterGridNet: An Electric Network Frequency Approach for Audio Source Location Classification Using Convolutional Neural Networks

2025-02-14 · Christos Korgialas, Ioannis Tsingalis, Georgios Tzolopoulos, Constantine Kotropoulos

A novel framework, called InterGridNet, is introduced, leveraging a shallow RawNet model for geolocation classification of Electric Network Frequency (ENF) signatures in the SP Cup 2016 dataset. During data preparation, …

Decision MakingNeural Architecture Search

Target-Aware Spatio-Temporal Reasoning via Answering Questions in Dynamics Audio-Visual Scenarios

2023-05-21 · Yuanyuan Jiang, Jianqin Yin

Audio-visual question answering (AVQA) is a challenging task that requires multistep spatio-temporal reasoning over multimodal contexts. Recent works rely on elaborate target-agnostic parsing of audio-visual scenes for s…

Audio-visual Question AnsweringAudio-Visual Question Answering (AVQA)Question AnsweringScene Understanding+1

Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription

2023-09-15 · Peter Vieting, Simon Berger, Thilo von Neumann, Christoph Boeddeker 외

Many real-life applications of automatic speech recognition (ASR) require processing of overlapped speech. A common method involves first separating the speech into overlap-free streams on which ASR is performed. Recentl…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

SwGridNet: A Deep Convolutional Neural Network based on Grid Topology for Image Classification

2017-09-22 · Atsushi Takeda

Deep convolutional neural networks (CNNs) achieve remarkable performance on image classification tasks. Recent studies, however, have demonstrated that generalization abilities are more important than the depth of neural…

ClassificationEnsemble LearningGeneral Classificationimage-classification+1