paper-with-me

Papers

Attention-Driven Multichannel Speech Enhancement in Moving Sound Source Scenarios

2023-12-17 · Yuzhu Wang, Archontis Politis, Tuomas Virtanen

Current multichannel speech enhancement algorithms typically assume a stationary sound source, a common mismatch with reality that limits their performance in real-world scenarios. This paper focuses on attention-driven spatial filtering techniques designed for dynamic settings. Specifically, we study the application of linear and nonlinear attention-based methods for estimating time-varying spatial covariance matrices used to design the filters. We also investigate the direct estimation of spatial filters by attention-based methods without explicitly estimating spatial statistics. The clean speech clips from WSJ0 are employed for simulating speech signals of moving speakers in a reverberant environment. The experimental dataset is built by mixing the simulated speech signals with multichannel real noise from CHiME-3. Evaluation results show that the attention-driven approaches are robust and consistently outperform conventional spatial filtering approaches in both static and dynamic sound environments.

📄 PDF Abstract BibTeX arXiv:2312.10756

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

Channel-Attention Dense U-Net for Multichannel Speech Enhancement

2020-01-30 · Bahareh Tolooshams, Ritwik Giri, Andrew H. Song, Umut Isik 외

Supervised deep learning has gained significant attention for speech enhancement recently. The state-of-the-art deep learning methods perform the task by learning a ratio/binary mask that is applied to the mixture in the…

Speech Enhancement

Student-Teacher Learning for BLSTM Mask-based Speech Enhancement

2018-03-27

Spectral mask estimation using bidirectional long short-term memory (BLSTM) neural networks has been widely used in various speech enhancement applications, and it has achieved great success when it is applied to multich…

Speech Enhancementspeech-recognitionSpeech Recognition

Cleanformer: A multichannel array configuration-invariant neural enhancement frontend for ASR in smart speakers

2022-04-25 · Joseph Caroselli, Arun Narayanan, Nathan Howard, Tom O'Malley

This work introduces the Cleanformer, a streaming multichannel neural based enhancement frontend for automatic speech recognition (ASR). This model has a conformer-based architecture which takes as inputs a single channe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Semi-Supervised Multichannel Speech Enhancement With a Deep Speech Prior

2019-10-07 · IEEE/ACM Transactions on Audio, Speech, and Language Processing 2019 10 · Kouhei Sekiguchi, Yoshiaki Bando, Aditya Arie Nugraha, Kazuyoshi Yoshii 외

This paper describes a semi-supervised multichannel speech enhancement method that uses clean speech data for prior training. Although multichannel nonnegative matrix factorization (MNMF) and its constrained variant call…

Speech Enhancement

Real-time Streaming Wave-U-Net with Temporal Convolutions for Multichannel Speech Enhancement

2021-04-05 · Vasiliy Kuzmin, Fyodor Kravchenko, Artem Sokolov, Jie Geng

In this paper, we describe the work that we have done to participate in Task1 of the ConferencingSpeech2021 challenge. This task set a goal to develop the solution for multi-channel speech enhancement in a real-time mann…

DecoderSpeech Enhancement