paper-with-me

Papers

Array Geometry-Robust Attention-Based Neural Beamformer for Moving Speakers

2024-02-05 · Marvin Tammen, Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani, Shoko Araki, Simon Doclo

Although mask-based beamforming is a powerful speech enhancement approach, it often requires manual parameter tuning to handle moving speakers. Recently, this approach was augmented with an attention-based spatial covariance matrix aggregator (ASA) module, enabling accurate tracking of moving speakers without manual tuning. However, the deep neural network model used in this module is limited to specific microphone arrays, necessitating a different model for varying channel permutations, numbers, or geometries. To improve the robustness of the ASA module against such variations, in this paper we investigate three approaches: training with random channel configurations, employing the transform-average-concatenate method to process multi-channel input features, and utilizing robust input features. Our experiments on the CHiME-3 and DEMAND datasets show that these approaches enable the ASA-augmented beamformer to track moving speakers across different microphone arrays unseen in training.

📄 PDF Abstract BibTeX arXiv:2402.03058

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

Weakly Guided and Autoregressive Beamformer Parameterization for Generalizable Moving Speaker Extraction in Higher-Order Ambisonics

2026-07-05 · Jakob Kienegger, Tal Peer, Sina Khanagha, Timo Gerkmann arxiv

Linear spatial filters (beamformers) enable robust, generalizable and interpretable speech enhancement with performance guarantees under ideal parameterization. Modern beamformers are often parameterized by deep neural n…

Speech Enhancement

Lightweight and Robust Multi-Channel End-to-End Speech Recognition with Spherical Harmonic Transform

2025-06-13 · Xiangzhu Kong, Huang Hao, Zhijian Ou

This paper presents SHTNet, a lightweight spherical harmonic transform (SHT) based framework, which is designed to address cross-array generalization challenges in multi-channel automatic speech recognition (ASR) through…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)channel selectionspeech-recognition+1

DFSNet: A Steerable Neural Beamformer Invariant to Microphone Array Configuration for Real-Time, Low-Latency Speech Enhancement

2023-02-26 · Anton Kovalyov, Kashyap Patel, Issa Panahi

Invariance to microphone array configuration is a rare attribute in neural beamformers. Filter-and-sum (FS) methods in this class define the target signal with respect to a reference channel. However, this not only compl…

AttributeSpeech Enhancement

Adaptive Dereverberation, Noise and Interferer Reduction Using Sparse Weighted Linearly Constrained Minimum Power Beamforming

2023-03-13 · Henri Gode, Simon Doclo

Interfering sources, background noise and reverberation degrade speech quality and intelligibility in hearing aid applications. In this paper, we present an adaptive algorithm aiming at dereverberation, noise and interfe…

Speech Enhancement

On Interference-Rejection Using Riemannian Geometry for Direction of Arrival Estimation

2023-01-09 · Amitay Bar, Ronen Talmon

We consider the problem of estimating the direction of arrival of desired acoustic sources in the presence of multiple acoustic interference sources. All the sources are located in noisy and reverberant environments and …

Direction of Arrival Estimation