paper-with-me

홈 › Papers

FaSNet: Low-latency Adaptive Beamforming for Multi-microphone Audio Processing

2019-09-29 · Yi Luo, Enea Ceolini, Cong Han, Shih-Chii Liu, Nima Mesgarani

Beamforming has been extensively investigated for multi-channel audio processing tasks. Recently, learning-based beamforming methods, sometimes called \textit{neural beamformers}, have achieved significant improvements in both signal quality (e.g. signal-to-noise ratio (SNR)) and speech recognition (e.g. word error rate (WER)). Such systems are generally non-causal and require a large context for robust estimation of inter-channel features, which is impractical in applications requiring low-latency responses. In this paper, we propose filter-and-sum network (FaSNet), a time-domain, filter-based beamforming approach suitable for low-latency scenarios. FaSNet has a two-stage system design that first learns frame-level time-domain adaptive beamforming filters for a selected reference channel, and then calculate the filters for all remaining channels. The filtered outputs at all channels are summed to generate the final output. Experiments show that despite its small model size, FaSNet is able to outperform several traditional oracle beamformers with respect to scale-invariant signal-to-noise ratio (SI-SNR) in reverberant speech enhancement and separation tasks. Moreover, when trained with a frequency-domain objective function on the CHiME-3 dataset, FaSNet achieves 14.3\% relative word error rate reduction (RWERR) compared with the baseline model. These results show the efficacy of FaSNet particularly in reverberant and noisy signal conditions.

📄 PDF Abstract BibTeX arXiv:1909.13387

Code (1)

yluo42/TAC pytorch

Tasks

Speech Enhancementspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

End-to-end Microphone Permutation and Number Invariant Multi-channel Speech Separation

2019-10-30 · Yi Luo, Zhuo Chen, Nima Mesgarani, Takuya Yoshioka

An important problem in ad-hoc microphone speech separation is how to guarantee the robustness of a system with respect to the locations and numbers of microphones. The former requires the system to be invariant to diffe…

Speech Separation

SRIB-LEAP submission to Far-field Multi-Channel Speech Enhancement Challenge for Video Conferencing

2021-06-24 · R G Prithvi Raj, Rohit Kumar, M K Jayesh, Anurenjan Purushothaman 외

This paper presents the details of the SRIB-LEAP submission to the ConferencingSpeech challenge 2021. The challenge involved the task of multi-channel speech enhancement to improve the quality of far field speech from mi…

Speech Enhancement

Deep Long Short-Term Memory Adaptive Beamforming Networks For Multichannel Robust Speech Recognition

2017-11-21 · Zhong Meng, Shinji Watanabe, John R. Hershey, Hakan Erdogan

Far-field speech recognition in noisy and reverberant conditions remains a challenging problem despite recent deep learning breakthroughs. This problem is commonly addressed by acquiring a speech signal from multiple mic…

Robust Speech Recognitionspeech-recognitionSpeech Recognition

Modified Parametric Multichannel Wiener Filter \\for Low-latency Enhancement of Speech Mixtures with Unknown Number of Speakers

2023-06-29 · Ning Guo, Tomohiro Nakatani, Shoko Araki, Takehiro Moriya

This paper introduces a novel low-latency online beamforming (BF) algorithm, named Modified Parametric Multichannel Wiener Filter (Mod-PMWF), for enhancing speech mixtures with unknown and varying number of speakers. Alt…

Low-latency processing

Optimal model-based beamforming and independent steering for spherical loudspeaker arrays

2023-10-06 · Boaz Rafaely, Dima Khaykin

Spherical loudspeaker arrays have been recently studied for directional sound radiation, where the compact arrangement of the loudspeaker units around a sphere facilitated the control of sound radiation in three-dimensio…