paper-with-me

Papers

Multi-Geometry Spatial Acoustic Modeling for Distant Speech Recognition

2019-04-28

The use of spatial information with multiple microphones can improve far-field automatic speech recognition (ASR) accuracy. However, conventional microphone array techniques degrade speech enhancement performance when there is an array geometry mismatch between design and test conditions. Moreover, such speech enhancement techniques do not always yield ASR accuracy improvement due to the difference between speech enhancement and ASR optimization objectives. In this work, we propose to unify an acoustic model framework by optimizing spatial filtering and long short-term memory (LSTM) layers from multi-channel (MC) input. Our acoustic model subsumes beamformers with multiple types of array geometry. In contrast to deep clustering methods that treat a neural network as a black box tool, the network encoding the spatial filters can process streaming audio data in real time without the accumulation of target signal statistics. We demonstrate the effectiveness of such MC neural networks through ASR experiments on the real-world far-field data. We show that our two-channel acoustic model can on average reduce word error rates (WERs) by~13.4 and~12.7% compared to a single channel ASR system with the log-mel filter bank energy (LFBE) feature under the matched and mismatched microphone placement conditions, respectively. Our result also shows that our two-channel network achieves a relative WER reduction of over~7.0% compared to conventional beamforming with seven microphones overall.

📄 PDF Abstract BibTeX arXiv:1903.06539

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Deep ClusteringDistant Speech RecognitionSpeech Enhancementspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

AV-Surf: Surface-Enhanced Geometry-Aware Novel-View Acoustic Synthesis

2025-03-17 · Hadam Baek, Hannie Shin, Jiyoung Seo, Chanwoo Kim 외

Accurately modeling sound propagation with complex real-world environments is essential for Novel View Acoustic Synthesis (NVAS). While previous studies have leveraged visual perception to estimate spatial acoustics, the…

3DGS

Frequency Domain Multi-channel Acoustic Modeling for Distant Speech Recognition

2019-04-28

Conventional far-field automatic speech recognition (ASR) systems typically employ microphone array techniques for speech enhancement in order to improve robustness against noise or reverberation. However, such speech en…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Distant Speech RecognitionSpeech Enhancement+2

NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant Meeting Transcription

2024-01-16 · Alon Vinnikov, Amir Ivry, Aviv Hurvitz, Igor Abramovski 외

We introduce the first Natural Office Talkers in Settings of Far-field Audio Recordings (``NOTSOFAR-1'') Challenge alongside datasets and baseline system. The challenge focuses on distant speaker diarization and automati…

Automatic Speech RecognitionBenchmarkingspeaker-diarizationSpeaker Diarization+2

Contaminated speech training methods for robust DNN-HMM distant speech recognition

2017-10-10 · Mirco Ravanelli, Maurizio Omologo

Despite the significant progress made in the last years, state-of-the-art speech recognition technologies provide a satisfactory performance only in the close-talking condition. Robustness of distant speech recognition i…

Distant Speech RecognitionSpeech Enhancementspeech-recognitionSpeech Recognition

ActiveRIR: Active Audio-Visual Exploration for Acoustic Environment Modeling

2024-04-24 · Arjun Somayazulu, Sagnik Majumder, Changan Chen, Kristen Grauman

An environment acoustic model represents how sound is transformed by the physical characteristics of an indoor environment, for any given source/receiver location. Traditional methods for constructing acoustic models inv…

Reinforcement Learning (RL)