paper-with-me

Papers

One model to enhance them all: array geometry agnostic multi-channel personalized speech enhancement

2021-10-20 · Hassan Taherian, Sefik Emre Eskimez, Takuya Yoshioka, Huaming Wang, Zhuo Chen, Xuedong Huang

With the recent surge of video conferencing tools usage, providing high-quality speech signals and accurate captions have become essential to conduct day-to-day business or connect with friends and families. Single-channel personalized speech enhancement (PSE) methods show promising results compared with the unconditional speech enhancement (SE) methods in these scenarios due to their ability to remove interfering speech in addition to the environmental noise. In this work, we leverage spatial information afforded by microphone arrays to improve such systems' performance further. We investigate the relative importance of speaker embeddings and spatial features. Moreover, we propose a new causal array-geometry-agnostic multi-channel PSE model, which can generate a high-quality enhanced signal from arbitrary microphone geometry. Experimental results show that the proposed geometry agnostic model outperforms the model trained on a specific microphone array geometry in both speech quality and automatic speech recognition accuracy. We also demonstrate the effectiveness of the proposed approach for unseen array geometries.

📄 PDF Abstract BibTeX arXiv:2110.10330

Code (0)

등록된 구현이 없습니다.

Tasks

AllAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Online Normalization Online Normalization is a normalization technique for training deep neural networks. To define Online Normalization. we replace arithmetic averages over the full dataset in…

Similar Papers 제목 키워드 기반

VarArray: Array-Geometry-Agnostic Continuous Speech Separation

2021-10-12 · Takuya Yoshioka, Xiaofei Wang, Dongmei Wang, Min Tang 외

Continuous speech separation using a microphone array was shown to be promising in dealing with the speech overlap problem in natural conversation transcription. This paper proposes VarArray, an array-geometry-agnostic s…

Speech Separation

Microphone Array Generalization for Multichannel Narrowband Deep Speech Enhancement

2021-07-27 · Siyuan Zhang, Xiaofei Li

This paper addresses the problem of microphone array generalization for deep-learning-based end-to-end multichannel speech enhancement. We aim to train a unique deep neural network (DNN) potentially performing well on un…

Speech Enhancement

AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition

2024-01-18 · Ju Lin, Niko Moritz, Yiteng Huang, Ruiming Xie 외

Wearable devices like smart glasses are approaching the compute capability to seamlessly generate real-time closed captions for live conversations. We build on our recently introduced directional Automatic Speech Recogni…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

VarArray Meets t-SOT: Advancing the State of the Art of Streaming Distant Conversational Speech Recognition

2022-09-12 · Naoyuki Kanda, Jian Wu, Xiaofei Wang, Zhuo Chen 외

This paper presents a novel streaming automatic speech recognition (ASR) framework for multi-talker overlapping speech captured by a distant microphone array with an arbitrary geometry. Our framework, named t-SOT-VA, cap…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Lightweight and Robust Multi-Channel End-to-End Speech Recognition with Spherical Harmonic Transform

2025-06-13 · Xiangzhu Kong, Huang Hao, Zhijian Ou

This paper presents SHTNet, a lightweight spherical harmonic transform (SHT) based framework, which is designed to address cross-array generalization challenges in multi-channel automatic speech recognition (ASR) through…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)channel selectionspeech-recognition+1