paper-with-me

Papers

AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition

2024-01-18 · Ju Lin, Niko Moritz, Yiteng Huang, Ruiming Xie, Ming Sun, Christian Fuegen, Frank Seide

Wearable devices like smart glasses are approaching the compute capability to seamlessly generate real-time closed captions for live conversations. We build on our recently introduced directional Automatic Speech Recognition (ASR) for smart glasses that have microphone arrays, which fuses multi-channel ASR with serialized output training, for wearer/conversation-partner disambiguation as well as suppression of cross-talk speech from non-target directions and noise. When ASR work is part of a broader system-development process, one may be faced with changes to microphone geometries as system development progresses. This paper aims to make multi-channel ASR insensitive to limited variations of microphone-array geometry. We show that a model trained on multiple similar geometries is largely agnostic and generalizes well to new geometries, as long as they are not too different. Furthermore, training the model this way improves accuracy for seen geometries by 15 to 28\% relative. Lastly, we refine the beamforming by a novel Non-Linearly Constrained Minimum Variance criterion.

📄 PDF Abstract BibTeX arXiv:2401.10411

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

One model to enhance them all: array geometry agnostic multi-channel personalized speech enhancement

2021-10-20 · Hassan Taherian, Sefik Emre Eskimez, Takuya Yoshioka, Huaming Wang 외

With the recent surge of video conferencing tools usage, providing high-quality speech signals and accurate captions have become essential to conduct day-to-day business or connect with friends and families. Single-chann…

AllAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancement+2

VarArray: Array-Geometry-Agnostic Continuous Speech Separation

2021-10-12 · Takuya Yoshioka, Xiaofei Wang, Dongmei Wang, Min Tang 외

Continuous speech separation using a microphone array was shown to be promising in dealing with the speech overlap problem in natural conversation transcription. This paper proposes VarArray, an array-geometry-agnostic s…

Speech Separation

VarArray Meets t-SOT: Advancing the State of the Art of Streaming Distant Conversational Speech Recognition

2022-09-12 · Naoyuki Kanda, Jian Wu, Xiaofei Wang, Zhuo Chen 외

This paper presents a novel streaming automatic speech recognition (ASR) framework for multi-talker overlapping speech captured by a distant microphone array with an arbitrary geometry. Our framework, named t-SOT-VA, cap…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Microphone Array Generalization for Multichannel Narrowband Deep Speech Enhancement

2021-07-27 · Siyuan Zhang, Xiaofei Li

This paper addresses the problem of microphone array generalization for deep-learning-based end-to-end multichannel speech enhancement. We aim to train a unique deep neural network (DNN) potentially performing well on un…

Speech Enhancement

ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior

2025-05-08 · Zhongweiyang Xu, Xulin Fan, Zhong-Qiu Wang, Xilin Jiang 외

Blind Speech Separation (BSS) aims to separate multiple speech sources from audio mixtures recorded by a microphone array. The problem is challenging because it is a blind inverse problem, i.e., the microphone array geom…

Room Impulse Response (RIR)Speech Separation