paper-with-me

Papers

CUSIDE-array: A Streaming Multi-Channel End-to-End Speech Recognition System with Realistic Evaluations

2024-07-13 · Xiangzhu Kong, Tianqi Ning, Hao Huang, Zhijian Ou

Recently multi-channel end-to-end (ME2E) ASR systems have emerged. While streaming single-channel end-to-end ASR has been extensively studied, streaming ME2E ASR is limited in exploration. Additionally, recent studies call attention to the gap between in-distribution (ID) and out-of-distribution (OOD) tests and doing realistic evaluations. This paper focuses on two research problems: realizing streaming ME2E ASR and improving OOD generalization. We propose the CUSIDE-array method, which integrates the recent CUSIDE methodology (Chunking, Simulating Future Context and Decoding) into the neural beamformer approach of ME2E ASR. It enables streaming processing of both front-end and back-end with a total latency of 402ms. The CUSIDE-array ME2E models are shown to achieve superior streaming results in both ID and OOD tests. Realistic evaluations confirm the advantage of CUSIDE-array in its capability to consume single-channel data to improve OOD generalization via back-end pre-training and ME2E fine-tuning.

📄 PDF Abstract BibTeX arXiv:2407.09807

Code (1)

thu-spmi/cat 공식 구현 pytorch

Tasks

Chunkingspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

CUSIDE: Chunking, Simulating Future Context and Decoding for Streaming ASR

2022-03-31 · Keyu An, Huahuan Zheng, Zhijian Ou, Hongyu Xiang 외

History and future contextual information are known to be important for accurate acoustic modeling. However, acquiring future context brings latency for streaming ASR. In this paper, we propose a new framework - Chunking…

Chunkingspeech-recognitionSpeech Recognition

Neural Directed Speech Enhancement with Dual Microphone Array in High Noise Scenario

2024-12-24 · Wen Wen, Qiang Zhou, Yu Xi, Haoyu Li 외

In multi-speaker scenarios, leveraging spatial features is essential for enhancing target speech. While with limited microphone arrays, developing a compact multi-channel speech enhancement system remains challenging, es…

Speech Enhancement

Cleanformer: A multichannel array configuration-invariant neural enhancement frontend for ASR in smart speakers

2022-04-25 · Joseph Caroselli, Arun Narayanan, Nathan Howard, Tom O'Malley

This work introduces the Cleanformer, a streaming multichannel neural based enhancement frontend for automatic speech recognition (ASR). This model has a conformer-based architecture which takes as inputs a single channe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Multi-Geometry Spatial Acoustic Modeling for Distant Speech Recognition

2019-04-28

The use of spatial information with multiple microphones can improve far-field automatic speech recognition (ASR) accuracy. However, conventional microphone array techniques degrade speech enhancement performance when th…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Deep ClusteringDistant Speech Recognition+3

VarArray Meets t-SOT: Advancing the State of the Art of Streaming Distant Conversational Speech Recognition

2022-09-12 · Naoyuki Kanda, Jian Wu, Xiaofei Wang, Zhuo Chen 외

This paper presents a novel streaming automatic speech recognition (ASR) framework for multi-talker overlapping speech captured by a distant microphone array with an arbitrary geometry. Our framework, named t-SOT-VA, cap…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1