paper-with-me

홈 › Papers

Streaming Noise Context Aware Enhancement For Automatic Speech Recognition in Multi-Talker Environments

2022-05-17 · Joe Caroselli, Arun Narayanan, Yiteng Huang

One of the most challenging scenarios for smart speakers is multi-talker, when target speech from the desired speaker is mixed with interfering speech from one or more speakers. A smart assistant needs to determine which voice to recognize and which to ignore and it needs to do so in a streaming, low-latency manner. This work presents two multi-microphone speech enhancement algorithms targeted at this scenario. Targeting on-device use-cases, we assume that the algorithm has access to the signal before the hotword, which is referred to as the noise context. First is the Context Aware Beamformer which uses the noise context and detected hotword to determine how to target the desired speaker. The second is an adaptive noise cancellation algorithm called Speech Cleaner which trains a filter using the noise context. It is demonstrated that the two algorithms are complementary in the signal-to-noise ratio conditions under which they work well. We also propose an algorithm to select which one to use based on estimated SNR. When using 3 microphone channels, the final system achieves a relative word error rate reduction of 55% at -12dB, and 43\% at 12dB.

📄 PDF Abstract BibTeX arXiv:2205.08555

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

Cleanformer: A multichannel array configuration-invariant neural enhancement frontend for ASR in smart speakers

2022-04-25 · Joseph Caroselli, Arun Narayanan, Nathan Howard, Tom O'Malley

This work introduces the Cleanformer, a streaming multichannel neural based enhancement frontend for automatic speech recognition (ASR). This model has a conformer-based architecture which takes as inputs a single channe…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

DPSNN: Spiking Neural Network for Low-Latency Streaming Speech Enhancement

2024-08-14 · Tao Sun, Sander Bohté

Speech enhancement (SE) improves communication in noisy environments, affecting areas such as automatic speech recognition, hearing aids, and telecommunications. With these domains typically being power-constrained and e…

Automatic Speech RecognitionSpeech Enhancementspeech-recognitionSpeech Recognition

Dynamic Context-Aware Streaming Pretrained Language Model For Inverse Text Normalization

2025-05-30 · Luong Ho, Khanh Le, Vinh Pham, Bao Nguyen 외

Inverse Text Normalization (ITN) is crucial for converting spoken Automatic Speech Recognition (ASR) outputs into well-formatted written text, enhancing both readability and usability. Despite its importance, the integra…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+3

Enhancement or Super-Resolution: Learning-based Adaptive Video Streaming with Client-Side Video Processing

2022-01-20 · Junyan Yang, Yang Jiang, Shuoyao Wang

The rapid development of multimedia and communication technology has resulted in an urgent need for high-quality video streaming. However, robust video streaming under fluctuating network conditions and heterogeneous cli…

Deep Reinforcement LearningSuper-ResolutionVideo CompressionVideo Enhancement

StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion

2024-08-05 · Zhichao Wang, Yuanzhe Chen, Xinsheng Wang, Lei Xie 외

StreamVoice has recently pushed the boundaries of zero-shot voice conversion (VC) in the streaming domain. It uses a streamable language model (LM) with a context-aware approach to convert semantic features from automati…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Modellingspeech-recognition+2