paper-with-me

홈 › Papers

Separator-Transducer-Segmenter: Streaming Recognition and Segmentation of Multi-party Speech

2022-05-10 · Ilya Sklyar, Anna Piunova, Christian Osendorfer

Streaming recognition and segmentation of multi-party conversations with overlapping speech is crucial for the next generation of voice assistant applications. In this work we address its challenges discovered in the previous work on multi-turn recurrent neural network transducer (MT-RNN-T) with a novel approach, separator-transducer-segmenter (STS), that enables tighter integration of speech separation, recognition and segmentation in a single model. First, we propose a new segmentation modeling strategy through start-of-turn and end-of-turn tokens that improves segmentation without recognition accuracy degradation. Second, we further improve both speech recognition and segmentation accuracy through an emission regularization method, FastEmit, and multi-task training with speech activity information as an additional training signal. Third, we experiment with end-of-turn emission latency penalty to improve end-point detection for each speaker turn. Finally, we establish a novel framework for segmentation analysis of multi-party conversations through emission latency metrics. With our best model, we report 4.6% abs. turn counting accuracy improvement and 17% rel. word error rate (WER) improvement on LibriCSS dataset compared to the previously published work.

📄 PDF Abstract BibTeX arXiv:2205.05199

Code (0)

등록된 구현이 없습니다.

Tasks

Segmentationspeech-recognitionSpeech RecognitionSpeech SeparationSTS

Similar Papers 제목 키워드 기반

Improving Fast-slow Encoder based Transducer with Streaming Deliberation

2022-12-15 · Ke Li, Jay Mahadeokar, Jinxi Guo, Yangyang Shi 외

This paper introduces a fast-slow encoder based transducer with streaming deliberation for end-to-end automatic speech recognition. We aim to improve the recognition accuracy of the fast-slow encoder based transducer whi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Globally Normalising the Transducer for Streaming Speech Recognition

2023-07-20 · Rogier Van Dalen

The Transducer (e.g. RNN-Transducer or Conformer-Transducer) generates an output label sequence as it traverses the input sequence. It is straightforward to use in streaming mode, where it generates partial hypotheses be…

speech-recognitionSpeech Recognition

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition

2024-11-26 · Hyeonseung Lee, Ji Won Yoon, Sungsoo Kim, Nam Soo Kim

Transducer neural networks have emerged as the mainstream approach for streaming automatic speech recognition (ASR), offering state-of-the-art performance in balancing accuracy and latency. In the conventional framework,…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Direct Segmentation Models for Streaming Speech Translation

2020-11-01 · EMNLP 2020 11 · Javier Iranzo-S{\'a}nchez, Adri{\`a} Gim{\'e}nez Pastor, Joan Albert Silvestre-Cerd{\`a}, Pau Baquero-Arnal 외

The cascade approach to Speech Translation (ST) is based on a pipeline that concatenates an Automatic Speech Recognition (ASR) system followed by a Machine Translation (MT) system. These systems are usually connected by …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationSegmentation+3

LAMASSU: Streaming Language-Agnostic Multilingual Speech Recognition and Translation Using Neural Transducers

2022-11-05 · Peidong Wang, Eric Sun, Jian Xue, Yu Wu 외

Automatic speech recognition (ASR) and speech translation (ST) can both use neural transducers as the model structure. It is thus possible to use a single transducer model to perform both tasks. In real-world application…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Identificationspeech-recognition+3