paper-with-me

홈 › Papers

WeNet 2.0: More Productive End-to-End Speech Recognition Toolkit

2022-03-29 · BinBin Zhang, Di wu, Zhendong Peng, Xingchen Song, Zhuoyuan Yao, Hang Lv, Lei Xie, Chao Yang, Fuping Pan, Jianwei Niu

Recently, we made available WeNet, a production-oriented end-to-end speech recognition toolkit, which introduces a unified two-pass (U2) framework and a built-in runtime to address the streaming and non-streaming decoding modes in a single model. To further improve ASR performance and facilitate various production requirements, in this paper, we present WeNet 2.0 with four important updates. (1) We propose U2++, a unified two-pass framework with bidirectional attention decoders, which includes the future contextual information by a right-to-left attention decoder to improve the representative ability of the shared encoder and the performance during the rescoring stage. (2) We introduce an n-gram based language model and a WFST-based decoder into WeNet 2.0, promoting the use of rich text data in production scenarios. (3) We design a unified contextual biasing framework, which leverages user-specific context (e.g., contact lists) to provide rapid adaptation ability for production and improves ASR accuracy in both with-LM and without-LM scenarios. (4) We design a unified IO to support large-scale data for effective model training. In summary, the brand-new WeNet 2.0 achieves up to 10\% relative recognition performance improvement over the original WeNet on various corpora and makes available several important production-oriented features.

📄 PDF Abstract BibTeX arXiv:2203.15455

Code (3)

wenet-e2e/wenet 공식 구현 pytorch
leonwlw/wenet pytorch
mobvoi/wenet pytorch

Tasks

DecoderLanguage Modellingspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

WeNet: Production oriented Streaming and Non-streaming End-to-End Speech Recognition Toolkit

2021-02-02 · Zhuoyuan Yao, Di wu, Xiong Wang, BinBin Zhang 외

In this paper, we propose an open source, production first, and production ready speech recognition toolkit called WeNet in which a new two-pass approach is implemented to unify streaming and non-streaming end-to-end (E2…

Decoderspeech-recognitionSpeech Recognition

WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition

2021-10-07 · BinBin Zhang, Hang Lv, Pengcheng Guo, Qijie Shao 외

In this paper, we present WenetSpeech, a multi-domain Mandarin corpus consisting of 10000+ hours high-quality labeled speech, 2400+ hours weakly labeled speech, and about 10000 hours unlabeled speech, with 22400+ hours i…

Label Error DetectionOptical Character RecognitionOptical Character Recognition (OCR)speech-recognition+2

TALCS: An Open-Source Mandarin-English Code-Switching Corpus and a Speech Recognition Baseline

2022-06-27 · Chengfei Li, Shuhao Deng, Yaoping Wang, Guangjing Wang 외

This paper introduces a new corpus of Mandarin-English code-switching speech recognition--TALCS corpus, suitable for training and evaluating code-switching speech recognition systems. TALCS corpus is derived from real on…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction

2024-09-24 · Shuai Wang, Ke Zhang, Shaoxiong Lin, Junjie Li 외

Target speaker extraction (TSE) focuses on isolating the speech of a specific target speaker from overlapped multi-talker speech, which is a typical setup in the cocktail party problem. In recent years, TSE draws increas…

Managementspeech-recognitionSpeech RecognitionTarget Speaker Extraction

WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction

2025-09-24 · Binbin Zhang, Chengdong Liang, Shuai Wang, Xuelong Geng 외 arxiv

In this paper, we present WEST(WE Speech Toolkit), a speech toolkit based on a large language model (LLM) for speech understanding, generation, and interaction. There are three key features of WEST: 1) Fully LLM-based: S…