paper-with-me

Papers

Flexi-Transducer: Optimizing Latency, Accuracy and Compute forMulti-Domain On-Device Scenarios

2021-04-06 · Jay Mahadeokar, Yangyang Shi, Yuan Shangguan, Chunyang Wu, Alex Xiao, Hang Su, Duc Le, Ozlem Kalinli, Christian Fuegen, Michael L. Seltzer

Often, the storage and computational constraints of embeddeddevices demand that a single on-device ASR model serve multiple use-cases / domains. In this paper, we propose aFlexibleTransducer(FlexiT) for on-device automatic speech recognition to flexibly deal with multiple use-cases / domains with different accuracy and latency requirements. Specifically, using a single compact model, FlexiT provides a fast response for voice commands, and accurate transcription but with more latency for dictation. In order to achieve flexible and better accuracy and latency trade-offs, the following techniques are used. Firstly, we propose using domain-specific altering of segment size for Emformer encoder that enables FlexiT to achieve flexible de-coding. Secondly, we use Alignment Restricted RNNT loss to achieve flexible fine-grained control on token emission latency for different domains. Finally, we add a domain indicator vector as an additional input to the FlexiT model. Using the combination of techniques, we show that a single model can be used to improve WERs and real time factor for dictation scenarios while maintaining optimal latency for voice commands use-cases

📄 PDF Abstract BibTeX arXiv:2104.02232

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Minimum Latency Training of Sequence Transducers for Streaming End-to-End Speech Recognition

2022-11-04 · Yusuke Shinohara, Shinji Watanabe

Sequence transducers, such as the RNN-T and the Conformer-T, are one of the most promising models of end-to-end speech recognition, especially in streaming scenarios where both latency and accuracy are important. Althoug…

speech-recognitionSpeech Recognition

Dynamic Encoder Transducer: A Flexible Solution For Trading Off Accuracy For Latency

2021-04-05 · Yangyang Shi, Varun Nagaraja, Chunyang Wu, Jay Mahadeokar 외

We propose a dynamic encoder transducer (DET) for on-device speech recognition. One DET model scales to multiple devices with different computation capacities without retraining or finetuning. To trading off accuracy and…

speech-recognitionSpeech Recognition

Delay-penalized CTC implemented based on Finite State Transducer

2023-05-19 · Zengwei Yao, Wei Kang, Fangjun Kuang, Liyong Guo 외

Connectionist Temporal Classification (CTC) suffers from the latency problem when applied to streaming models. We argue that in CTC lattice, the alignments that can access more future context are preferred during trainin…

Attribute

FastEmit: Low-latency Streaming ASR with Sequence-level Emission Regularization

2020-10-21 · Jiahui Yu, Chung-Cheng Chiu, Bo Li, Shuo-Yiin Chang 외

Streaming automatic speech recognition (ASR) aims to emit each hypothesized word as quickly and accurately as possible. However, emitting fast without degrading quality, as measured by word error rate (WER), is highly ch…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

PEtra: A Flexible and Open-Source PE Loop Tracer for Polymer Thin-Film Transducers

2024-10-21 · Marc-Andre Wessner, Federico Villani, Sofia Papa, Kirill Keller 외

Accurate characterization of ferroelectric properties in polymer piezoelectrics is critical for optimizing the performance of flexible and wearable ultrasound transducers, such as screen-printed PVDF devices. Standard ch…