paper-with-me

홈 › Papers

One-pass Multiple Conformer and Foundation Speech Systems Compression and Quantization Using An All-in-one Neural Model

2024-06-14 · Zhaoqing Li, Haoning Xu, Tianzi Wang, Shoukang Hu, Zengrui Jin, Shujie Hu, Jiajun Deng, Mingyu Cui, Mengzhe Geng, Xunying Liu

We propose a novel one-pass multiple ASR systems joint compression and quantization approach using an all-in-one neural model. A single compression cycle allows multiple nested systems with varying Encoder depths, widths, and quantization precision settings to be simultaneously constructed without the need to train and store individual target systems separately. Experiments consistently demonstrate the multiple ASR systems compressed in a single all-in-one model produced a word error rate (WER) comparable to, or lower by up to 1.01\% absolute (6.98\% relative) than individually trained systems of equal complexity. A 3.4x overall system compression and training time speed-up was achieved. Maximum model size compression ratios of 12.8x and 3.93x were obtained over the baseline Switchboard-300hr Conformer and LibriSpeech-100hr fine-tuned wav2vec2.0 models, respectively, incurring no statistically significant WER increase.

📄 PDF Abstract BibTeX arXiv:2406.10160

Code (0)

등록된 구현이 없습니다.

Tasks

AllQuantization

Similar Papers 제목 키워드 기반

Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition

2024-07-03 · Shujie Hu, Xurong Xie, Mengzhe Geng, Zengrui Jin 외

Self-supervised learning (SSL) based speech foundation models have been applied to a wide range of ASR tasks. However, their application to dysarthric and elderly speech via data-intensive parameter fine-tuning is confro…

Alzheimer's Disease DetectionSelf-Supervised Learningspeech-recognitionSpeech Recognition

Two-pass Decoding and Cross-adaptation Based System Combination of End-to-end Conformer and Hybrid TDNN ASR Systems

2022-06-23 · Mingyu Cui, Jiajun Deng, Shoukang Hu, Xurong Xie 외

Fundamental modelling differences between hybrid and end-to-end (E2E) automatic speech recognition (ASR) systems create large diversity and complementarity among them. This paper investigates multi-pass rescoring and cro…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Diversityspeech-recognition+1

Conformer Based Elderly Speech Recognition System for Alzheimer's Disease Detection

2022-06-23 · Tianzi Wang, Jiajun Deng, Mengzhe Geng, Zi Ye 외

Early diagnosis of Alzheimer's disease (AD) is crucial in facilitating preventive care to delay further progression. This paper presents the development of a state-of-the-art Conformer based speech recognition system bui…

Alzheimer's Disease DetectionData AugmentationNeural Architecture Searchspeech-recognition+1

Exploring Self-supervised Pre-trained ASR Models For Dysarthric and Elderly Speech Recognition

2023-02-28 · Shujie Hu, Xurong Xie, Zengrui Jin, Mengzhe Geng 외

Automatic recognition of disordered and elderly speech remains a highly challenging task to date due to the difficulty in collecting such data in large quantities. This paper explores a series of approaches to integrate …

speech-recognitionSpeech Recognition

Conformer-Based Self-Supervised Learning for Non-Speech Audio Tasks

2021-10-14 · Sangeeta Srivastava, Yun Wang, Andros Tjandra, Anurag Kumar 외

Representation learning from unlabeled data has been of major interest in artificial intelligence research. While self-supervised speech representation learning has been popular in the speech research community, very few…

Audio ClassificationRepresentation LearningSelf-Supervised LearningSpeech Representation Learning