paper-with-me

홈 › Papers

TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch

2024-12-20 · Xingchen Song, Chengdong Liang, BinBin Zhang, Pengshen Zhang, Ziyu Wang, Youcheng Ma, Menglong Xu, Lin Wang, Di wu, Fuping Pan, Dinghao Zhou, Zhendong Peng

Large Automatic Speech Recognition (ASR) models demand a vast number of parameters, copious amounts of data, and significant computational resources during the training process. However, such models can merely be deployed on high-compute cloud platforms and are only capable of performing speech recognition tasks. This leads to high costs and restricted capabilities. In this report, we initially propose the elastic mixture of the expert (eMoE) model. This model can be trained just once and then be elastically scaled in accordance with deployment requirements. Secondly, we devise an unsupervised data creation and validation procedure and gather millions of hours of audio data from diverse domains for training. Using these two techniques, our system achieves elastic deployment capabilities while reducing the Character Error Rate (CER) on the SpeechIO testsets from 4.98\% to 2.45\%. Thirdly, our model is not only competent in Mandarin speech recognition but also proficient in multilingual, multi-dialect, emotion, gender, and sound event perception. We refer to this as Automatic Speech Perception (ASP), and the perception results are presented in the experimental section.

📄 PDF Abstract BibTeX arXiv:2412.15622

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Speech Technology for Everyone: Automatic Speech Recognition for Non-Native English

2021-11-01 · ICNLSP 2021 11 · Toshiko Shibano, Xinyi Zhang, Mia Taige Li, Haejin Cho 외
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Everyone deserves their voice to be heard: Analyzing Predictive Gender Bias in ASR Models Applied to Dutch Speech Data

2024-11-14 · Rik Raes, Saskia Lensink, Mykola Pechenizkiy

Recent research has shown that state-of-the-art (SotA) Automatic Speech Recognition (ASR) systems, such as Whisper, often exhibit predictive biases that disproportionately affect various demographic groups. This study fo…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)FairnessSemantic Similarity+3

Speech Technology for Everyone: Automatic Speech Recognition for Non-Native English with Transfer Learning

2021-10-01 · Toshiko Shibano, Xinyi Zhang, Mia Taige Li, Haejin Cho 외

To address the performance gap of English ASR models on L2 English speakers, we evaluate fine-tuning of pretrained wav2vec 2.0 models (Baevski et al., 2020; Xu et al., 2021) on L2-ARCTIC, a non-native English speech corp…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+3

Is the price right? Reconceptualizing price and income elasticity to anticipate price perception issues

2024-02-07 · Shawn Berry

Price perception by consumers represents a challenge to the ability of a business to correctly and profitably price and sell their products or services in a given market and any new target market. Complicating the percep…

Predictive coding and stochastic resonance as fundamental principles of auditory perception

2022-04-07 · Achim Schilling, William Sedley, Richard Gerum, Claus Metzner 외

How is information processed in the brain during perception? Mechanistic insight is achieved only when experiments are employed to test formal or computational models. In analogy to lesion studies, phantom perception may…

Bayesian Inference