paper-with-me

Papers

A Robust Frame-based Nonlinear Prediction System for Automatic Speech Coding

2016-01-22 · Mahmood Yousefi-Azar, Farbod Razzazi

In this paper, we propose a neural-based coding scheme in which an artificial neural network is exploited to automatically compress and decompress speech signals by a trainable approach. Having a two-stage training phase, the system can be fully specified to each speech frame and have robust performance across different speakers and wide range of spoken utterances. Indeed, Frame-based nonlinear predictive coding (FNPC) would code a frame in the procedure of training to predict the frame samples. The motivating objective is to analyze the system behavior in regenerating not only the envelope of spectra, but also the spectra phase. This scheme has been evaluated in time and discrete cosine transform (DCT) domains and the output of predicted phonemes show the potentiality of the FNPC to reconstruct complicated signals. The experiments were conducted on three voiced plosive phonemes, b/d/g/ in time and DCT domains versus the number of neurons in the hidden layer. Experiments approve the FNPC capability as an automatic coding system by which /b/d/g/ phonemes have been reproduced with a good accuracy. Evaluations revealed that the performance of FNPC system, trained to predict DCT coefficients is more desirable, particularly for frames with the wider distribution of energy, compared to time samples.

📄 PDF Abstract BibTeX arXiv:1601.06008

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multi-channel Multi-frame ADL-MVDR for Target Speech Separation

2020-12-24 · Zhuohuang Zhang, Yong Xu, Meng Yu, Shi-Xiong Zhang 외

Many purely neural network based speech separation approaches have been proposed to improve objective assessment scores, but they often introduce nonlinear distortions that are harmful to modern automatic speech recognit…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

ADPCM with nonlinear prediction

2022-02-24 · Marcos Faundez-Zanuy, Oscar Oliva-Suarez

Many speech coders are based on linear prediction coding (LPC), nevertheless with LPC is not possible to model the nonlinearities present in the speech signal. Because of this there is a growing interest for nonlinear te…

Prediction

Nonlinear predictive models computation in ADPCM schemes

2022-03-03 · Marcos Faundez-Zanuy

Recently several papers have been published on nonlinear prediction applied to speech coding. At ICASSP98 we presented a system based on an ADPCM scheme with a nonlinear predictor based on a neural net. The most critical…

Intelligibility prediction with a pretrained noise-robust automatic speech recognition model

2023-10-20 · Zehai Tu, Ning Ma, Jon Barker

This paper describes two intelligibility prediction systems derived from a pretrained noise-robust automatic speech recognition (ASR) model for the second Clarity Prediction Challenge (CPC2). One system is intrusive and …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Predictionspeech-recognition+1

Comparison of Speech Representations for Automatic Quality Estimation in Multi-Speaker Text-to-Speech Synthesis

2020-02-28 · Jennifer Williams, Joanna Rownicka, Pilar Oplustil, Simon King

We aim to characterize how different speakers contribute to the perceived output quality of multi-speaker Text-to-Speech (TTS) synthesis. We automatically rate the quality of TTS using a neural network (NN) trained on hu…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis+1