paper-with-me

Papers

A Wavelet Transform Based Scheme to Extract Speech Pitch and Formant Frequencies

2022-09-01 · Seyedamiryousef Hosseini Goki, Mahdieh Ghazvini, Sajad Hamzenejadi

Pitch and Formant frequencies are important features in speech processing applications. The period of the vocal cord's output for vowels is known as the pitch or the fundamental frequency, and formant frequencies are essentially resonance frequencies of the vocal tract. These features vary among different persons and even words, but they are within a certain frequency range. In practice, just the first three formants are enough for the most of speech processing. Feature extraction and classification are the main components of each speech recognition system. In this article, two wavelet based approaches are proposed to extract the mentioned features with help of the filter bank idea. By comparing the results of the presented feature extraction methods on several speech signals, it was found out that the wavelet transform has a good accuracy compared to the cepstrum method and it has no sensitivity to noise. In addition, several fuzzy based classification techniques for speech processing are reviewed.

📄 PDF Abstract BibTeX arXiv:2209.00733

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Audio Signal Processing Using Time Domain Mel-Frequency Wavelet Coefficient

2025-10-28 · Rinku Sebastian, Simon O'Keefe, Martin Trefzer arxiv

Extracting features from the speech is the most critical process in speech signal processing. Mel Frequency Cepstral Coefficients (MFCC) are the most widely used features in the majority of the speaker and speech recogni…

Speech Recognition

Wavelet Scattering on the Pitch Spiral

2016-01-03 · Vincent Lostanlen, Stéphane Mallat

We present a new representation of harmonic sounds that linearizes the dynamics of pitch and spectral envelope, while remaining stable to deformations in the time-frequency plane. It is an instance of the scattering tran…

Comparative Analysis of Mel-Frequency Cepstral Coefficients and Wavelet Based Audio Signal Processing for Emotion Detection and Mental Health Assessment in Spoken Speech

2024-12-12 · Idoko Agbo, Dr Hoda El-Sayed, M. D Kamruzzan Sarker

The intersection of technology and mental health has spurred innovative approaches to assessing emotional well-being, particularly through computational techniques applied to audio data analysis. This study explores the …

Audio Signal ProcessingData Augmentation

Training a Neural Speech Waveform Model using Spectral Losses of Short-Time Fourier Transform and Continuous Wavelet Transform

2019-03-29 · Shinji Takaki, Hirokazu Kameoka, Junichi Yamagishi

Recently, we proposed short-time Fourier transform (STFT)-based loss functions for training a neural speech waveform model. In this paper, we generalize the above framework and propose a training scheme for such models b…

hf0: A hybrid pitch extraction method for multimodal voice

2019-04-22 · Pradeep Rengaswamy, Gurunath Reddy M, Krothapalli Sreenivasa Rao

Pitch or fundamental frequency (f0) extraction is a fundamental problem studied extensively for its potential applications in speech and clinical applications. In literature, explicit mode specific (modal speech or singi…

Deep Learning