Performance of using Mel-Frequency Cepstrum Based Features in Nonlinear Classifiers for Phonocardiography Recordings
Cardiovascular system diseases can be identified by using a specialized diagnostic process utilizing a digital stethoscope. Digital stethoscopes provide phonocardiography (PCG) recordings for further inspection, besides filtering and amplification of heart sounds. In this paper, a framework that is useful to develop feature extraction and classification of PCG recordings is presented. This framework is built upon a previously proposed segmentation algorithm that processes a feature vector produced by the agglutinate application of Mel-frequency cepstrum and discrete wavelet transform (DWT). The performance of the segmentation algorithm is also tested on a new data set and compared to the previously reported results. After identifying the fundamental heart sounds and segmenting the PCG recordings, five principal features are extracted from the time domain signal and Mel-Frequency cepstral coefficients (MFCC) of each cardiac cycle. Classification outcomes are reported for three nonlinear models: k nearest neighbor (k-NN), support vector machine (SVM), and multilayer perceptrons (MLP) classifiers in comparison with a linear approach, namely Mahalanobis distance linear classifier. The results underline that although neural networks and linear classifier show compatible performance in basic classification problems, with the increase in the nonlinearity of the classification problem their performance significantly vary.
Code (0)
등록된 구현이 없습니다.
Tasks
ClassificationDiagnosticMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Multi-layered Cepstrum for Instantaneous Frequency Estimation
We propose the multi-layered cepstrum (MLC) method to estimate multiple fundamental frequencies (MF0) of a signal under challenging contamination such as high-pass filter noise. Taking the operation of cepstrum (i.e., Fo…
Improved Frame Level Features and SVM Supervectors Approach for the Recogniton of Emotional States from Speech: Application to categorical and dimensional states
The purpose of speech emotion recognition system is to classify speakers utterances into different emotional states such as disgust, boredom, sadness, neutral and happiness. Speech features that are commonly used in spee…
Emotion RecognitionSpeech Emotion RecognitionCycleGAN-VC3: Examining and Improving CycleGAN-VCs for Mel-spectrogram Conversion
Non-parallel voice conversion (VC) is a technique for learning mappings between source and target speeches without using a parallel corpus. Recently, cycle-consistent adversarial network (CycleGAN)-VC and CycleGAN-VC2 ha…
Voice ConversionVoiceprint recognition of Parkinson patients based on deep learning
More than 90% of the Parkinson Disease (PD) patients suffer from vocal disorders. Speech impairment is already indicator of PD. This study focuses on PD diagnosis through voiceprint features. In this paper, a method base…
Deep LearningGeneral ClassificationPseudo-Cepstrum: Pitch Modification for Mel-Based Neural Vocoders
This paper introduces a cepstrum-based pitch modification method that can be applied to any mel-spectrogram representation. As a result, this method is compatible with any mel-based vocoder without requiring any addition…