paper-with-me

홈 › Papers

Unsupervised Classification of Voiced Speech and Pitch Tracking Using Forward-Backward Kalman Filtering

2021-03-01 · Benedikt Boenninghoff, Robert M. Nickel, Steffen Zeiler, Dorothea Kolossa

The detection of voiced speech, the estimation of the fundamental frequency, and the tracking of pitch values over time are crucial subtasks for a variety of speech processing techniques. Many different algorithms have been developed for each of the three subtasks. We present a new algorithm that integrates the three subtasks into a single procedure. The algorithm can be applied to pre-recorded speech utterances in the presence of considerable amounts of background noise. We combine a collection of standard metrics, such as the zero-crossing rate, for example, to formulate an unsupervised voicing classifier. The estimation of pitch values is accomplished with a hybrid autocorrelation-based technique. We propose a forward-backward Kalman filter to smooth the estimated pitch contour. In experiments, we are able to show that the proposed method compares favorably with current, state-of-the-art pitch detection algorithms.

📄 PDF Abstract BibTeX arXiv:2103.01173

Code (0)

등록된 구현이 없습니다.

Tasks

General Classification

Similar Papers 제목 키워드 기반

A Fast and Accurate Pitch Estimation Algorithm Based on the Pseudo Wigner-Ville Distribution

2022-10-27 · Yisi Liu, Peter Wu, Alan W Black, Gopala K. Anumanchipalli

Estimation of fundamental frequency (F0) in voiced segments of speech signals, also known as pitch tracking, plays a crucial role in pitch synchronous speech analysis, speech synthesis, and speech manipulation. In this p…

Speech Synthesis

Whispered-to-voiced Alaryngeal Speech Conversion with Generative Adversarial Networks

2018-08-31 · Santiago Pascual, Antonio Bonafonte, Joan Serrà, Jose A. Gonzalez

Most methods of voice restoration for patients suffering from aphonia either produce whispered or monotone speech. Apart from intelligibility, this type of speech lacks expressiveness and naturalness due to the absence o…

Speech EnhancementSpeech Recognition

rVAD: An Unsupervised Segment-Based Robust Voice Activity Detection Method

2019-06-09 · Zheng-Hua Tan, Achintya Kr. Sarkar, Najim Dehak

This paper presents an unsupervised segment-based method for robust voice activity detection (rVAD). The method consists of two passes of denoising followed by a voice activity detection (VAD) stage. In the first pass, h…

Action DetectionActivity DetectionDenoisingSpeaker Verification+1

Cross-domain Neural Pitch and Periodicity Estimation

2023-01-28 · Max Morrison, Caedon Hsieh, Nathan Pruyne, Bryan Pardo

Pitch is a foundational aspect of our perception of audio signals. Pitch contours are commonly used to analyze speech and music signals and as input features for many audio tasks, including music transcription, singing v…

CPUGPUMusic TranscriptionSinging Voice Synthesis

Improving Automatic Emotion Recognition from speech using Rhythm and Temporal feature

2013-03-07 · Mayank Bhargava, Tim Polzehl

This paper is devoted to improve automatic emotion recognition from speech by incorporating rhythm and temporal features. Research on automatic emotion recognition so far has mostly been based on applying features like M…

Emotion RecognitionRhythm