paper-with-me

홈 › Papers

DEEPF0: End-To-End Fundamental Frequency Estimation for Music and Speech Signals

2021-02-11 · Satwinder Singh, Ruili Wang, Yuanhang Qiu

We propose a novel pitch estimation technique called DeepF0, which leverages the available annotated data to directly learns from the raw audio in a data-driven manner. F0 estimation is important in various speech processing and music information retrieval applications. Existing deep learning models for pitch estimations have relatively limited learning capabilities due to their shallow receptive field. The proposed model addresses this issue by extending the receptive field of a network by introducing the dilated convolutional blocks into the network. The dilation factor increases the network receptive field exponentially without increasing the parameters of the model exponentially. To make the training process more efficient and faster, DeepF0 is augmented with residual blocks with residual connections. Our empirical evaluation demonstrates that the proposed model outperforms the baselines in terms of raw pitch accuracy and raw chroma accuracy even using 77.4% fewer network parameters. We also show that our model can capture reasonably well pitch estimation even under the various levels of accompaniment noise.

📄 PDF Abstract BibTeX arXiv:2102.06306

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalMusic Information RetrievalRetrieval

Similar Papers 제목 키워드 기반

Harmonic enhancement using learnable comb filter for light-weight full-band speech enhancement model

2023-06-01 · Xiaohuai Le, Tong Lei, Li Chen, Yiqing Guo 외

With fewer feature dimensions, filter banks are often used in light-weight full-band speech enhancement models. In order to further enhance the coarse speech in the sub-band domain, it is necessary to apply a post-filter…

RetrievalSpeech Enhancement

Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory Perception

2025-04-09 · Yuankun Xie, Ruibo Fu, Zhiyong Wang, Xiaopeng Wang 외

The rapid advancement of audio generation technologies has escalated the risks of malicious deepfake audio across speech, sound, singing voice, and music, threatening multimedia security and trust. While existing counter…

AllAudio Deepfake DetectionAudio GenerationDeepFake Detection+2

Audio Deepfake Detection Based on a Combination of F0 Information and Real Plus Imaginary Spectrogram Features

2022-08-02 · Jun Xue, Cunhang Fan, Zhao Lv, JianHua Tao 외

Recently, pioneer research works have proposed a large number of acoustic features (log power spectrogram, linear frequency cepstral coefficients, constant Q cepstral coefficients, etc.) for audio deepfake detection, obt…

Audio Deepfake DetectionDeepFake DetectionFace Swapping

PHAST-Net: Attention-Guided, Physics-Informed Network for Unified Estimation of Ideal Time-Frequency Representations

2026-06-22 · James M. Cozens, Simon J. Godsill arxiv

We introduce PHAST-Net, an attention-guided, physics-informed network for unified estimation of Ideal Time-Frequency Representations (ITFRs), spanning spectral, tempo-based, metrical, and harmonic representations such as…

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan

2026-04-09 · Yuankun Xie, Haonan Cheng, Jiayi Zhou, Xiaoxuan Guo 외 arxiv

The rapid advancement of Audio Large Language Models (ALLMs) has enabled cost-effective, high-fidelity generation and manipulation of both speech and non-speech audio, including sound effects, singing voices, and music. …

Audio Deepfake Detection