paper-with-me

홈 › Papers

PSVRF: Learning to restore Pitch-Shifted Voice without reference

2022-10-06 · Yangfu Li, Xiaodan Lin, Jiaxin Yang

Pitch scaling algorithms have a significant impact on the security of Automatic Speaker Verification (ASV) systems. Although numerous anti-spoofing algorithms have been proposed to identify the pitch-shifted voice and even restore it to the original version, they either have poor performance or require the original voice as a reference, limiting the prospects of applications. In this paper, we propose a no-reference approach termed PSVRF$^1$ for high-quality restoration of pitch-shifted voice. Experiments on AISHELL-1 and AISHELL-3 demonstrate that PSVRF can restore the voice disguised by various pitch-scaling techniques, which obviously enhances the robustness of ASV systems to pitch-scaling attacks. Furthermore, the performance of PSVRF even surpasses that of the state-of-the-art reference-based approach.

📄 PDF Abstract BibTeX arXiv:2210.02731

Code (1)

ychenl/pssrf 공식 구현 pytorch

Tasks

Speaker Verification

Similar Papers 제목 키워드 기반

When Automatic Voice Disguise Meets Automatic Speaker Verification

2020-09-15 · Linlin Zheng, Jiakang Li, Meng Sun, Xiongwei Zhang 외

The technique of transforming voices in order to hide the real identity of a speaker is called voice disguise, among which automatic voice disguise (AVD) by modifying the spectral and temporal characteristics of voices w…

MiscellaneousSpeaker VerificationVoice Conversion

SPICE: Self-supervised Pitch Estimation

2019-10-25 · Beat Gfeller, Christian Frank, Dominik Roblek, Matt Sharifi 외

We propose a model to estimate the fundamental frequency in monophonic audio, often referred to as pitch estimation. We acknowledge the fact that obtaining ground truth annotations at the required temporal and frequency …

Self-Supervised LearningTranslation

Deep Autotuner: A Data-Driven Approach to Natural-Sounding Pitch Correction for Singing Voice in Karaoke Performances

2019-02-03 · Sanna Wager, George Tzanetakis, Cheng-i Wang, Lijiang Guo 외

We describe a machine-learning approach to pitch correcting a solo singing performance in a karaoke setting, where the solo voice and accompaniment are on separate tracks. The proposed approach addresses the situation wh…

Personalized Voice Synthesis through Human-in-the-Loop Coordinate Descent

2024-08-30 · Yusheng Tian, Junbin Liu, Tan Lee

This paper describes a human-in-the-loop approach to personalized voice synthesis in the absence of reference speech data from the target speaker. It is intended to help vocally disabled individuals restore their lost vo…

Whispered-to-voiced Alaryngeal Speech Conversion with Generative Adversarial Networks

2018-08-31 · Santiago Pascual, Antonio Bonafonte, Joan Serrà, Jose A. Gonzalez

Most methods of voice restoration for patients suffering from aphonia either produce whispered or monotone speech. Apart from intelligibility, this type of speech lacks expressiveness and naturalness due to the absence o…

Speech EnhancementSpeech Recognition