paper-with-me

Papers

Can we steal your vocal identity from the Internet?: Initial investigation of cloning Obama's voice using GAN, WaveNet and low-quality found data

2018-03-02 · Jaime Lorenzo-Trueba, Fuming Fang, Xin Wang, Isao Echizen, Junichi Yamagishi, Tomi Kinnunen

Thanks to the growing availability of spoofing databases and rapid advances in using them, systems for detecting voice spoofing attacks are becoming more and more capable, and error rates close to zero are being reached for the ASVspoof2015 database. However, speech synthesis and voice conversion paradigms that are not considered in the ASVspoof2015 database are appearing. Such examples include direct waveform modelling and generative adversarial networks. We also need to investigate the feasibility of training spoofing systems using only low-quality found data. For that purpose, we developed a generative adversarial network-based speech enhancement system that improves the quality of speech data found in publicly available sources. Using the enhanced data, we trained state-of-the-art text-to-speech and voice conversion models and evaluated them in terms of perceptual speech quality and speaker similarity. The results show that the enhancement models significantly improved the SNR of low-quality degraded data found in publicly available sources and that they significantly improved the perceptual cleanliness of the source speech without significantly degrading the naturalness of the voice. However, the results also show limitations when generating speech with the low-quality found data.

📄 PDF Abstract BibTeX arXiv:1803.00860

Code (0)

등록된 구현이 없습니다.

Tasks

Generative Adversarial NetworkSpeech EnhancementSpeech Synthesistext-to-speechText to SpeechVoice Conversion

Similar Papers 제목 키워드 기반

ANTI-PHISHING IN ANDROID PHONE PROJECT REPORT.

2025-07-07 · Zenodo 2025 7 · Kamal Acharya

Phishing is a new word produced from 'fishing', it refers to the act that the attacker allure users to visit a faked Web site by sending them faked e-mails (or instant messages), and stealthily get victim's personal …

Turing Test for the Internet of Things

2014-12-11 · Neil Rubens

How smart is your kettle? How smart are things in your kitchen, your house, your neighborhood, on the internet? With the advent of Internet of Things, and the move of making devices `smart' by utilizing AI, a natural que…

Singing Voice Conversion with Disentangled Representations of Singer and Vocal Technique Using Variational Autoencoders

2019-12-03 · Yin-Jyun Luo, Chin-Chen Hsu, Kat Agres, Dorien Herremans

We propose a flexible framework that deals with both singer conversion and singers vocal technique conversion. The proposed model is trained on non-parallel corpora, accommodates many-to-many conversion, and leverages re…

DecoderVoice Conversion

Vocal Style Factorization for Effective Speaker Recognition in Affective Scenarios

2023-05-13 · Morgan Sandler, Arun Ross

The accuracy of automated speaker recognition is negatively impacted by change in emotions in a person's speech. In this paper, we hypothesize that speaker identity is composed of various vocal style factors that may be …

Speaker Recognition

How Google Search Works

2024-09-20 · Authorea 2024 10 · Kamal Acharya

Internet telephony consists of a combination of hardware and software that enables you to use the Internet as the transmission medium for telephone calls. For users who have free, or fixed-price Internet access, Internet…

Form