paper-with-me

홈 › Papers

Hiding speaker's sex in speech using zero-evidence speaker representation in an analysis/synthesis pipeline

2022-11-29 · Paul-Gauthier Noé, Xiaoxiao Miao, Xin Wang, Junichi Yamagishi, Jean-François Bonastre, Driss Matrouf

The use of modern vocoders in an analysis/synthesis pipeline allows us to investigate high-quality voice conversion that can be used for privacy purposes. Here, we propose to transform the speaker embedding and the pitch in order to hide the sex of the speaker. ECAPA-TDNN-based speaker representation fed into a HiFiGAN vocoder is protected using a neural-discriminant analysis approach, which is consistent with the zero-evidence concept of privacy. This approach significantly reduces the information in speech related to the speaker's sex while preserving speech content and some consistency in the resulting protected voices.

📄 PDF Abstract BibTeX arXiv:2211.16065

Code (1)

nii-yamagishilab/speaker_sex_attribute_privacy 공식 구현 pytorch

Tasks

Voice Conversion

Methods 이 논문이 사용한 방법론

HiFi-GAN HiFi-GAN is a generative adversarial network for speech synthesis. HiFi-GAN consists of one generator and two discriminators: multi-scale and multi-period discriminators. The…

Similar Papers 제목 키워드 기반

Enhancing Zero-Shot Many to Many Voice Conversion with Self-Attention VAE

2022-03-30 · Ziang Long, Yunling Zheng, Meng Yu, Jack Xin

Variational auto-encoder (VAE) is an effective neural network architecture to disentangle a speech utterance into speaker identity and linguistic content latent embeddings, then generate an utterance for a target speaker…

DecoderSentenceVoice Conversion

PersonaGesture: Single-Reference Co-Speech Gesture Personalization for Unseen Speakers

2026-05-07 · Xiangyue Zhang, Yiyi Cai, Kunhang Li, Kaixing Yang 외 arxiv

We propose PersonaGesture, a diffusion-based pipeline for single-reference co-speech gesture personalization of unseen speakers. Given target speech and one motion clip from a new speaker, the model must synthesize gestu…

Anonymizing Speech with Generative Adversarial Networks to Preserve Speaker Privacy

2022-10-13 · Sarina Meyer, Pascal Tilli, Pavel Denisov, Florian Lux 외

In order to protect the privacy of speech data, speaker anonymization aims for hiding the identity of a speaker by changing the voice in speech recordings. This typically comes with a privacy-utility trade-off between pr…

Generative Adversarial NetworkSpeaker anonymizationSpeech-to-Texttext-to-speech+1

Multi-level Temporal-channel Speaker Retrieval for Zero-shot Voice Conversion

2023-05-12 · Zhichao Wang, Liumeng Xue, Qiuqiang Kong, Lei Xie 외

Zero-shot voice conversion (VC) converts source speech into the voice of any desired speaker using only one utterance of the speaker without requiring additional model updates. Typical methods use a speaker representatio…

DisentanglementRetrievalSpeaker VerificationVoice Conversion

SNAC: Speaker-normalized affine coupling layer in flow-based architecture for zero-shot multi-speaker text-to-speech

2022-11-30 · Byoung Jin Choi, Myeonghun Jeong, Joun Yeop Lee, Nam Soo Kim

Zero-shot multi-speaker text-to-speech (ZSM-TTS) models aim to generate a speech sample with the voice characteristic of an unseen speaker. The main challenge of ZSM-TTS is to increase the overall speaker similarity for …

Speech Synthesistext-to-speechText to Speech