paper-with-me

Papers

AC-VC: Non-parallel Low Latency Phonetic Posteriorgrams Based Voice Conversion

2021-11-12 · Damien Ronssin, Milos Cernak

This paper presents AC-VC (Almost Causal Voice Conversion), a phonetic posteriorgrams based voice conversion system that can perform any-to-many voice conversion while having only 57.5 ms future look-ahead. The complete system is composed of three neural networks trained separately with non-parallel data. While most of the current voice conversion systems focus primarily on quality irrespective of algorithmic latency, this work elaborates on designing a method using a minimal amount of future context thus allowing a future real-time implementation. According to a subjective listening test organized in this work, the proposed AC-VC system achieves parity with the non-causal ASR-TTS baseline of the Voice Conversion Challenge 2020 in naturalness with a MOS of 3.5. In contrast, the results indicate that missing future context impacts speaker similarity. Obtained similarity percentage of 65% is lower than the similarity of current best voice conversion systems.

📄 PDF Abstract BibTeX arXiv:2111.06601

Code (0)

등록된 구현이 없습니다.

Tasks

Voice Conversion

Similar Papers 제목 키워드 기반

ALO-VC: Any-to-any Low-latency One-shot Voice Conversion

2023-06-01 · Bohan Wang, Damien Ronssin, Milos Cernak

This paper presents ALO-VC, a non-parallel low-latency one-shot phonetic posteriorgrams (PPGs) based voice conversion method. ALO-VC enables any-to-any voice conversion using only one utterance from the target speaker, w…

CPUVoice Conversion

V2S attack: building DNN-based voice conversion from automatic speaker verification

2019-08-05 · Taiki Nakamura, Yuki Saito, Shinnosuke Takamichi, Yusuke Ijima 외

This paper presents a new voice impersonation attack using voice conversion (VC). Enrolling personal voices for automatic speaker verification (ASV) offers natural and flexible biometric authentication systems. Basically…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Verificationspeech-recognition+2

Investigating self-supervised features for expressive, multilingual voice conversion

2025-05-13 · Álvaro Martín-Cortinas, Daniel Sáez-Trigueros, Grzegorz Beringer, Iván Vallés-Pérez 외

Voice conversion (VC) systems are widely used for several applications, from speaker anonymisation to personalised speech synthesis. Supervised approaches learn a mapping between different speakers using parallel data, w…

Self-Supervised LearningSpeech SynthesisVoice Conversion

Investigating Inter- and Intra-speaker Voice Conversion using Audiobooks

2022-06-01 · LREC 2022 6 · Aghilas Sini, Damien Lolive, Nelly Barbot, Pierre Alain

Audiobook readers play with their voices to emphasize some text passages, highlight discourse changes or significant events, or in order to make listening easier and entertaining. A dialog is a central passage in audiobo…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis+1

Taco-VC: A Single Speaker Tacotron based Voice Conversion with Limited Data

2019-04-06 · Roee Levy Leshem, Raja Giryes

This paper introduces Taco-VC, a novel architecture for voice conversion based on Tacotron synthesizer, which is a sequence-to-sequence with attention model. The training of multi-speaker voice conversion systems require…

Phoneme RecognitionSpeech EnhancementVoice Conversion