Assem-VC: Realistic Voice Conversion by Assembling Modern Speech Synthesis Techniques
Recent works on voice conversion (VC) focus on preserving the rhythm and the intonation as well as the linguistic content. To preserve these features from the source, we decompose current non-parallel VC systems into two encoders and one decoder. We analyze each module with several experiments and reassemble the best components to propose Assem-VC, a new state-of-the-art any-to-many non-parallel VC system. We also examine that PPG and Cotatron features are speaker-dependent, and attempt to remove speaker identity with adversarial training. Code and audio samples are available at https://github.com/mindslab-ai/assem-vc.
Code (1)
Tasks
DecoderRhythmSpeech SynthesisVoice ConversionSimilar Papers 제목 키워드 기반
Voice Conversion for Stuttered Speech, Instruments, Unseen Languages and Textually Described Voices
Voice conversion aims to convert source speech into a target voice using recordings of the target speaker as a reference. Newer models are producing increasingly realistic output. But what happens when models are fed wit…
Voice ConversionHiFi-VC: High Quality ASR-Based Voice Conversion
The goal of voice conversion (VC) is to convert input voice to match the target speaker's voice while keeping text and prosody intact. VC is usually used in entertainment and speaking-aid systems, as well as applied for …
speech-recognitionSpeech RecognitionVocal Bursts Intensity PredictionVoice ConversionDefense Against Synthetic Speech: Real-Time Detection of RVC Voice Conversion Attacks
Generative audio technologies now enable highly realistic voice cloning and real-time voice conversion, increasing the risk of impersonation, fraud, and misinformation in communication channels such as phone and video ca…
Voice ConversionArraMon: A Joint Navigation-Assembly Instruction Interpretation Task in Dynamic Environments
For embodied agents, navigation is an important ability but not an isolated goal. Agents are also expected to perform specific tasks after reaching the target location, such as picking up objects and assembling them into…
Referring ExpressionReferring Expression ComprehensionVision and Language NavigationDisassembling Object Representations without Labels
In this paper, we study a new representation-learning task, which we termed as disassembling object representations. Given an image featuring multiple objects, the goal of disassembling is to acquire a latent representat…
General ClassificationGenerative Adversarial NetworkObjectRepresentation Learning+1