paper-with-me

Papers

TriAAN-VC: Triple Adaptive Attention Normalization for Any-to-Any Voice Conversion

2023-03-16 · Hyun Joon Park, Seok Woo Yang, Jin Sob Kim, WooSeok Shin, Sung Won Han

Voice Conversion (VC) must be achieved while maintaining the content of the source speech and representing the characteristics of the target speaker. The existing methods do not simultaneously satisfy the above two aspects of VC, and their conversion outputs suffer from a trade-off problem between maintaining source contents and target characteristics. In this study, we propose Triple Adaptive Attention Normalization VC (TriAAN-VC), comprising an encoder-decoder and an attention-based adaptive normalization block, that can be applied to non-parallel any-to-any VC. The proposed adaptive normalization block extracts target speaker representations and achieves conversion while minimizing the loss of the source content with siamese loss. We evaluated TriAAN-VC on the VCTK dataset in terms of the maintenance of the source content and target speaker similarity. Experimental results for one-shot VC suggest that TriAAN-VC achieves state-of-the-art performance while mitigating the trade-off problem encountered in the existing VC methods.

📄 PDF Abstract BibTeX arXiv:2303.09057

Code (1)

winddori2002/TriAAN-VC 공식 구현 pytorch

Tasks

DecoderVoice Conversion

Similar Papers 제목 키워드 기반

DiVA-DocRE: A Discriminative and Voice-Aware Paradigm for Document-Level Relation Extraction

2024-09-07 · YiHeng Wu, Roman Yangarber, Xian Mao

The remarkable capabilities of Large Language Models (LLMs) in text comprehension and generation have revolutionized Information Extraction (IE). One such advancement is in Document-level Relation Triplet Extraction (Doc…

Document-level Relation ExtractionReading ComprehensionRelationRelation Extraction+2

AdaSpeech 4: Adaptive Text to Speech in Zero-Shot Scenarios

2022-04-01 · Yihan Wu, Xu Tan, Bohan Li, Lei He 외

Adaptive text to speech (TTS) can synthesize new voices in zero-shot scenarios efficiently, by using a well-trained source TTS model without adapting it on the speech data of new speakers. Considering seen and unseen spe…

Speech Synthesistext-to-speechText to Speech

Speech Enhancement using Separable Polling Attention and Global Layer Normalization followed with PReLU

2021-05-06 · Dengfeng Ke, Jinsong Zhang, Yanlu Xie, Yanyan Xu 외

Single channel speech enhancement is a challenging task in speech community. Recently, various neural networks based methods have been applied to speech enhancement. Among these models, PHASEN and T-GSA achieve state-of-…

Speech Enhancement

Triplet-Trained Vector Space and Sieve-Based Search Improve Biomedical Concept Normalization

2021-06-01 · NAACL (BioNLP) 2021 6 · Dongfang Xu, Steven Bethard

Concept normalization, the task of linking textual mentions of concepts to concepts in an ontology, is critical for mining and analyzing biomedical texts. We propose a vector-space model for concept normalization, where …

Triplet

TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control

2024-09-24 · Yu Zhang, Ziyue Jiang, RuiQi Li, Changhao Pan 외

Zero-shot singing voice synthesis (SVS) with style transfer and style control aims to generate high-quality singing voices with unseen timbres and styles (including singing method, emotion, rhythm, technique, and pronunc…

ClusteringLanguage ModellingQuantizationRhythm+2