Papers Voice Similarity
“Voice Similarity” 태그가 달린 논문 14편 · 필터 해제
AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
In this paper, we address the task of multimodal-to-speech generation, which aims to synthesize high-quality speech from multiple input modalities: text, video, and reference audio. This task has gained increasing attent…
In-Context LearningSpeech SynthesisVideo SynchronizationVoice SimilarityDMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis
Diffusion models have demonstrated significant potential in speech synthesis tasks, including text-to-speech (TTS) and voice cloning. However, their iterative denoising processes are computationally intensive, and previo…
DenoisingSpeaker VerificationSpeech Synthesistext-to-speech+3Disentangling segmental and prosodic factors to non-native speech comprehensibility
Current accent conversion (AC) systems do not disentangle the two main sources of non-native accent: segmental and prosodic characteristics. Being able to manipulate a non-native speaker's segmental and/or prosodic chann…
QuantizationVoice SimilarityVoxSim: A perceptual voice similarity dataset
This paper introduces VoxSim, a dataset of perceptual voice similarity ratings. Recent efforts to automate the assessment of speech synthesis technologies have primarily focused on predicting mean opinion score of natura…
BenchmarkingSpeaker RecognitionSpeech SynthesisVoice SimilaritySVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models
Representations from pre-trained speech foundation models (SFMs) have shown impressive performance in many downstream tasks. However, the potential benefits of incorporating pre-trained SFM representations into speaker v…
Voice ConversionVoice SimilaritySinger Identity Representation Learning using Self-Supervised Techniques
Significant strides have been made in creating voice identity representations using speech data. However, the same level of progress has not been achieved for singing voices. To bridge this gap, we suggest a framework fo…
Domain GeneralizationRepresentation LearningSelf-Supervised LearningSpeaker Verification+1YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone
YourTTS brings the power of a multilingual approach to the task of zero-shot multi-speaker TTS. Our method builds upon the VITS model and adds several novel modifications for zero-shot multi-speaker and multilingual trai…
Speech SynthesisText-To-Speech SynthesisVoice ConversionVoice Similarity+2SVSNet: An End-to-end Speaker Voice Similarity Assessment Model
Neural evaluation metrics derived for numerous speech generation tasks have recently attracted great attention. In this paper, we propose SVSNet, the first end-to-end neural network model to assess the speaker voice simi…
Voice ConversionVoice SimilarityDiffSVC: A Diffusion Probabilistic Model for Singing Voice Conversion
Singing voice conversion (SVC) is one promising technique which can enrich the way of human-computer interaction by endowing a computer the ability to produce high-fidelity and expressive singing voice. In this paper, we…
DenoisingVoice ConversionVoice SimilarityAn Adaptive Learning based Generative Adversarial Network for One-To-One Voice Conversion
Voice Conversion (VC) emerged as a significant domain of research in the field of speech synthesis in recent years due to its emerging application in voice-assisting technology, automated movie dubbing, and speech-to-sin…
Generative Adversarial NetworkSpeech SynthesisVoice ConversionVoice SimilarityPPG-based singing voice conversion with adversarial representation learning
Singing voice conversion (SVC) aims to convert the voice of one singer to that of other singers while keeping the singing content and melody. On top of recent voice conversion works, we propose a novel model to steadily …
Representation LearningVoice ConversionVoice SimilaritySpeech Pseudonymisation Assessment Using Voice Similarity Matrices
The proliferation of speech technologies and rising privacy legislation calls for the development of privacy preservation solutions for speech applications. These are essential since speech signals convey a wealth of ric…
De-identificationVoice SimilarityWaveform generation for text-to-speech synthesis using pitch-synchronous multi-scale generative adversarial networks
The state-of-the-art in text-to-speech synthesis has recently improved considerably due to novel neural waveform generation methods, such as WaveNet. However, these methods suffer from their slow sequential inference pro…
Image GenerationSpeech Synthesistext-to-speechText to Speech+2Sample Efficient Adaptive Text-to-Speech
We present a meta-learning approach for adaptive text-to-speech (TTS) with few data. During training, we learn a multi-speaker model using a shared conditional WaveNet core and independent learned embeddings for each spe…
Meta-Learningtext-to-speechText to SpeechVoice Similarity