paper-with-me

Papers Voice Similarity

“Voice Similarity” 태그가 달린 논문 14편 · 필터 해제

AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation

2025-04-29 · Jeongsoo Choi, Ji-Hoon Kim, Kim Sung-Bin, Tae-Hyun Oh 외

In this paper, we address the task of multimodal-to-speech generation, which aims to synthesize high-quality speech from multiple input modalities: text, video, and reference audio. This task has gained increasing attent…

In-Context LearningSpeech SynthesisVideo SynchronizationVoice Similarity

DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis

2024-10-14 · Yingahao Aaron Li, Rithesh Kumar, Zeyu Jin

Diffusion models have demonstrated significant potential in speech synthesis tasks, including text-to-speech (TTS) and voice cloning. However, their iterative denoising processes are computationally intensive, and previo…

DenoisingSpeaker VerificationSpeech Synthesistext-to-speech+3

Disentangling segmental and prosodic factors to non-native speech comprehensibility

2024-08-20 · Waris Quamer, Ricardo Gutierrez-Osuna

Current accent conversion (AC) systems do not disentangle the two main sources of non-native accent: segmental and prosodic characteristics. Being able to manipulate a non-native speaker's segmental and/or prosodic chann…

QuantizationVoice Similarity

VoxSim: A perceptual voice similarity dataset

2024-07-26 · Junseok Ahn, Youkyum Kim, Yeunju Choi, Doyeop Kwak 외

This paper introduces VoxSim, a dataset of perceptual voice similarity ratings. Recent efforts to automate the assessment of speech synthesis technologies have primarily focused on predicting mean opinion score of natura…

BenchmarkingSpeaker RecognitionSpeech SynthesisVoice Similarity

SVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models

2024-06-12 · Chun Yin, Tai-Shih Chi, Yu Tsao, Hsin-Min Wang

Representations from pre-trained speech foundation models (SFMs) have shown impressive performance in many downstream tasks. However, the potential benefits of incorporating pre-trained SFM representations into speaker v…

Voice ConversionVoice Similarity

Singer Identity Representation Learning using Self-Supervised Techniques

2024-01-10 · International Society of Music Information Retrieval 2023 8 · Bernardo Torres, Stefan Lattner, Gaël Richard

Significant strides have been made in creating voice identity representations using speech data. However, the same level of progress has not been achieved for singing voices. To bridge this gap, we suggest a framework fo…

Domain GeneralizationRepresentation LearningSelf-Supervised LearningSpeaker Verification+1

YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone

2021-12-04 · Edresson Casanova, Julian Weber, Christopher Shulby, Arnaldo Candido Junior 외

YourTTS brings the power of a multilingual approach to the task of zero-shot multi-speaker TTS. Our method builds upon the VITS model and adds several novel modifications for zero-shot multi-speaker and multilingual trai…

Speech SynthesisText-To-Speech SynthesisVoice ConversionVoice Similarity+2

SVSNet: An End-to-end Speaker Voice Similarity Assessment Model

2021-07-20 · Cheng-Hung Hu, Yu-Huai Peng, Junichi Yamagishi, Yu Tsao 외

Neural evaluation metrics derived for numerous speech generation tasks have recently attracted great attention. In this paper, we propose SVSNet, the first end-to-end neural network model to assess the speaker voice simi…

Voice ConversionVoice Similarity

DiffSVC: A Diffusion Probabilistic Model for Singing Voice Conversion

2021-05-28 · Songxiang Liu, Yuewen Cao, Dan Su, Helen Meng

Singing voice conversion (SVC) is one promising technique which can enrich the way of human-computer interaction by endowing a computer the ability to produce high-fidelity and expressive singing voice. In this paper, we…

DenoisingVoice ConversionVoice Similarity

An Adaptive Learning based Generative Adversarial Network for One-To-One Voice Conversion

2021-04-25 · Sandipan Dhar, Nanda Dulal Jana, Swagatam Das

Voice Conversion (VC) emerged as a significant domain of research in the field of speech synthesis in recent years due to its emerging application in voice-assisting technology, automated movie dubbing, and speech-to-sin…

Generative Adversarial NetworkSpeech SynthesisVoice ConversionVoice Similarity

PPG-based singing voice conversion with adversarial representation learning

2020-10-28 · Zhonghao Li, Benlai Tang, Xiang Yin, Yuan Wan 외

Singing voice conversion (SVC) aims to convert the voice of one singer to that of other singers while keeping the singing content and melody. On top of recent voice conversion works, we propose a novel model to steadily …

Representation LearningVoice ConversionVoice Similarity

Speech Pseudonymisation Assessment Using Voice Similarity Matrices

2020-08-30 · Paul-Gauthier Noé, Jean-François Bonastre, Driss Matrouf, Natalia Tomashenko 외

The proliferation of speech technologies and rising privacy legislation calls for the development of privacy preservation solutions for speech applications. These are essential since speech signals convey a wealth of ric…

De-identificationVoice Similarity

Waveform generation for text-to-speech synthesis using pitch-synchronous multi-scale generative adversarial networks

2018-10-30 · Lauri Juvela, Bajibabu Bollepalli, Junichi Yamagishi, Paavo Alku

The state-of-the-art in text-to-speech synthesis has recently improved considerably due to novel neural waveform generation methods, such as WaveNet. However, these methods suffer from their slow sequential inference pro…

Image GenerationSpeech Synthesistext-to-speechText to Speech+2

Sample Efficient Adaptive Text-to-Speech

2018-09-27 · ICLR 2019 5 · Yutian Chen, Yannis Assael, Brendan Shillingford, David Budden 외

We present a meta-learning approach for adaptive text-to-speech (TTS) with few data. During training, we learn a multi-speaker model using a shared conditional WaveNet core and independent learned embeddings for each spe…

Meta-Learningtext-to-speechText to SpeechVoice Similarity
1–14 / 14