paper-with-me

홈 › Papers

CrossVoice: Crosslingual Prosody Preserving Cascade-S2ST using Transfer Learning

2024-05-23 · Medha Hira, Arnav Goel, Anubha Gupta

This paper presents CrossVoice, a novel cascade-based Speech-to-Speech Translation (S2ST) system employing advanced ASR, MT, and TTS technologies with cross-lingual prosody preservation through transfer learning. We conducted comprehensive experiments comparing CrossVoice with direct-S2ST systems, showing improved BLEU scores on tasks such as Fisher Es-En, VoxPopuli Fr-En and prosody preservation on benchmark datasets CVSS-T and IndicTTS. With an average mean opinion score of 3.75 out of 4, speech synthesized by CrossVoice closely rivals human speech on the benchmark, highlighting the efficacy of cascade-based systems and transfer learning in multilingual S2ST with prosody transfer.

📄 PDF Abstract BibTeX arXiv:2406.00021

Code (0)

등록된 구현이 없습니다.

Tasks

es-enfr-enSpeech-to-Speech TranslationTransfer LearningTranslation

Similar Papers 제목 키워드 기반

A unified one-shot prosody and speaker conversion system with self-supervised discrete speech units

2022-11-12 · Li-Wei Chen, Shinji Watanabe, Alexander Rudnicky

We present a unified system to realize one-shot voice conversion (VC) on the pitch, rhythm, and speaker attributes. Existing works generally ignore the correlation between prosody and language content, leading to the deg…

RhythmVoice Conversion

Expressive Machine Dubbing Through Phrase-level Cross-lingual Prosody Transfer

2023-06-20 · Jakub Swiatkowski, Duo Wang, Mikolaj Babianski, Giuseppe Coccia 외

Speech generation for machine dubbing adds complexity to conventional Text-To-Speech solutions as the generated output is required to match the expressiveness, emotion and speaking rate of the source content. Capturing a…

text-to-speechText to Speech

Direct Speech to Speech Translation: A Review

2025-03-03 · Mohammad Sarim, Saim Shakeel, Laeeba Javed, Jamaluddin 외

Speech to speech translation (S2ST) is a transformative technology that bridges global communication gaps, enabling real time multilingual interactions in diplomacy, tourism, and international trade. Our review examines …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+5

A Neural TTS System with Parallel Prosody Transfer from Unseen Speakers

2023-09-20 · Slava Shechtman, Raul Fernandez

Modern neural TTS systems are capable of generating natural and expressive speech when provided with sufficient amounts of training data. Such systems can be equipped with prosody-control functionality, allowing for more…

CopyCat: Many-to-Many Fine-Grained Prosody Transfer for Neural Text-to-Speech

2020-04-30

Prosody Transfer (PT) is a technique that aims to use the prosody from a source audio as a reference while synthesising speech. Fine-grained PT aims at capturing prosodic aspects like rhythm, emphasis, melody, duration, …

Rhythmtext-to-speechText to Speech