paper-with-me

홈 › Papers

Direct Speech to Speech Translation: A Review

2025-03-03 · Mohammad Sarim, Saim Shakeel, Laeeba Javed, Jamaluddin, Mohammad Nadeem

Speech to speech translation (S2ST) is a transformative technology that bridges global communication gaps, enabling real time multilingual interactions in diplomacy, tourism, and international trade. Our review examines the evolution of S2ST, comparing traditional cascade models which rely on automatic speech recognition (ASR), machine translation (MT), and text to speech (TTS) components with newer end to end and direct speech translation (DST) models that bypass intermediate text representations. While cascade models offer modularity and optimized components, they suffer from error propagation, increased latency, and loss of prosody. In contrast, direct S2ST models retain speaker identity, reduce latency, and improve translation naturalness by preserving vocal characteristics and prosody. However, they remain limited by data sparsity, high computational costs, and generalization challenges for low-resource languages. The current work critically evaluates these approaches, their tradeoffs, and future directions for improving real time multilingual communication.

📄 PDF Abstract BibTeX arXiv:2503.04799

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognitionSpeech RecognitionSpeech-to-Speech Translationtext-to-speechText to SpeechTranslation

Similar Papers 제목 키워드 기반

Speech to Speech Translation with Translatotron: A State of the Art Review

2025-02-09 · Jules R. Kala, Emmanuel Adetiba, Abdultaofeek Abayom, Oluwatobi E. Dare 외

A cascade-based speech-to-speech translation has been considered a benchmark for a very long time, but it is plagued by many issues, like the time taken to translate a speech from one language to another and compound err…

speech-recognitionSpeech RecognitionSpeech-to-Speech TranslationSpeech-to-Text+5

End-to-End Speech-to-Text Translation: A Survey

2023-12-02 · Nivedita Sethiya, Chandresh Kumar Maurya

Speech-to-text translation pertains to the task of converting speech signals in a language to text in another language. It finds its application in various domains, such as hands-free communication, dictation, video lect…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+5

Does Joint Training Really Help Cascaded Speech Translation?

2022-10-24 · Viet Anh Khoa Tran, David Thulke, Yingbo Gao, Christian Herold 외

Currently, in speech translation, the straightforward approach - cascading a recognition system with a translation system - delivers state-of-the-art results. However, fundamental challenges such as error propagation fro…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Direct Speech-to-Speech Neural Machine Translation: A Survey

2024-11-13 · Mahendra Gupta, Maitreyee Dutta, Chandresh Kumar Maurya

Speech-to-Speech Translation (S2ST) models transform speech from one language to another target language with the same linguistic information. S2ST is important for bridging the communication gap among communities and ha…

Machine TranslationSpeech-to-Speech TranslationSurveyText Generation+1

Prosody in Cascade and Direct Speech-to-Text Translation: a case study on Korean Wh-Phrases

2024-02-01 · Giulio Zhou, Tsz Kin Lam, Alexandra Birch, Barry Haddow

Speech-to-Text Translation (S2TT) has typically been addressed with cascade systems, where speech recognition systems generate a transcription that is subsequently passed to a translation model. While there has been a gr…

speech-recognitionSpeech RecognitionSpeech-to-TextSpeech-to-Text Translation+1