paper-with-me

Papers

Direct Speech-to-Speech Neural Machine Translation: A Survey

2024-11-13 · Mahendra Gupta, Maitreyee Dutta, Chandresh Kumar Maurya

Speech-to-Speech Translation (S2ST) models transform speech from one language to another target language with the same linguistic information. S2ST is important for bridging the communication gap among communities and has diverse applications. In recent years, researchers have introduced direct S2ST models, which have the potential to translate speech without relying on intermediate text generation, have better decoding latency, and the ability to preserve paralinguistic and non-linguistic features. However, direct S2ST has yet to achieve quality performance for seamless communication and still lags behind the cascade models in terms of performance, especially in real-world translation. To the best of our knowledge, no comprehensive survey is available on the direct S2ST system, which beginners and advanced researchers can look upon for a quick survey. The present work provides a comprehensive review of direct S2ST models, data and application issues, and performance metrics. We critically analyze the models' performance over the benchmark datasets and provide research challenges and future directions.

📄 PDF Abstract BibTeX arXiv:2411.14453

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSpeech-to-Speech TranslationSurveyText GenerationTranslation

Similar Papers 제목 키워드 기반

End-to-End Speech-to-Text Translation: A Survey

2023-12-02 · Nivedita Sethiya, Chandresh Kumar Maurya

Speech-to-text translation pertains to the task of converting speech signals in a language to text in another language. It finds its application in various domains, such as hands-free communication, dictation, video lect…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+5

A Survey of Voice Translation Methodologies - Acoustic Dialect Decoder

2016-10-13 · Hans Krupakar, Keerthika Rajvel, Bharathi B, Angel Deborah S 외

Speech Translation has always been about giving source text or audio input and waiting for system to give translated output in desired form. In this paper, we present the Acoustic Dialect Decoder (ADD) - a voice to voice…

DecoderSentenceSpeech SynthesisSurvey+1

Neural Speech Translation: From Neural Machine Translation to Direct Speech Translation

2022-06-01 · EAMT 2022 6 · Mattia Antonino Di Gangi
Machine TranslationTranslation

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

2026-04-13 · Bingzheng Qu, Kehai Chen, Xuefeng Bai, Min Zhang arxiv

Recent progress in multimodal large language models (MLLMs) is reshaping video translation from a cascaded pipeline of automatic speech recognition, machine translation, text-to-speech, and lip synchronization into a uni…

Multimodal ReasoningMachine TranslationSpeech Recognition

Cascaded Models With Cyclic Feedback For Direct Speech Translation

2020-10-21 · Tsz Kin Lam, Shigehiko Schamoni, Stefan Riezler

Direct speech translation describes a scenario where only speech inputs and corresponding translations are available. Such data are notoriously limited. We present a technique that allows cascades of automatic speech rec…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+2