paper-with-me

홈 › Papers

Prosody in Cascade and Direct Speech-to-Text Translation: a case study on Korean Wh-Phrases

2024-02-01 · Giulio Zhou, Tsz Kin Lam, Alexandra Birch, Barry Haddow

Speech-to-Text Translation (S2TT) has typically been addressed with cascade systems, where speech recognition systems generate a transcription that is subsequently passed to a translation model. While there has been a growing interest in developing direct speech translation systems to avoid propagating errors and losing non-verbal content, prior work in direct S2TT has struggled to conclusively establish the advantages of integrating the acoustic signal directly into the translation process. This work proposes using contrastive evaluation to quantitatively measure the ability of direct S2TT systems to disambiguate utterances where prosody plays a crucial role. Specifically, we evaluated Korean-English translation systems on a test set containing wh-phrases, for which prosodic features are necessary to produce translations with the correct intent, whether it's a statement, a yes/no question, a wh-question, and more. Our results clearly demonstrate the value of direct translation systems over cascade translation models, with a notable 12.9% improvement in overall accuracy in ambiguous cases, along with up to a 15.6% increase in F1 scores for one of the major intent categories. To the best of our knowledge, this work stands as the first to provide quantitative evidence that direct S2TT models can effectively leverage prosody. The code for our evaluation is openly accessible and freely available for review and utilisation.

📄 PDF Abstract BibTeX arXiv:2402.00632

Code (0)

등록된 구현이 없습니다.

Tasks

speech-recognitionSpeech RecognitionSpeech-to-TextSpeech-to-Text TranslationTranslation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?

2024-10-31 · Ioannis Tsiamas, Matthias Sperber, Andrew Finch, Sarthak Garg

The prosody of a spoken utterance, including features like stress, intonation and rhythm, can significantly affect the underlying semantics, and as a consequence can also affect its textual translation. Nevertheless, pro…

Rhythmspeech-recognitionSpeech RecognitionSpeech-to-Text+4

Direct Speech to Speech Translation: A Review

2025-03-03 · Mohammad Sarim, Saim Shakeel, Laeeba Javed, Jamaluddin 외

Speech to speech translation (S2ST) is a transformative technology that bridges global communication gaps, enabling real time multilingual interactions in diplomacy, tourism, and international trade. Our review examines …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+5

CrossVoice: Crosslingual Prosody Preserving Cascade-S2ST using Transfer Learning

2024-05-23 · Medha Hira, Arnav Goel, Anubha Gupta

This paper presents CrossVoice, a novel cascade-based Speech-to-Speech Translation (S2ST) system employing advanced ASR, MT, and TTS technologies with cross-lingual prosody preservation through transfer learning. We cond…

es-enfr-enSpeech-to-Speech TranslationTransfer Learning+1

Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation

2025-10-03 · Jacobo Romero-Díaz, Gerard I. Gállego, Oriol Pareras, Federico Costa 외 arxiv

Speech-to-Text Translation (S2TT) systems built from Automatic Speech Recognition (ASR) and Text-to-Text Translation (T2TT) modules face two major limitations: error propagation and the inability to exploit prosodic or o…

Speech-to-Text TranslationSpeech Recognition

A Holistic Cascade System, benchmark, and Human Evaluation Protocol for Expressive Speech-to-Speech Translation

2023-01-25 · Wen-Chin Huang, Benjamin Peloquin, Justine Kao, Changhan Wang 외

Expressive speech-to-speech translation (S2ST) aims to transfer prosodic attributes of source speech to target speech while maintaining translation accuracy. Existing research in expressive S2ST is limited, typically foc…

Speech-to-Speech TranslationTranslation