paper-with-me

홈 › Papers

Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?

2025-10-03 · Oriol Pareras, Gerard I. Gállego, Federico Costa, Cristina España-Bonet, Javier Hernando arxiv

Recent work on Speech-to-Text Translation (S2TT) has focused on LLM-based models, introducing the increasingly adopted Chain-of-Thought (CoT) prompting, where the model is guided to first transcribe the speech and then translate it. CoT typically outperforms direct prompting primarily because it can exploit abundant Automatic Speech Recognition (ASR) and Text-to-Text Translation (T2TT) datasets to explicitly model its steps. In this paper, we systematically compare CoT and Direct prompting under increasing amounts of S2TT data. To this end, we pseudo-label an ASR corpus by translating its transcriptions into six European languages, and train LLM-based S2TT systems with both prompting strategies at different data scales. Our results show that Direct improves more consistently as the amount of data increases, suggesting that it may become a more effective approach as larger S2TT resources are created.

📄 PDF Abstract BibTeX arXiv:2510.03093

Code (0)

등록된 구현이 없습니다.

Tasks

Speech-to-Text TranslationSpeech Recognition

Similar Papers 제목 키워드 기반

Revisiting End-to-End Speech-to-Text Translation From Scratch

2022-06-09 · Biao Zhang, Barry Haddow, Rico Sennrich

End-to-end (E2E) speech-to-text translation (ST) often depends on pretraining its encoder and/or decoder using source transcripts via speech recognition or text translation tasks, without which translation performance dr…

Decoderspeech-recognitionSpeech RecognitionSpeech-to-Text+2

Direct speech-to-speech translation with discrete units

2021-07-12 · ACL 2022 5 · Ann Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu 외

We present a direct speech-to-speech translation (S2ST) model that translates speech from one language to speech in another language without relying on intermediate text generation. We tackle the problem by first applyin…

Speech-to-Speech TranslationText GenerationTranslation

Direct Speech-to-speech Translation without Textual Annotation using Bottleneck Features

2022-12-12 · Junhui Zhang, Junjie Pan, Xiang Yin, Zejun Ma

Speech-to-speech translation directly translates a speech utterance to another between different languages, and has great potential in tasks such as simultaneous interpretation. State-of-art models usually contains an au…

Speech-to-Speech TranslationTranslation

Joint Pre-Training with Speech and Bilingual Text for Direct Speech to Speech Translation

2022-10-31 · Kun Wei, Long Zhou, Ziqiang Zhang, Liping Chen 외

Direct speech-to-speech translation (S2ST) is an attractive research topic with many advantages compared to cascaded S2ST. However, direct S2ST suffers from the data scarcity problem because the corpora from speech of th…

Speech-to-Speech TranslationTranslation

Automatic Labelling of Speech Translation Errors

2026-06-04 · Dominik Macháček, Maike Züfle, Ondrej Klejch arxiv

Errors in speech translations reduce trustworthiness of Speech Translation (ST) systems and can have serious consequences. Yet currently there is no established methodology for evaluating confidence and quality estimatio…