paper-with-me

홈 › Papers

End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs

2025-10-11 · Nam Luu, Ondřej Bojar arxiv

Speech Translation (ST) is a machine translation task that involves converting speech signals from one language to the corresponding text in another language; this task has two different approaches, namely the traditional cascade and the more recent end-to-end. This paper explores a combined end-to-end architecture of pre-trained speech encoders and Large Language Models (LLMs) for performing both Automatic Speech Recognition (ASR) and ST simultaneously. Experiments with the English-to-German language pair show that our best model not only can achieve better translation results than SeamlessM4T, a large foundational end-to-end, multi-modal translation model, but can also match the performance of a cascaded system with Whisper and NLLB, with up to a score gain of 8% in $\text{COMET}^{\text{DA}}_{22}$ metric.

📄 PDF Abstract BibTeX arXiv:2510.10329

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSpeech Recognition

Similar Papers 제목 키워드 기반

Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs?

2026-06-24 · Tomoya Mizumoto, Yusuke Fujita arxiv

Connecting a pre-trained speech encoder to a Large Language Model (LLM) is the standard architecture for building Speech LLMs. However, a structural misalignment exists between the encoder and the LLM. Unlike encoders ba…

Speech Recognition

ESPnet-SE++: Speech Enhancement for Robust Speech Recognition, Translation, and Understanding

2022-07-19 · Yen-Ju Lu, Xuankai Chang, Chenda Li, Wangyou Zhang 외

This paper presents recent progress on integrating speech separation and enhancement (SSE) into the ESPnet toolkit. Compared with the previous ESPnet-SE work, numerous features have been added, including recent state-of-…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Robust Speech RecognitionSpeech Enhancement+4

Enhancements in statistical spoken language translation by de-normalization of ASR results

2015-11-18 · Agnieszka Wołk, Krzysztof Wołk, Krzysztof Marasek

Spoken language translation (SLT) has become very important in an increasingly globalized world. Machine translation (MT) for automatic speech recognition (ASR) systems is a major challenge of great interest. This resear…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationSegmentation+5

Overcoming Latency Bottlenecks in On-Device Speech Translation: A Cascaded Approach with Alignment-Based Streaming MT

2025-08-18 · Zeeshan Ahmed, Frank Seide, Niko Moritz, Ju Lin 외 arxiv

This paper tackles several challenges that arise when integrating Automatic Speech Recognition (ASR) and Machine Translation (MT) for real-time, on-device streaming speech translation. Although state-of-the-art ASR syste…

Machine TranslationSpeech Recognition

LibriVoxDeEn: A Corpus for German-to-English Speech Translation and German Speech Recognition

2019-10-17 · LREC 2020 5 · Benjamin Beilharz, Xin Sun, Sariya Karimova, Stefan Riezler

We present a corpus of sentence-aligned triples of German audio, German text, and English translation, based on German audiobooks. The speech translation data consist of 110 hours of audio material aligned to over 50k pa…

Sentencespeech-recognitionSpeech RecognitionTranslation