paper-with-me

홈 › Papers

Using Phonemes in cascaded S2S translation pipeline

2025-04-22 · Rene Pilz, Johannes Schneider

This paper explores the idea of using phonemes as a textual representation within a conventional multilingual simultaneous speech-to-speech translation pipeline, as opposed to the traditional reliance on text-based language representations. To investigate this, we trained an open-source sequence-to-sequence model on the WMT17 dataset in two formats: one using standard textual representation and the other employing phonemic representation. The performance of both approaches was assessed using the BLEU metric. Our findings shows that the phonemic approach provides comparable quality but offers several advantages, including lower resource requirements or better suitability for low-resource languages.

📄 PDF Abstract BibTeX arXiv:2504.16234

Code (1)

fungus75/phonemes_s2s_pipeline 공식 구현

Tasks

Simultaneous Speech-to-Speech TranslationSpeech-to-Speech TranslationTranslation

Similar Papers 제목 키워드 기반

Improving Cascaded Unsupervised Speech Translation with Denoising Back-translation

2023-05-12 · Yu-Kuan Fu, Liang-Hsuan Tseng, Jiatong Shi, Chen-An Li 외

Most of the speech translation models heavily rely on parallel data, which is hard to collect especially for low-resource languages. To tackle this issue, we propose to build a cascaded speech translation system without …

DenoisingMachine TranslationTranslation

CUNI Neural ASR with Phoneme-Level Intermediate Step for\textasciitildeNon-Native\textasciitildeSLT at IWSLT 2020

2020-07-01 · WS 2020 7 · Peter Pol{\'a}k, Sangeet Sagar, Dominik Mach{\'a}{\v{c}}ek, Ond{\v{r}}ej Bojar

In this paper, we present our submission to the Non-Native Speech Translation Task for IWSLT 2020. Our main contribution is a proposed speech recognition pipeline that consists of an acoustic model and a phoneme-to-graph…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Mitigating Structural Noise in Low-Resource S2TT: An Optimized Cascaded Nepali-English Pipeline with Punctuation Restoration

2026-02-25 · Tangsang Chongbang, Pranesh Pyara Shrestha, Amrit Sarki, Anku Jaiswal arxiv

Cascaded speech-to-text translation (S2TT) systems for low-resource languages can suffer from structural noise, particularly the loss of punctuation during the Automatic Speech Recognition (ASR) phase. This research inve…

Speech-to-Text TranslationSpeech Recognition

OmniFusion: Simultaneous Multilingual Multimodal Translations via Modular Fusion

2025-11-28 · Sai Koneru, Matthias Huck, Jan Niehues arxiv

There has been significant progress in open-source text-only translation large language models (LLMs) with better language coverage and quality. However, these models can be only used in cascaded pipelines for speech tra…

Speech Recognition

Improving Isochronous Machine Translation with Target Factors and Auxiliary Counters

2023-05-22 · Proyag Pal, Brian Thompson, Yogesh Virkar, Prashant Mathur 외

To translate speech for automatic dubbing, machine translation needs to be isochronous, i.e. translated speech needs to be aligned with the source in terms of speech durations. We introduce target factors in a transforme…

DecoderMachine TranslationTranslation