Using Phonemes in cascaded S2S translation pipeline
This paper explores the idea of using phonemes as a textual representation within a conventional multilingual simultaneous speech-to-speech translation pipeline, as opposed to the traditional reliance on text-based language representations. To investigate this, we trained an open-source sequence-to-sequence model on the WMT17 dataset in two formats: one using standard textual representation and the other employing phonemic representation. The performance of both approaches was assessed using the BLEU metric. Our findings shows that the phonemic approach provides comparable quality but offers several advantages, including lower resource requirements or better suitability for low-resource languages.
Code (1)
Tasks
Simultaneous Speech-to-Speech TranslationSpeech-to-Speech TranslationTranslationSimilar Papers 제목 키워드 기반
Improving Cascaded Unsupervised Speech Translation with Denoising Back-translation
Most of the speech translation models heavily rely on parallel data, which is hard to collect especially for low-resource languages. To tackle this issue, we propose to build a cascaded speech translation system without …
DenoisingMachine TranslationTranslationCUNI Neural ASR with Phoneme-Level Intermediate Step for\textasciitildeNon-Native\textasciitildeSLT at IWSLT 2020
In this paper, we present our submission to the Non-Native Speech Translation Task for IWSLT 2020. Our main contribution is a proposed speech recognition pipeline that consists of an acoustic model and a phoneme-to-graph…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1Mitigating Structural Noise in Low-Resource S2TT: An Optimized Cascaded Nepali-English Pipeline with Punctuation Restoration
Cascaded speech-to-text translation (S2TT) systems for low-resource languages can suffer from structural noise, particularly the loss of punctuation during the Automatic Speech Recognition (ASR) phase. This research inve…
Speech-to-Text TranslationSpeech RecognitionOmniFusion: Simultaneous Multilingual Multimodal Translations via Modular Fusion
There has been significant progress in open-source text-only translation large language models (LLMs) with better language coverage and quality. However, these models can be only used in cascaded pipelines for speech tra…
Speech RecognitionImproving Isochronous Machine Translation with Target Factors and Auxiliary Counters
To translate speech for automatic dubbing, machine translation needs to be isochronous, i.e. translated speech needs to be aligned with the source in terms of speech durations. We introduce target factors in a transforme…
DecoderMachine TranslationTranslation