End-to-end Speech Translation with Spoken-to-Written Style Conversion
End-to-end speech translation (ST), which translates speech in source language directly into text in target language by a single model, has attracted a great deal of attention in recent years. Compared to the cascade ST, it has the advantages of easier deployment, better efficiency, and less error propagation. Meanwhile, spoken-to-written style conversion has been proved to be able to improve cascaded ST by reducing the gap between the language style of speech transcription and bilingual corpora used for machine translation training. Therefore, it is desirable to integrate the conversion into end-to-end ST. In this paper, we propose a joint task of speech-to-written-style-text conversion and end-to-end ST, as well as an interactive-attention-based multi-decoder model for the joint task to improve end-to-end ST. Experiments on a Japanese-English lecture ST dataset and CoVoST 2 Native Japanese show that our models outperform a strong baseline on Japanese-English ST.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderMachine TranslationTranslationSimilar Papers 제목 키워드 기반
Parallel Corpus for Japanese Spoken-to-Written Style Conversion
With the increase of automatic speech recognition (ASR) applications, spoken-to-written style conversion that transforms spoken-style text into written-style text is becoming an important technology to increase the reada…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Punctuation Restorationspeech-recognition+1Recording for Eyes, Not Echoing to Ears: Contextualized Spoken-to-Written Conversion of ASR Transcripts
Automatic Speech Recognition (ASR) transcripts exhibit recognition errors and various spoken language phenomena such as disfluencies, ungrammatical sentences, and incomplete sentences, hence suffering from poor readabili…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)In-Context Learningspeech-recognition+1End-to-End Speech-to-Text Translation: A Survey
Speech-to-text translation pertains to the task of converting speech signals in a language to text in another language. It finds its application in various domains, such as hands-free communication, dictation, video lect…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+5Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization
SpeechLLMs typically combine ASR-trained encoders with text-based LLM backbones, leading them to inherit written-style output patterns unsuitable for text-to-speech synthesis. This mismatch is particularly pronounced in …
Text-To-Speech SynthesisZero-Shot Joint Modeling of Multiple Spoken-Text-Style Conversion Tasks using Switching Tokens
In this paper, we propose a novel spoken-text-style conversion method that can simultaneously execute multiple style conversion modules such as punctuation restoration and disfluency deletion without preparing matched da…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Punctuation Restorationspeech-recognition+2