Improving Stability in Simultaneous Speech Translation: A Revision-Controllable Decoding Approach
Simultaneous Speech-to-Text translation serves a critical role in real-time crosslingual communication. Despite the advancements in recent years, challenges remain in achieving stability in the translation process, a concern primarily manifested in the flickering of partial results. In this paper, we propose a novel revision-controllable method designed to address this issue. Our method introduces an allowed revision window within the beam search pruning process to screen out candidate translations likely to cause extensive revisions, leading to a substantial reduction in flickering and, crucially, providing the capability to completely eliminate flickering. The experiments demonstrate the proposed method can significantly improve the decoding stability without compromising substantially on the translation quality.
Code (0)
등록된 구현이 없습니다.
Tasks
Simultaneous Speech-to-Text TranslationSpeech-to-TextSpeech-to-Text TranslationTranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Re-Translation Strategies For Long Form, Simultaneous, Spoken Language Translation
We investigate the problem of simultaneous machine translation of long-form speech content. We target a continuous speech-to-text scenario, generating translated captions for a live audio feed, such as a lecture or play-…
FormMachine Translationspeech-recognitionSpeech Recognition+2Presenting Simultaneous Translation in Limited Space
Some methods of automatic simultaneous translation of a long-form speech allow revisions of outputs, trading accuracy for low latency. Deploying these systems for users faces the problem of presenting subtitles in a limi…
TranslationIncremental Blockwise Beam Search for Simultaneous Speech Translation with Controllable Quality-Latency Tradeoff
Blockwise self-attentional encoder models have recently emerged as one promising end-to-end approach to simultaneous speech translation. These models employ a blockwise beam search with hypothesis reliability scoring to …
TranslationBreaking the Data Barrier: Towards Robust Speech Translation via Adversarial Stability Training
In a pipeline speech translation system, automatic speech recognition (ASR) system will transmit errors in recognition to the downstream machine translation (MT) system. A standard machine translation system is usually t…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationDecoder+4FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech
Current speech generation research can be categorized into two primary classes: non-autoregressive and autoregressive. The fundamental distinction between these approaches lies in the duration prediction strategy employe…
Style Transfertext-to-speechText to Speech