paper-with-me

홈 › Papers

Long-form Simultaneous Speech Translation: Thesis Proposal

2023-10-17 · Peter Polák

Simultaneous speech translation (SST) aims to provide real-time translation of spoken language, even before the speaker finishes their sentence. Traditionally, SST has been addressed primarily by cascaded systems that decompose the task into subtasks, including speech recognition, segmentation, and machine translation. However, the advent of deep learning has sparked significant interest in end-to-end (E2E) systems. Nevertheless, a major limitation of most approaches to E2E SST reported in the current literature is that they assume that the source speech is pre-segmented into sentences, which is a significant obstacle for practical, real-world applications. This thesis proposal addresses end-to-end simultaneous speech translation, particularly in the long-form setting, i.e., without pre-segmentation. We present a survey of the latest advancements in E2E SST, assess the primary obstacles in SST and its relevance to long-form scenarios, and suggest approaches to tackle these challenges.

📄 PDF Abstract BibTeX arXiv:2310.11141

Code (0)

등록된 구현이 없습니다.

Tasks

FormMachine TranslationSegmentationSentencespeech-recognitionSpeech RecognitionTranslation

Similar Papers 제목 키워드 기반

Simultaneous Speech-to-Speech Translation System with Neural Incremental ASR, MT, and TTS

2020-11-10 · Katsuhito Sudoh, Takatomo Kano, Sashi Novitasari, Tomoya Yanagita 외

This paper presents a newly developed, simultaneous neural speech-to-speech translation system and its evaluation. The system consists of three fully-incremental neural processing modules for automatic speech recognition…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationSimultaneous Speech-to-Speech Translation+8

Simultaneous Translation

2020-11-01 · EMNLP 2020 11 · Liang Huang, Colin Cherry, Mingbo Ma, Naveen Arivazhagan 외

Simultaneous translation, which performs translation concurrently with the source speech, is widely useful in many scenarios such as international conferences, negotiations, press releases, legal proceedings, and medicin…

Machine Translationspeech-recognitionSpeech RecognitionSpeech Synthesis+1

MLLP-VRAIN UPV systems for the IWSLT 2022 Simultaneous Speech Translation and Speech-to-Speech Translation tasks

2022-05-01 · IWSLT (ACL) 2022 5 · Javier Iranzo-Sánchez, Javier Jorge Cano, Alejandro Pérez-González-de-Martos, Adrián Giménez Pastor 외

This work describes the participation of the MLLP-VRAIN research group in the two shared tasks of the IWSLT 2022 conference: Simultaneous Speech Translation and Speech-to-Speech Translation. We present our streaming-read…

Simultaneous Speech-to-Text TranslationSpeech-to-Speech TranslationTranslation

StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning

2024-06-05 · Shaolei Zhang, Qingkai Fang, Shoutao Guo, Zhengrui Ma 외

Simultaneous speech-to-speech translation (Simul-S2ST, a.k.a streaming speech translation) outputs target speech while receiving streaming speech inputs, which is critical for real-time communication. Beyond accomplishin…

Automatic Speech Recognition (ASR)de-enes-enfr-en+11

Direct Simultaneous Speech-to-Speech Translation with Variational Monotonic Multihead Attention

2021-10-15 · Xutai Ma, Hongyu Gong, Danni Liu, Ann Lee 외

We present a direct simultaneous speech-to-speech translation (Simul-S2ST) model, Furthermore, the generation of translation is independent from intermediate text representations. Our approach leverages recent progress o…

Simultaneous Speech-to-Speech TranslationSpeech SynthesisSpeech-to-Speech TranslationTranslation