paper-with-me

Papers

Direct Segmentation Models for Streaming Speech Translation

2020-11-01 · EMNLP 2020 11 · Javier Iranzo-S{\'a}nchez, Adri{\`a} Gim{\'e}nez Pastor, Joan Albert Silvestre-Cerd{\`a}, Pau Baquero-Arnal, Jorge Civera Saiz, Alfons Juan

The cascade approach to Speech Translation (ST) is based on a pipeline that concatenates an Automatic Speech Recognition (ASR) system followed by a Machine Translation (MT) system. These systems are usually connected by a segmenter that splits the ASR output into hopefully, semantically self-contained chunks to be fed into the MT system. This is specially challenging in the case of streaming ST, where latency requirements must also be taken into account. This work proposes novel segmentation models for streaming ST that incorporate not only textual, but also acoustic information to decide when the ASR output is split into a chunk. An extensive and throughly experimental setup is carried out on the Europarl-ST dataset to prove the contribution of acoustic information to the performance of the segmentation model in terms of BLEU score in a streaming ST scenario. Finally, comparative results with previous work also show the superiority of the segmentation models proposed in this work.

📄 PDF Abstract BibTeX

Code (1)

jairsan/speech_translation_segmenter 공식 구현 pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationSegmentationspeech-recognitionSpeech RecognitionTranslation

Similar Papers 제목 키워드 기반

End-to-End Simultaneous Speech Translation with Differentiable Segmentation

2023-05-25 · Shaolei Zhang, Yang Feng

End-to-end simultaneous speech translation (SimulST) outputs translation while receiving the streaming speech inputs (a.k.a. streaming speech translation), and hence needs to segment the speech inputs and then translate …

SegmentationTranslation

Learning Adaptive Segmentation Policy for End-to-End Simultaneous Translation

2022-05-01 · ACL 2022 5 · Ruiqing Zhang, Zhongjun He, Hua Wu, Haifeng Wang

End-to-end simultaneous speech-to-text translation aims to directly perform translation from streaming source speech to target text with high translation quality and low latency. A typical simultaneous translation (ST) s…

SegmentationSimultaneous Speech-to-Text TranslationSpeech-to-TextSpeech-to-Text Translation+1

Learning When to Translate for Streaming Speech

2021-09-15 · ACL 2022 5 · Qianqian Dong, Yaoming Zhu, Mingxuan Wang, Lei LI

How to find proper moments to generate partial sentence translation given a streaming speech input? Existing approaches waiting-and-translating for a fixed duration often break the acoustic units in speech, since the bou…

DecoderSentenceSpeech-to-Text TranslationTranslation

Streaming Punctuation: A Novel Punctuation Technique Leveraging Bidirectional Context for Continuous Speech Recognition

2023-01-10 · Piyush Behre, Sharman Tan, Padma Varadharajan, Shuangyu Chang

While speech recognition Word Error Rate (WER) has reached human parity for English, continuous speech recognition scenarios such as voice typing and meeting transcriptions still suffer from segmentation and punctuation …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationSegmentation+2

Segmentation Strategies for Streaming Speech Translation

2013-06-01 · NAACL 2013 6 · Vivek Kumar Rangarajan Sridhar, John Chen, Srinivas Bangalore, Andrej Ljolje 외
ChunkingLanguage ModellingMachine TranslationSegmentation+2