paper-with-me

홈 › Papers

GigaST: A 10,000-hour Pseudo Speech Translation Corpus

2022-04-08 · Rong Ye, Chengqi Zhao, Tom Ko, Chutong Meng, Tao Wang, Mingxuan Wang, Jun Cao

This paper introduces GigaST, a large-scale pseudo speech translation (ST) corpus. We create the corpus by translating the text in GigaSpeech, an English ASR corpus, into German and Chinese. The training set is translated by a strong machine translation system and the test set is translated by human. ST models trained with an addition of our corpus obtain new state-of-the-art results on the MuST-C English-German benchmark test set. We provide a detailed description of the translation process and verify its quality. We make the translated text data public and hope to facilitate research in speech translation. Additionally, we also release the training scripts on NeurST to make it easy to replicate our systems. GigaST dataset is available at https://st-benchmark.github.io/resources/GigaST.

📄 PDF Abstract BibTeX arXiv:2204.03939

Code (1)

speechtranslation/gigas2s

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

LibriVoxDeEn: A Corpus for German-to-English Speech Translation and German Speech Recognition

2019-10-17 · LREC 2020 5 · Benjamin Beilharz, Xin Sun, Sariya Karimova, Stefan Riezler

We present a corpus of sentence-aligned triples of German audio, German text, and English translation, based on German audiobooks. The speech translation data consist of 110 hours of audio material aligned to over 50k pa…

Sentencespeech-recognitionSpeech RecognitionTranslation

Improving speech translation by fusing speech and text

2023-05-23 · Wenbiao Yin, Zhicheng Liu, Chengqi Zhao, Tao Wang 외

In speech translation, leveraging multimodal data to improve model performance and address limitations of individual modalities has shown significant effectiveness. In this paper, we harness the complementary strengths o…

cross-modal alignmentMachine TranslationTranslation

MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation

2023-03-01 · Mohamed Anwar, Bowen Shi, Vedanuj Goswami, Wei-Ning Hsu 외

We introduce MuAViC, a multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation providing 1200 hours of audio-visual speech in 9 languages. It is fully transcribed and covers 6…

Audio-Visual Speech RecognitionRobust Speech Recognitionspeech-recognitionSpeech Recognition+4

BSTC: A Large-Scale Chinese-English Speech Translation Dataset

2021-04-08 · NAACL (AutoSimTrans) 2021 6 · Ruiqing Zhang, Xiyang Wang, Chuanqiang Zhang, Zhongjun He 외

This paper presents BSTC (Baidu Speech Translation Corpus), a large-scale Chinese-English speech translation dataset. This dataset is constructed based on a collection of licensed videos of talks or lectures, including a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations

2022-11-08 · arXiv 2022 10 · Paul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du 외

We present SpeechMatrix, a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings. It contains speech alignments in 136 language pairs with a total of 41…

Mixture-of-ExpertsSpeech-to-Speech TranslationTranslation