paper-with-me

홈 › Papers

Collection of a Simultaneous Translation Corpus for Comparative Analysis

2014-05-01 · LREC 2014 5 · Hiroaki Shimizu, Graham Neubig, Sakriani Sakti, Tomoki Toda, Satoshi Nakamura

This paper describes the collection of an English-Japanese/Japanese-English simultaneous interpretation corpus. There are two main features of the corpus. The first is that professional simultaneous interpreters with different amounts of experience cooperated with the collection. By comparing data from simultaneous interpretation of each interpreter, it is possible to compare better interpretations to those that are not as good. The second is that for part of our corpus there are already translation data available. This makes it possible to compare translation data with simultaneous interpretation data. We recorded the interpretations of lectures and news, and created time-aligned transcriptions. A total of 387k words of transcribed data were collected. The corpus will be helpful to analyze differences in interpretations styles and to construct simultaneous interpretation systems.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Translation

Similar Papers 제목 키워드 기반

BSTC: A Large-Scale Chinese-English Speech Translation Dataset

2021-04-08 · NAACL (AutoSimTrans) 2021 6 · Ruiqing Zhang, Xiyang Wang, Chuanqiang Zhang, Zhongjun He 외

This paper presents BSTC (Baidu Speech Translation Corpus), a large-scale Chinese-English speech translation dataset. This dataset is constructed based on a collection of licensed videos of talks or lectures, including a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Semantic Relatedness for All (Languages): A Comparative Analysis of Multilingual Semantic Relatedness Using Machine Translation

2018-05-16 · Andre Freitas, Siamak Barzegar, Juliano Efson Sales, Siegfried Handschuh 외

This paper provides a comparative analysis of the performance of four state-of-the-art distributional semantic models (DSMs) over 11 languages, contrasting the native language-specific models with the use of machine tran…

AllMachine TranslationTranslation

Statistical Analysis of Missing Translation in Simultaneous Interpretation Using A Large-scale Bilingual Speech Corpus

2018-05-01 · LREC 2018 5 · Zhongxi Cai, Koichiro Ryu, Shigeki Matsubara
Translation

A parallel collection of clinical trials in Portuguese and English

2017-08-01 · WS 2017 8 · Mariana Neves

Parallel collections of documents are crucial resources for training and evaluating machine translation (MT) systems. Even though large collections are available for certain domains and language pairs, these are still sc…

Machine TranslationSentenceTranslation

Extending the MuST-C Corpus for a Comparative Evaluation of Speech Translation Technology

2022-06-01 · EAMT 2022 6 · Luisa Bentivogli, Mauro Cettolo, Marco Gaido, Alina Karakanta 외

This project aimed at extending the test sets of the MuST-C speech translation (ST) corpus with new reference translations. The new references were collected from professional post-editors working on the output of differ…

Machine TranslationTranslation