paper-with-me

홈 › Papers

Spoken Language Translation for Polish

2015-11-24 · Krzysztof Marasek, Łukasz Brocki, Danijel Korzinek, Krzysztof Wołk, Ryszard Gubrynowicz

Spoken language translation (SLT) is becoming more important in the increasingly globalized world, both from a social and economic point of view. It is one of the major challenges for automatic speech recognition (ASR) and machine translation (MT), driving intense research activities in these areas. While past research in SLT, due to technology limitations, dealt mostly with speech recorded under controlled conditions, today's major challenge is the translation of spoken language as it can be found in real life. Considered application scenarios range from portable translators for tourists, lectures and presentations translation, to broadcast news and shows with live captioning. We would like to present PJIIT's experiences in the SLT gained from the Eu-Bridge 7th framework project and the U-Star consortium activities for the Polish/English language pair. Presented research concentrates on ASR adaptation for Polish (state-of-the-art acoustic models: DBN-BLSTM training, Kaldi: LDA+MLLT+SAT+MMI), language modeling for ASR & MT (text normalization, RNN-based LMs, n-gram model domain interpolation) and statistical translation techniques (hierarchical models, factored translation models, automatic casing and punctuation, comparable and bilingual corpora preparation). While results for the well-defined domains (phrases for travelers, parliament speeches, medical documentation, movie subtitling) are very encouraging, less defined domains (presentation, lectures) still form a challenge. Our progress in the IWSLT TED task (MT only) will be presented, as well as current progress in the Polish ASR.

📄 PDF Abstract BibTeX arXiv:1511.07788

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage ModellingMachine Translationspeech-recognitionSpeech RecognitionText NormalizationTranslation

Similar Papers 제목 키워드 기반

Polish to English Statistical Machine Translation

2015-09-30 · Krzysztof Wołk

This research explores the effects of various training settings on a Polish to English Statistical Machine Translation system for spoken language. Various elements of the TED, Europarl, and OPUS parallel text corpora wer…

Machine TranslationTranslation

Polish - English Speech Statistical Machine Translation Systems for the IWSLT 2013

2015-09-30 · Krzysztof Wołk, Krzysztof Marasek

This research explores the effects of various training settings from Polish to English Statistical Machine Translation system for spoken language. Various elements of the TED parallel text corpora for the IWSLT 2013 eval…

Machine TranslationTranslation

Polish - English Speech Statistical Machine Translation Systems for the IWSLT 2014

2015-09-29 · Krzysztof Wołk, Krzysztof Marasek

This research explores effects of various training settings between Polish and English Statistical Machine Translation systems for spoken language. Various elements of the TED parallel text corpora for the IWSLT 2014 eva…

LEMMAMachine TranslationTranslation

Enhancements in statistical spoken language translation by de-normalization of ASR results

2015-11-18 · Agnieszka Wołk, Krzysztof Wołk, Krzysztof Marasek

Spoken language translation (SLT) has become very important in an increasingly globalized world. Machine translation (MT) for automatic speech recognition (ASR) systems is a major challenge of great interest. This resear…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine TranslationSegmentation+5

Numbers Normalisation in the Inflected Languages: a Case Study of Polish

2019-08-01 · WS 2019 8 · Rafa{\l} Po{\'s}wiata, Micha{\l} Pere{\l}kiewicz

Text normalisation in Text-to-Speech systems is a process of converting written expressions to their spoken forms. This task is complicated because in many cases the normalised form depends on the context. Furthermore, w…

text-to-speechText to Speech