IESTAC: English-Italian Parallel Corpus for End-to-End Speech-to-Text Machine Translation
We discuss a set of methods for the creation of IESTAC: a English-Italian speech and text parallel corpus designed for the training of end-to-end speech-to-text machine translation models and publicly released as part of this work. We first mapped English LibriVox audiobooks and their corresponding English Gutenberg Project e-books to Italian e-books with a set of three complementary methods. Then we aligned the English and the Italian texts using both traditional Gale-Church based alignment methods and a recently proposed tool to perform bilingual sentences alignment computing the cosine similarity of multilingual sentence embeddings. Finally, we forced the alignment between the English audiobooks and the English side of our textual parallel corpus with a text-to-speech and dynamic time warping based forced alignment tool. For each step, we provide the reader with a critical discussion based on detailed evaluation and comparison of the results of the different methods.
Code (1)
Tasks
Dynamic Time WarpingMachine TranslationSentenceSentence EmbeddingsSpeech-to-Texttext-to-speechText to SpeechTranslationSimilar Papers 제목 키워드 기반
Using English as Pivot to Extract Persian-Italian Parallel Sentences from Non-Parallel Corpora
The effectiveness of a statistical machine translation system (SMT) is very dependent upon the amount of parallel corpus used in the training phase. For low-resource language pairs there are not enough parallel corpora t…
Machine TranslationSentenceSentence SimilarityTranslationSentiment recognition of Italian elderly through domain adaptation on cross-corpus speech dataset
The aim of this work is to define a speech emotion recognition (SER) model able to recognize positive, neutral and negative emotions in natural conversations of Italian elderly people. Several datasets for SER are availa…
Cross-corpusDomain AdaptationEmotion RecognitionSpeech Emotion RecognitionSwissAdmin: A multilingual tagged parallel corpus of press releases
SwissAdmin is a new multilingual corpus of press releases from the Swiss Federal Administration, available in German, French, Italian and English. We provide SwissAdmin in three versions: (i) plain texts of approximately…
Language IdentificationSentenceDIETA: A Decoder-only transformer-based model for Italian-English machine TrAnslation
In this paper, we present DIETA, a small, decoder-only Transformer model with 0.5 billion parameters, specifically designed and trained for Italian-English machine translation. We collect and curate a large parallel corp…
Machine TranslationA small Griko-Italian speech translation corpus
This paper presents an extension to a very low-resource parallel corpus collected in an endangered language, Griko, making it useful for computational research. The corpus consists of 330 utterances (about 20 minutes of …
DiversityTranslation