paper-with-me

Papers

Generating Multilingual Parallel Corpus Using Subtitles

2018-04-11 · Farshad Jafari

Neural Machine Translation with its significant results, still has a great problem: lack or absence of parallel corpus for many languages. This article suggests a method for generating considerable amount of parallel corpus for any language pairs, extracted from open source materials existing on the Internet. Parallel corpus contents will be derived from video subtitles. It needs a set of video titles, with some attributes like release date, rating, duration and etc. Process of finding and downloading subtitle pairs for desired language pairs is automated by using a crawler. Finally sentence pairs will be extracted from synchronous dialogues in subtitles. The main problem of this method is unsynchronized subtitle pairs. Therefore subtitles will be verified before downloading. If two subtitle were not synchronized, then another subtitle of that video will be processed till it finds the matching subtitle. Using this approach gives ability to make context based parallel corpus through filtering videos by genre. Context based corpus can be used in complex translators which decode sentences by different networks after determining contents subject. Languages have many differences in their formal and informal styles, including words and syntax. Other advantage of this method is to make corpus of informal style of languages. Because most of movies dialogues are parts of a conversation. So they had informal style. This feature of generated corpus can be used in real-time translators to have more accurate conversation translations.

📄 PDF Abstract BibTeX arXiv:1804.03923

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentence

Similar Papers 제목 키워드 기반

The AMARA Corpus: Building Parallel Language Resources for the Educational Domain

2014-05-01 · LREC 2014 5 · Ahmed Abdelali, Francisco Guzman, Hassan Sajjad, Stephan Vogel

This paper presents the AMARA corpus of on-line educational content: a new parallel corpus of educational video subtitles, multilingually aligned for 20 languages, i.e. 20 monolingual corpora and 190 parallel corpora. Th…

Machine TranslationTranslation

Building a Hebrew Semantic Role Labeling Lexical Resource from Parallel Movie Subtitles

2020-05-17 · LREC 2020 5 · Ben Eyal, Michael Elhadad

We present a semantic role labeling resource for Hebrew built semi-automatically through annotation projection from English. This corpus is derived from the multilingual OpenSubtitles dataset and includes short informal …

Morphological AnalysisSemantic Role Labeling

Innovations in Parallel Corpus Search Tools

2014-05-01 · LREC 2014 5 · Martin Volk, Johannes Gra{\"e}n, Elena Callegaro

Recent years have seen an increased interest in and availability of parallel corpora. Large corpora from international organizations (e.g. European Union, United Nations, European Patent Office), or from multilingual Int…

Machine TranslationSentenceTranslationWord Alignment+1

EduMT: Developing Machine Translation System for Educational Content in Indian Languages

2021-12-01 · ICON 2021 12 · Ramakrishna Appicharla, Asif Ekbal, Pushpak Bhattacharyya

In this paper, we explore various approaches to build Hindi to Bengali Neural Machine Translation (NMT) systems for the educational domain. Translation of educational content poses several challenges, such as unavailabil…

Data AugmentationDomain AdaptationMachine TranslationNMT+1

Breaking Bad: Extraction of Verb-Particle Constructions from a Parallel Subtitles Corpus

2014-04-01 · WS 2014 4 · Aaron Smith
Machine TranslationPart-Of-Speech TaggingWord Alignment