paper-with-me

Papers

Open Subtitles Paraphrase Corpus for Six Languages

2018-09-17 · LREC 2018 5 · Mathias Creutz

This paper accompanies the release of Opusparcus, a new paraphrase corpus for six European languages: German, English, Finnish, French, Russian, and Swedish. The corpus consists of paraphrases, that is, pairs of sentences in the same language that mean approximately the same thing. The paraphrases are extracted from the OpenSubtitles2016 corpus, which contains subtitles from movies and TV shows. The informal and colloquial genre that occurs in subtitles makes such data a very interesting language resource, for instance, from the perspective of computer assisted language learning. For each target language, the Opusparcus data have been partitioned into three types of data sets: training, development and test sets. The training sets are large, consisting of millions of sentence pairs, and have been compiled automatically, with the help of probabilistic ranking functions. The development and test sets consist of sentence pairs that have been checked manually; each set contains approximately 1000 sentence pairs that have been verified to be acceptable paraphrases by two annotators.

📄 PDF Abstract BibTeX arXiv:1809.06142

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

Finding Alternative Translations in a Large Corpus of Movie Subtitle

2016-05-01 · LREC 2016 5 · J{\"o}rg Tiedemann

OpenSubtitles.org provides a large collection of user contributed subtitles in various languages for movies and TV programs. Subtitle translations are valuable resources for cross-lingual studies and machine translation …

Machine TranslationSentenceTranslation

Paraphrase Detection on Noisy Subtitles in Six Languages

2018-09-21 · WS 2018 11 · Eetu Sjöblom, Mathias Creutz, Mikko Aulamo

We perform automatic paraphrase detection on subtitle data from the Opusparcus corpus comprising six European languages: German, English, Finnish, French, Russian, and Swedish. We train two types of supervised sentence e…

SentenceSentence EmbeddingSentence-Embedding

A contrastive review of paraphrase acquisition techniques

2012-05-01 · LREC 2012 5 · Houda Bouamor, Aur{\'e}lien Max, Gabriel Illouz, Anne Vilnat

This paper addresses the issue of what approach should be used for building a corpus of sententential paraphrases depending on one's requirements. Six strategies are studied: (1) multiple translations into a single langu…

ArticlesInformation RetrievalMachine Translation

Finnish Paraphrase Corpus

2021-03-24 · NoDaLiDa 2021 5 · Jenna Kanerva, Filip Ginter, Li-Hsin Chang, Iiro Rastas 외

In this paper, we introduce the first fully manually annotated paraphrase corpus for Finnish containing 53,572 paraphrase pairs harvested from alternative subtitles and news headings. Out of all paraphrase pairs in our c…

All

Generating Multilingual Parallel Corpus Using Subtitles

2018-04-11 · Farshad Jafari

Neural Machine Translation with its significant results, still has a great problem: lack or absence of parallel corpus for many languages. This article suggests a method for generating considerable amount of parallel cor…

Machine TranslationSentence