paper-with-me

Papers

Finnish Paraphrase Corpus

2021-03-24 · NoDaLiDa 2021 5 · Jenna Kanerva, Filip Ginter, Li-Hsin Chang, Iiro Rastas, Valtteri Skantsi, Jemina Kilpeläinen, Hanna-Mari Kupari, Jenna Saarni, Maija Sevón, Otto Tarkka

In this paper, we introduce the first fully manually annotated paraphrase corpus for Finnish containing 53,572 paraphrase pairs harvested from alternative subtitles and news headings. Out of all paraphrase pairs in our corpus 98% are manually classified to be paraphrases at least in their given context, if not in all contexts. Additionally, we establish a manual candidate selection method and demonstrate its feasibility in high quality paraphrase selection in terms of both cost and quality.

📄 PDF Abstract BibTeX arXiv:2103.13103

Code (1)

TurkuNLP/Turku-paraphrase-corpus 공식 구현

Tasks

All

Similar Papers 제목 키워드 기반

Quantitative Evaluation of Alternative Translations in a Corpus of Highly Dissimilar Finnish Paraphrases

2021-05-06 · MoTra (NoDaLiDa) 2021 5 · Li-Hsin Chang, Sampo Pyysalo, Jenna Kanerva, Filip Ginter

In this paper, we present a quantitative evaluation of differences between alternative translations in a large recently released Finnish paraphrase corpus focusing in particular on non-trivial variation in translation. W…

Translation

Annotation Guidelines for the Turku Paraphrase Corpus

2021-08-17 · Jenna Kanerva, Filip Ginter, Li-Hsin Chang, Iiro Rastas 외

This document describes the annotation guidelines used to construct the Turku Paraphrase Corpus. These guidelines were developed together with the corpus annotation, revising and extending the guidelines regularly during…

Coping with Noisy Training Data Labels in Paraphrase Detection

2021-11-01 · WNUT (ACL) 2021 11 · Teemu Vahtola, Mathias Creutz, Eetu Sjöblom, Sami Itkonen

We present new state-of-the-art benchmarks for paraphrase detection on all six languages in the Opusparcus sentential paraphrase corpus: English, Finnish, French, German, Russian, and Swedish. We reach these baselines by…

Translation

Open Subtitles Paraphrase Corpus for Six Languages

2018-09-17 · LREC 2018 5 · Mathias Creutz

This paper accompanies the release of Opusparcus, a new paraphrase corpus for six European languages: German, English, Finnish, French, Russian, and Swedish. The corpus consists of paraphrases, that is, pairs of sentence…

Sentence

Paraphrase Generation and Evaluation on Colloquial-Style Sentences

2020-05-01 · LREC 2020 5 · Eetu Sj{\"o}blom, Mathias Creutz, Yves Scherrer

In this paper, we investigate paraphrase generation in the colloquial domain. We use state-of-the-art neural machine translation models trained on the Opusparcus corpus to generate paraphrases in six languages: German, E…

Machine TranslationParaphrase GenerationSentenceTranslation