On Paraphrase Identification Corpora
We analyze in this paper a number of data sets proposed over the last decade or so for the task of paraphrase identification. The goal of the analysis is to identify the advantages as well as shortcomings of the previously proposed data sets. Based on the analysis, we then make recommendations about how to improve the process of creating and using such data sets for evaluating in the future approaches to the task of paraphrase identification or the more general task of semantic similarity. The recommendations are meant to improve our understanding of what a paraphrase is, offer a more fair ground for comparing approaches, increase the diversity of actual linguistic phenomena that future data sets will cover, and offer ways to improve our understanding of the contributions of various modules or approaches proposed for solving the task of paraphrase identification or similar tasks.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversityNatural Language InferenceParaphrase IdentificationSemantic SimilaritySemantic Textual SimilaritySimilar Papers 제목 키워드 기반
Improving Large-scale Paraphrase Acquisition and Generation
This paper addresses the quality issues in existing Twitter-based paraphrase datasets, and discusses the necessity of using two separate definitions of paraphrase for identification and generation tasks. We present a new…
Language ModelingLanguage ModellingParaphrase GenerationParaphrase Identification+1A Continuously Growing Dataset of Sentential Paraphrases
A major challenge in paraphrase research is the lack of parallel corpora. In this paper, we present a new method to collect large-scale sentential paraphrases from Twitter by linking tweets through shared URLs. The main …
BenchmarkingParaphrase IdentificationSentencePARADE: A New Dataset for Paraphrase Identification Requiring Computer Science Domain Knowledge
We present a new benchmark dataset called PARADE for paraphrase identification that requires specialized domain knowledge. PARADE contains paraphrases that overlap very little at the lexical and syntactic level but are s…
Paraphrase IdentificationPredicate-Argument Based Bi-Encoder for Paraphrase Identification
Paraphrase identification involves identifying whether a pair of sentences express the same or similar meanings. While cross-encoders have achieved high performances across several benchmarks, bi-encoders such as SBERT h…
Paraphrase IdentificationSentencePredicate-Argument Based Bi-Encoder for Paraphrase Identification
Paraphrase identification involves identifying whether a pair of sentences express the same or similar meanings. While cross-encoders have achieved high performances across several benchmarks, bi-encoders such as SBERT h…
Paraphrase IdentificationSentence