paper-with-me

Papers

Mining Naturally-occurring Corrections and Paraphrases from Wikipedia's Revision History

2022-02-25 · Aurélien Max, Guillaume Wisniewski

Naturally-occurring instances of linguistic phenomena are important both for training and for evaluating automatic processes on text. When available in large quantities, they also prove interesting material for linguistic studies. In this article, we present a new resource built from Wikipedia's revision history, called WiCoPaCo (Wikipedia Correction and Paraphrase Corpus), which contains numerous editings by human contributors, including various corrections and rewritings. We discuss the main motivations for building such a resource, describe how it was built and present initial applications on French.

📄 PDF Abstract BibTeX arXiv:2202.12575

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Paraphrase Acquisition from Image Captions

2023-01-26 · Marcel Gohsen, Matthias Hagen, Martin Potthast, Benno Stein

We propose to use image captions from the Web as a previously underutilized resource for paraphrases (i.e., texts with the same "message") and to create and analyze a corresponding dataset. When an image is reused on the…

ArticlesImage Captioning

Learning To Split and Rephrase From Wikipedia Edit History

2018-08-28 · EMNLP 2018 10 · Jan A. Botha, Manaal Faruqui, John Alex, Jason Baldridge 외

Split and rephrase is the task of breaking down a sentence into shorter ones that together convey the same meaning. We extract a rich new dataset for this task by mining Wikipedia's edit history: WikiSplit contains one m…

SentenceSplit and Rephrase

Multilingual Whispers: Generating Paraphrases with Translation

2019-11-01 · WS 2019 11 · Christian Federmann, Oussama Elachqar, Chris Quirk

Naturally occurring paraphrase data, such as multiple news stories about the same event, is a useful but rare resource. This paper compares translation-based paraphrase gathering using human, automatic, or hybrid techniq…

DiversityMachine TranslationTranslation

CHEW: A Dataset of CHanging Events in Wikipedia

2024-06-27 · Hsuvas Borkakoty, Luis Espinosa-Anke

We introduce CHEW, a novel dataset of changing events in Wikipedia expressed in naturally occurring text. We use CHEW for probing LLMs for their timeline understanding of Wikipedia entities and events in generative and c…

What Did This Castle Look like before? Exploring Referential Relations in Naturally Occurring Multimodal Texts

2021-04-01 · EACL (LANTERN) 2021 4 · Ronja Utescher, Sina Zarrieß

Multi-modal texts are abundant and diverse in structure, yet Language & Vision research of these naturally occurring texts has mostly focused on genres that are comparatively light on text, like tweets. In this paper, we…

ArticlesSurvey