paper-with-me

홈 › Papers

A Massive Scale Semantic Similarity Dataset of Historical English

2023-06-30 · NeurIPS 2023 11

A diversity of tasks use language models trained on semantic similarity data. While there are a variety of datasets that capture semantic similarity, they are either constructed from modern web data or are relatively small datasets created in the past decade by human annotators. This study utilizes a novel source, newly digitized articles from off-copyright, local U.S. newspapers, to assemble a massive-scale semantic similarity dataset spanning 70 years from 1920 to 1989 and containing nearly 400M positive semantic similarity pairs. Historically, around half of articles in U.S. local newspapers came from newswires like the Associated Press. While local papers reproduced articles from the newswire, they wrote their own headlines, which form abstractive summaries of the associated articles. We associate articles and their headlines by exploiting document layouts and language understanding. We then use deep neural methods to detect which articles are from the same underlying source, in the presence of substantial noise and abridgement. The headlines of reproduced articles form positive semantic similarity pairs. The resulting publicly available HEADLINES dataset is significantly larger than most existing semantic similarity datasets and covers a much longer span of time. It will facilitate the application of contrastively trained semantic similarity models to a variety of tasks, including the study of semantic change across space and time.

📄 PDF Abstract BibTeX arXiv:2306.17810

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesSemantic SimilaritySemantic Textual Similarity

Similar Papers 제목 키워드 기반

Contrastive Entity Coreference and Disambiguation for Historical Texts

2024-06-21 · Abhishek Arora, Emily Silcock, Leander Heldring, Melissa Dell

Massive-scale historical document collections are crucial for social science research. Despite increasing digitization, these documents typically lack unique cross-document identifiers for individuals mentioned within th…

Articlescoreference-resolutionCoreference ResolutionCross Document Coreference Resolution+1

News Deja Vu: Connecting Past and Present with Semantic Search

2024-06-21 · Brevin Franklin, Emily Silcock, Abhishek Arora, Tom Bryan 외

Social scientists and the general public often analyze contemporary events by drawing parallels with the past, a process complicated by the vast, noisy, and unstructured nature of historical texts. For example, hundreds …

ArticlesOptical Character Recognition (OCR)

Advancing Knowledge Tracing by Exploring Follow-up Performance Trends

2025-08-11 · Hengyu Liu, Yushuai Li, Minghe Yu, Tiancheng Zhang 외 arxiv

Intelligent Tutoring Systems (ITS), such as Massive Open Online Courses, offer new opportunities for human learning. At the core of such systems, knowledge tracing (KT) predicts students' future performance by analyzing …

Knowledge Tracing

Deep Learning-based Online Alternative Product Recommendations at Scale

2021-04-15 · WS 2020 7 · Mingming Guo, Nian Yan, Xiquan Cui, San He Wu 외

Alternative recommender systems are critical for ecommerce companies. They guide customers to explore a massive product catalog and assist customers to find the right products among an overwhelming number of options. How…

Deep LearningRecommendation Systems

Sentence Embedding Models for Ancient Greek Using Multilingual Knowledge Distillation

2023-08-24 · Kevin Krahn, Derrick Tate, Andrew C. Lamicela

Contextual language models have been trained on Classical languages, including Ancient Greek and Latin, for tasks such as lemmatization, morphological tagging, part of speech tagging, authorship attribution, and detectio…

Authorship AttributionKnowledge DistillationLemmatizationMorphological Tagging+10