paper-with-me

홈 › Papers

News Deja Vu: Connecting Past and Present with Semantic Search

2024-06-21 · Brevin Franklin, Emily Silcock, Abhishek Arora, Tom Bryan, Melissa Dell

Social scientists and the general public often analyze contemporary events by drawing parallels with the past, a process complicated by the vast, noisy, and unstructured nature of historical texts. For example, hundreds of millions of page scans from historical newspapers have been noisily transcribed. Traditional sparse methods for searching for relevant material in these vast corpora, e.g., with keywords, can be brittle given complex vocabularies and OCR noise. This study introduces News Deja Vu, a novel semantic search tool that leverages transformer large language models and a bi-encoder approach to identify historical news articles that are most similar to modern news queries. News Deja Vu first recognizes and masks entities, in order to focus on broader parallels rather than the specific named entities being discussed. Then, a contrastively trained, lightweight bi-encoder retrieves historical articles that are most similar semantically to a modern query, illustrating how phenomena that might seem unique to the present have varied historical precedents. Aimed at social scientists, the user-friendly News Deja Vu package is designed to be accessible for those who lack extensive familiarity with deep learning. It works with large text datasets, and we show how it can be deployed to a massive scale corpus of historical, open-source news articles. While human expertise remains important for drawing deeper insights, News Deja Vu provides a powerful tool for exploring parallels in how people have perceived past and present.

📄 PDF Abstract BibTeX arXiv:2406.15593

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesOptical Character Recognition (OCR)

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Been There, Scanned That: Nostalgia-Driven LiDAR Compression for Self-Driving Cars

2025-11-01 · Ali Khalid, Jaiaid Mobin, Sumanth Rao Appala, Avinash Maurya 외 arxiv

An autonomous vehicle can generate several terabytes of sensor data per day. A significant portion of this data consists of 3D point clouds produced by depth sensors such as LiDARs. This data must be transferred to cloud…

Autonomous VehiclesPoint Clouds

Dejavu: Towards Experience Feedback Learning for Embodied Intelligence

2025-10-11 · Shaokai Wu, Yanbiao Ji, Qiuchang Li, Zhiyi Zhang 외 arxiv

Embodied agents face a fundamental limitation: once deployed in real-world environments, they cannot easily acquire new knowledge to improve task performance. In this paper, we propose Dejavu, a general post-deployment l…

Reinforcement LearningSemantic Similarity

DejaVu: Conditional Regenerative Learning to Enhance Dense Prediction

2023-03-02 · CVPR 2023 1 · Shubhankar Borse, Debasmit Das, Hyojin Park, Hong Cai 외

We present DejaVu, a novel framework which leverages conditional image regeneration as additional supervision during training to improve deep networks for dense prediction tasks such as segmentation, depth estimation, an…

Depth EstimationPrediction

DejAIvu: Identifying and Explaining AI Art on the Web in Real-Time with Saliency Maps

2025-02-12 · Jocelyn Dzuong

The recent surge in advanced generative models, such as diffusion models and generative adversarial networks (GANs), has led to an alarming rise in AI-generated images across various domains on the web. While such techno…

MarketingMisinformation

Graph with Sequence: Broad-Range Semantic Modeling for Fake News Detection

2024-12-07 · Junwei Yin, Min Gao, Kai Shu, Wentao Li 외

The rapid proliferation of fake news on social media threatens social stability, creating an urgent demand for more effective detection methods. While many promising approaches have emerged, most rely on content analysis…

DenoisingFake News Detection