paper-with-me

홈 › Papers

Tracking the Newsworthiness of Public Documents

2023-11-16 · Alexander Spangher, Emilio Ferrara, Ben Welsh, Nanyun Peng, Serdar Tumgoren, Jonathan May

Journalists must find stories in huge amounts of textual data (e.g. leaks, bills, press releases) as part of their jobs: determining when and why text becomes news can help us understand coverage patterns and help us build assistive tools. Yet, this is challenging because very few labelled links exist, language use between corpora is very different, and text may be covered for a variety of reasons. In this work we focus on news coverage of local public policy in the San Francisco Bay Area by the San Francisco Chronicle. First, we gather news articles, public policy documents and meeting recordings and link them using probabilistic relational modeling, which we show is a low-annotation linking methodology that outperforms other retrieval-based baselines. Second, we define a new task: newsworthiness prediction, to predict if a policy item will get covered. We show that different aspects of public policy discussion yield different newsworthiness signals. Finally we perform human evaluation with expert journalists and show our systems identify policies they consider newsworthy with 68% F1 and our coverage recommendations are helpful with an 84% win-rate.

📄 PDF Abstract BibTeX arXiv:2311.09734

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesRetrieval

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Modeling "Newsworthiness" for Lead-Generation Across Corpora

2021-04-19 · Alexander Spangher, Nanyun Peng, Jonathan May, Emilio Ferrara

Journalists obtain "leads", or story ideas, by reading large corpora of government records: court cases, proposed bills, etc. However, only a small percentage of such records are interesting documents. We propose a model…

Articles

Modeling Public Perceptions of Science in Media

2025-06-19 · Jiaxin Pei, Dustin Wright, Isabelle Augenstin, David Jurgens

Effectively engaging the public with science is vital for fostering trust and understanding in our scientific community. Yet, with an ever-growing volume of information, science communicators struggle to anticipate how a…

Multiple Document Representations from News Alerts for Automated Bio-surveillance Event Detection

2019-02-17 · Aaron Tuor, Fnu Anubhav, Lauren Charles

Due to globalization, geographic boundaries no longer serve as effective shields for the spread of infectious diseases. In order to aid bio-surveillance analysts in disease tracking, recent research has been devoted to d…

ClassificationEvent DetectionGeneral ClassificationInformation Retrieval+1

RESIN: A Dockerized Schema-Guided Cross-document Cross-lingual Cross-media Information Extraction and Event Tracking System

2021-06-01 · NAACL 2021 4 · Haoyang Wen, Ying Lin, Tuan Lai, Xiaoman Pan 외

We present a new information extraction system that can automatically construct temporal event graphs from a collection of news documents from multiple sources, multiple languages (English and Spanish for our experiment)…

coreference-resolutionCoreference ResolutionEvent ExtractionSentence

The Birth of Collective Memories: Analyzing Emerging Entities in Text Streams

2017-01-15 · David Graus, Daan Odijk, Maarten de Rijke

We study how collective memories are formed online. We do so by tracking entities that emerge in public discourse, that is, in online text streams such as social media and news streams, before they are incorporated into …