paper-with-me

홈 › Papers

Towards Corpus-Scale Discovery of Selection Biases in News Coverage: Comparing What Sources Say About Entities as a Start

2023-04-06 · Sihao Chen, William Bruno, Dan Roth

News sources undergo the process of selecting newsworthy information when covering a certain topic. The process inevitably exhibits selection biases, i.e. news sources' typical patterns of choosing what information to include in news coverage, due to their agenda differences. To understand the magnitude and implications of selection biases, one must first discover (1) on what topics do sources typically have diverging definitions of "newsworthy" information, and (2) do the content selection patterns correlate with certain attributes of the news sources, e.g. ideological leaning, etc. The goal of the paper is to investigate and discuss the challenges of building scalable NLP systems for discovering patterns of media selection biases directly from news content in massive-scale news corpora, without relying on labeled data. To facilitate research in this domain, we propose and study a conceptual framework, where we compare how sources typically mention certain controversial entities, and use such as indicators for the sources' content selection preferences. We empirically show the capabilities of the framework through a case study on NELA-2020, a corpus of 1.8M news articles in English from 519 news sources worldwide. We demonstrate an unsupervised representation learning method to capture the selection preferences for how sources typically mention controversial entities. Our experiments show that that distributional divergence of such representations, when studied collectively across entities and news sources, serve as good indicators for an individual source's ideological leaning. We hope our findings will provide insights for future research on media selection biases.

📄 PDF Abstract BibTeX arXiv:2304.03414

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesRepresentation Learning

Similar Papers 제목 키워드 기반

Impacts of Racial Bias in Historical Training Data for News AI

2025-12-18 · Rahul Bhargava, Malene Hornstrup Jespersen, Emily Boardman Ndulue, Vivica Dsouza arxiv

AI technologies have rapidly moved into business and research applications that involve large text corpora, including computational journalism research and newsroom settings. These models, trained on extant data from var…

MLSUM: The Multilingual Summarization Corpus

2020-04-30 · EMNLP 2020 11 · Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski 외

We present MLSUM, the first large-scale MultiLingual SUMmarization dataset. Obtained from online newspapers, it contains 1.5M+ article/summary pairs in five different languages -- namely, French, German, Spanish, Russian…

Text Summarization

Crowdsourcing a Large Corpus of Clickbait on Twitter

2018-08-01 · COLING 2018 8 · Martin Potthast, Tim Gollub, Kristof Komlossy, Sebastian Schuster 외

Clickbait has become a nuisance on social media. To address the urging task of clickbait detection, we constructed a new corpus of 38,517 annotated Twitter tweets, the Webis Clickbait Corpus 2017. To avoid biases in term…

Clickbait Detection

Automatic Extraction of News Values from Headline Text

2017-04-01 · EACL 2017 4 · Alicja Piotrkowicz, Vania Dimitrova, Katja Markert

Headlines play a crucial role in attracting audiences{'} attention to online artefacts (e.g. news articles, videos, blogs). The ability to carry out an automatic, large-scale analysis of headlines is critical to facilita…

ArticlesKeyword SpottingRecommendation Systems

Exploratory Analysis of News Sentiment Using Subgroup Discovery

2021-04-01 · EACL (BSNLP) 2021 4 · Anita Valmarska, Luis Adrián Cabrera-Diego, Elvys Linhares Pontes, Senja Pollak

In this study, we present an exploratory analysis of a Slovenian news corpus, in which we investigate the association between named entities and sentiment in the news. We propose a methodology that combines Named Entity …

Descriptivenamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1