paper-with-me

Papers

Should we trust web-scraped data?

2023-08-04 · Jens Foerderer

The increasing adoption of econometric and machine-learning approaches by empirical researchers has led to a widespread use of one data collection method: web scraping. Web scraping refers to the use of automated computer programs to access websites and download their content. The key argument of this paper is that na\"ive web scraping procedures can lead to sampling bias in the collected data. This article describes three sources of sampling bias in web-scraped data. More specifically, sampling bias emerges from web content being volatile (i.e., being subject to change), personalized (i.e., presented in response to request characteristics), and unindexed (i.e., abundance of a population register). In a series of examples, I illustrate the prevalence and magnitude of sampling bias. To support researchers and reviewers, this paper provides recommendations on anticipating, detecting, and overcoming sampling bias in web-scraped data.

📄 PDF Abstract BibTeX arXiv:2308.02231

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An Interpretable Graph-based Mapping of Trustworthy Machine Learning Research

2021-05-13 · Noemi Derzsy, Subhabrata Majumdar, Rajat Malik

There is an increasing interest in ensuring machine learning (ML) frameworks behave in a socially responsible manner and are deemed trustworthy. Although considerable progress has been made in the field of Trustworthy ML…

BIG-bench Machine LearningCommunity Detection

Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining

2022-12-13 · Florian Tramèr, Gautam Kamath, Nicholas Carlini

The performance of differentially private machine learning can be boosted significantly by leveraging the transfer learning capabilities of non-private models pretrained on large public datasets. We critically review thi…

PositionPrivacy PreservingTransfer Learning

Collecting Verified COVID-19 Question Answer Pairs

2020-12-01 · EMNLP (NLP-COVID19) 2020 12 · Adam Poliak, Max Fleming, Cash Costello, Kenton Murray 외

We release a dataset of over 2,100 COVID19 related Frequently asked Question-Answer pairs scraped from over 40 trusted websites. We include an additional 24, 000 questions pulled from online sources that have been aligne…

Chatbot

TeSum: Human-Generated Abstractive Summarization Corpus for Telugu

2022-06-01 · LREC 2022 6 · Ashok Urlana, Nirmal Surange, Pavan Baswani, Priyanka Ravva 외

Expert human annotation for summarization is definitely an expensive task, and can not be done on huge scales. But with this work, we show that even with a crowd sourced summary generation approach, quality can be contro…

Abstractive Text Summarization

MegaWika 2: A More Comprehensive Multilingual Collection of Articles and their Sources

2025-08-05 · Samuel Barham, Chandler May, Benjamin Van Durme arxiv

We introduce MegaWika 2, a large, multilingual dataset of Wikipedia articles with their citations and scraped web sources; articles are represented in a rich data structure, and scraped source texts are stored inline wit…

Question AnsweringFact Checking