paper-with-me

홈 › Papers

AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters

2024-01-12 · Li Lucy, Suchin Gururangan, Luca Soldaini, Emma Strubell, David Bamman, Lauren F. Klein, Jesse Dodge

Large language models' (LLMs) abilities are drawn from their pretraining data, and model development begins with data curation. However, decisions around what data is retained or removed during this initial stage are under-scrutinized. In our work, we ground web text, which is a popular pretraining data source, to its social and geographic contexts. We create a new dataset of 10.3 million self-descriptions of website creators, and extract information about who they are and where they are from: their topical interests, social roles, and geographic affiliations. Then, we conduct the first study investigating how ten "quality" and English language identification (langID) filters affect webpages that vary along these social dimensions. Our experiments illuminate a range of implicit preferences in data curation: we show that some quality classifiers act like topical domain filters, and langID can overlook English content from some regions of the world. Overall, we hope that our work will encourage a new line of research on pretraining data curation practices and its social implications.

📄 PDF Abstract BibTeX arXiv:2401.06408

Code (1)

lucy3/whos_filtered 공식 구현

Tasks

Language Identification

Similar Papers 제목 키워드 기반

Predicting Visual Attention in Graphic Design Documents

2024-07-02 · Souradeep Chakraborty, Zijun Wei, Conor Kelton, Seoyoung Ahn 외

We present a model for predicting visual attention during the free viewing of graphic design documents. While existing works on this topic have aimed at predicting static saliency of graphic designs, our work is the firs…

Scanpath prediction

Online information on medical cannabis may rise unrealistic expectations and downplay potential side effects

2020-04-05 · Arthur Cassa Macedo, André Oliveira Vilela de Faria, Isabella Bizzi, Fabrício A. Moreira 외

There is a growing literature on the potential medical uses of Cannabis sativa and cannabinoid compounds. Although these have only been approved by regulatory agencies for few indications, there is a hype about their pos…

Exploring Paracrawl for Document-level Neural Machine Translation

2023-04-20 · Yusser Al Ghussin, Jingyi Zhang, Josef van Genabith

Document-level neural machine translation (NMT) has outperformed sentence-level NMT on a number of datasets. However, document-level NMT is still not widely adopted in real-world translation systems mainly due to the lac…

Machine TranslationNMTSentenceTranslation

Hierarchical Multimodal Pre-training for Visually Rich Webpage Understanding

2024-02-28 · Hongshen Xu, Lu Chen, Zihan Zhao, Da Ma 외

The growing prevalence of visually rich documents, such as webpages and scanned/digital-born documents (images, PDFs, etc.), has led to increased interest in automatic document understanding and information extraction ac…

document understandingInformation RetrievalRetrieval

Layout-aware Webpage Quality Assessment

2023-01-28 · Anfeng Cheng, Yiding Liu, Weibin Li, Qian Dong 외

Identifying high-quality webpages is fundamental for real-world search engines, which can fulfil users' information need with the less cognitive burden. Early studies of \emph{webpage quality assessment} usually design h…

Graph Neural Network