paper-with-me

홈 › Papers

Southern Newswire Corpus: A Large-Scale Dataset of Mid-Century Wire Articles Beyond the Front Page

2025-02-17 · Michael McRae

I introduce a new large-scale dataset of historical wire articles from U.S. Southern newspapers, spanning 1960-1975 and covering multiple wire services: The Associated Press, United Press International, Newspaper Enterprise Association. Unlike prior work focusing on front-page content, this dataset captures articles across the entire newspaper, offering broader insight into mid-century Southern coverage. The dataset includes a version that has undergone an LLM-based text cleanup pipeline to reduce OCR noise, enhancing its suitability for quantitative text analysis. Additionally, duplicate versions of articles are retained to enable analysis of editorial differences in language and framing across newspapers. Each article is tagged by wire service, facilitating comparative studies of editorial patterns across agencies. This resource opens new avenues for research in computational social science, digital humanities, and historical linguistics, providing a detailed perspective on how Southern newspapers relayed national and international news during a transformative period in American history. The dataset will be made available upon publication or request for research purposes.

📄 PDF Abstract BibTeX arXiv:2502.11866

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesOptical Character Recognition (OCR)

Methods 이 논문이 사용한 방법론

Golden Queue Managers 설명 없음
American 설명 없음

Similar Papers 제목 키워드 기반

Neural Machine Translation System using a Content-equivalently Translated Parallel Corpus for the Newswire Translation Tasks at WAT 2019

2019-11-01 · WS 2019 11 · Hideya Mino, Hitoshi Ito, Isao Goto, Ichiro Yamada 외

This paper describes NHK and NHK Engineering System (NHK-ES){'}s submission to the newswire translation tasks of WAT 2019 in both directions of Japanese→English and English→Japanese. In addition to the JIJI Corpus that w…

Machine TranslationSentenceTranslation

Web-scale Surface and Syntactic n-gram Features for Dependency Parsing

2015-02-25 · Dominick Ng, Mohit Bansal, James R. Curran

We develop novel first- and second-order features for dependency parsing based on the Google Syntactic Ngrams corpus, a collection of subtree counts of parsed sentences from scanned books. We also extend previous work on…

Dependency Parsing

Cross-Level Semantic Similarity for Serbian Newswire Texts

2022-06-01 · LREC 2022 6 · Vuk Batanović, Maja Miličević Petrović

Cross-Level Semantic Similarity (CLSS) is a measure of the level of semantic overlap between texts of different lengths. Although this problem was formulated almost a decade ago, research on it has been sparse, and limit…

Semantic SimilaritySemantic Textual SimilaritySentence

Introducing QuBERT: A Large Monolingual Corpus and BERT Model for Southern Quechua

2022-07-01 · DeepLo 2022 7 · Rodolfo Zevallos, John Ortega, William Chen, Richard Castro 외

t

Newswire: A Large-Scale Structured Database of a Century of Historical News

2024-06-13 · Emily Silcock, Abhishek Arora, Luca D'Amico-Wong, Melissa Dell

In the U.S. historically, local newspapers drew their content largely from newswires like the Associated Press. Historians argue that newswires played a pivotal role in creating a national identity and shared understandi…

ArticlesEntity DisambiguationLanguage ModelingLanguage Modelling+1