paper-with-me

홈 › Papers

Clustering-Based Article Identification in Historical Newspapers

2019-06-01 · WS 2019 6 · Martin Riedl, Daniela Betz, Sebastian Pad{\'o}

This article focuses on the problem of identifying articles and recovering their text from within and across newspaper pages when OCR just delivers one text file per page. We frame the task as a segmentation plus clustering step. Our results on a sample of 1912 New York Tribune magazine shows that performing the clustering based on similarities computed with word embeddings outperforms a similarity measure based on character n-grams and words. Furthermore, the automatic segmentation based on the text results in low scores, due to the low quality of some OCRed documents.

📄 PDF Abstract BibTeX

Code (1)

riedlma/cluster_identification 공식 구현

Tasks

ArticlesClusteringOptical Character Recognition (OCR)SegmentationWord Embeddings

Similar Papers 제목 키워드 기반

Word Clustering for Historical Newspapers Analysis

2019-09-01 · RANLP 2019 9 · Lidia Pivovarova, Elaine Zosa, Jani Marjanen

This paper is a part of a collaboration between computer scientists and historians aimed at development of novel tools and methods to improve analysis of historical newspapers. We present a case study of ideological term…

Clustering

American Stories: A Large-Scale Structured Text Dataset of Historical U.S. Newspapers

2023-08-24 · NeurIPS 2023 11 · Melissa Dell, Jacob Carlson, Tom Bryan, Emily Silcock 외

Existing full text datasets of U.S. public domain newspapers do not recognize the often complex layouts of newspaper scans, and as a result the digitized content scrambles texts from articles, headlines, captions, advert…

ArticlesLanguage ModelingLanguage ModellingLarge Language Model+3

Southern Newswire Corpus: A Large-Scale Dataset of Mid-Century Wire Articles Beyond the Front Page

2025-02-17 · Michael McRae

I introduce a new large-scale dataset of historical wire articles from U.S. Southern newspapers, spanning 1960-1975 and covering multiple wire services: The Associated Press, United Press International, Newspaper Enterpr…

ArticlesOptical Character Recognition (OCR)

Combining Visual and Textual Features for Semantic Segmentation of Historical Newspapers

2020-02-14 · Raphaël Barman, Maud Ehrmann, Simon Clematide, Sofia Ares Oliveira 외

The massive amounts of digitized historical documents acquired over the last decades naturally lend themselves to automatic processing and exploration. Research work seeking to automatically process facsimiles and extrac…

Document Layout AnalysisSemantic Segmentation

Media Manipulations in the Coverage of Events of the Ukrainian Revolution of Dignity: Historical, Linguistic, and Psychological Approaches

2024-07-09 · Ivan Khoma, Solomia Fedushko, Zoryana Kunch

This article examines the use of manipulation in the coverage of events of the Ukrainian Revolution of Dignity in the mass media, namely in the content of the online newspaper Ukrainian Truth (Ukrainska pravda), online n…