paper-with-me

홈 › Papers

Word Clustering for Historical Newspapers Analysis

2019-09-01 · RANLP 2019 9 · Lidia Pivovarova, Elaine Zosa, Jani Marjanen

This paper is a part of a collaboration between computer scientists and historians aimed at development of novel tools and methods to improve analysis of historical newspapers. We present a case study of ideological terms ending with -ism suffix in nineteenth century Finnish newspapers. We propose a two-step procedure to trace differences in word usages over time: training of diachronic embeddings on several time slices and when clustering embeddings of selected words together with their neighbours to obtain historical context. The obtained clusters turn out to be useful for historical studies. The paper also discuss specific difficulties related to development historian-oriented tools.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Clustering

Similar Papers 제목 키워드 기반

A Quantitative Discourse Analysis of Asian Workers in the US Historical Newspapers

2024-02-04 · Jaihyun Park, Ryan Cordell

Warning: This paper contains examples of offensive language targetting marginalized population. The digitization of historical texts invites researchers to explore the large-scale corpus of historical texts with computat…

Clustering-Based Article Identification in Historical Newspapers

2019-06-01 · WS 2019 6 · Martin Riedl, Daniela Betz, Sebastian Pad{\'o}

This article focuses on the problem of identifying articles and recovering their text from within and across newspaper pages when OCR just delivers one text file per page. We frame the task as a segmentation plus cluster…

ArticlesClusteringOptical Character Recognition (OCR)Segmentation+1

Using Word Embeddings to Examine Gender Bias in Dutch Newspapers, 1950-1990

2019-07-21 · WS 2019 8 · Melvin Wevers

Contemporary debates on filter bubbles and polarization in public and social media raise the question to what extent news media of the past exhibited biases. This paper specifically examines bias related to gender in six…

Word Embeddings

Measuring Intersectional Biases in Historical Documents

2023-05-21 · Nadav Borenstein, Karolina Stańczak, Thea Rolskov, Natália da Silva Perez 외

Data-driven analyses of biases in historical texts can help illuminate the origin and development of biases prevailing in modern society. However, digitised historical documents pose a challenge for NLP practitioners as …

Optical Character RecognitionOptical Character Recognition (OCR)Word Embeddings

Combining Visual and Textual Features for Semantic Segmentation of Historical Newspapers

2020-02-14 · Raphaël Barman, Maud Ehrmann, Simon Clematide, Sofia Ares Oliveira 외

The massive amounts of digitized historical documents acquired over the last decades naturally lend themselves to automatic processing and exploration. Research work seeking to automatically process facsimiles and extrac…

Document Layout AnalysisSemantic Segmentation