paper-with-me

홈 › Papers

Text Preprocessing and its Implications in a Digital Humanities Project

2021-09-01 · RANLP 2021 9 · Maria Kunilovskaya, Alistair Plum

This paper focuses on data cleaning as part of a preprocessing procedure applied to text data retrieved from the web. Although the importance of this early stage in a project using NLP methods is often highlighted by researchers, the details, general principles and techniques are usually left out due to consideration of space. At best, they are dismissed with a comment “The usual data cleaning and preprocessing procedures were applied”. More coverage is usually given to automatic text annotation such as lemmatisation, part-of-speech tagging and parsing, which is often included in preprocessing. In the literature, the term ‘preprocessing’ is used to refer to a wide range of procedures, from filtering and cleaning to data transformation such as stemming and numeric representation, which might create confusion. We argue that text preprocessing might skew original data distribution with regard to the metadata, such as types, locations and times of registered datapoints. In this paper we describe a systematic approach to cleaning text data mined by a data-providing company for a Digital Humanities (DH) project focused on cultural analytics. We reveal the types and amount of noise in the data coming from various web sources and estimate the changes in the size of the data associated with preprocessing. We also compare the results of a text classification experiment run on the raw and preprocessed data. We hope that our experience and approaches will help the DH community to diagnose the quality of textual data collected from the web and prepare it for further natural language processing.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Part-Of-Speech Taggingtext annotationtext-classificationText Classification

Similar Papers 제목 키워드 기반

Topic Modeling the Hàn diăn Ancient Classics

2017-02-02 · Colin Allen, Hongliang Luo, Jaimie Murdock, Jianghuai Pu 외

Ancient Chinese texts present an area of enormous challenge and opportunity for humanities scholars interested in exploiting computational methods to assist in the development of new insights and interpretations of cultu…

Philosophy

Flexible and Reliable Text Analytics in the Digital Humanities -- Some Methodological Considerations

2016-12-01 · WS 2016 12 · Jonas Kuhn

The availability of Language Technology Resources and Tools generates a considerable methodological potential in the Digital Humanities: aspects of research questions from the Humanities and Social Sciences can be addres…

Cultural Vocal Bursts Intensity Prediction

From Digital Humanities to Quantum Humanities: Potentials and Applications

2021-03-17 · Johanna Barzen

Quantum computers are becoming real. Therefore, it is promising to use their potentials in different applications areas, which includes research in the humanities. Due to an increasing amount of data that needs to be pro…

ClusteringFeature Engineering

CLARIAH in the Netherlands

2016-05-01 · LREC 2016 5 · Jan Odijk

I introduce CLARIAH in the Netherlands, which aims to contribute the Netherlands part of a Europe-wide humanities research infrastructure. I describe the digital turn in the humanities, the background and context of CLAR…

GutenTag: an NLP-driven Tool for Digital Humanities Research in the Project Gutenberg Corpus

2015-06-01 · WS 2015 6 · Julian Brooke, Adam Hammond, Graeme Hirst