paper-with-me

Papers

Leipzig Corpus Miner - A Text Mining Infrastructure for Qualitative Data Analysis

2017-07-11 · Andreas Niekler, Gregor Wiedemann, Gerhard Heyer

This paper presents the "Leipzig Corpus Miner", a technical infrastructure for supporting qualitative and quantitative content analysis. The infrastructure aims at the integration of 'close reading' procedures on individual documents with procedures of 'distant reading', e.g. lexical characteristics of large document collections. Therefore information retrieval systems, lexicometric statistics and machine learning procedures are combined in a coherent framework which enables qualitative data analysts to make use of state-of-the-art Natural Language Processing techniques on very large document collections. Applicability of the framework ranges from social sciences to media studies and market research. As an example we introduce the usage of the framework in a political science study on post-democracy and neoliberalism.

📄 PDF Abstract BibTeX arXiv:1707.03253

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalRetrieval

Similar Papers 제목 키워드 기반

Application of the interactive Leipzig Corpus Miner as a generic research platform for the use in the social sciences

2021-10-06 · Christian Kahmann, Andreas Niekler, Gregor Wiedemann

This article introduces to the interactive Leipzig Corpus Miner (iLCM) - a newly released, open-source software to perform automatic content analysis. Since the iLCM is based on the R-programming language, its generic te…

iLCM - A Virtual Research Infrastructure for Large-Scale Qualitative Data

2018-05-11 · LREC 2018 5 · Andreas Niekler, Arnim Bleier, Christian Kahmann, Lisa Posch 외

The iLCM project pursues the development of an integrated research environment for the analysis of structured and unstructured data in a "Software as a Service" architecture (SaaS). The research environment addresses req…

SiliconHealth: A Complete Low-Cost Blockchain Healthcare Infrastructure for Resource-Constrained Regions Using Repurposed Bitcoin Mining ASICs

2026-01-14 · Francisco Angulo de Lafuente, Seid Mehammed Abdu, Nirmal Tej arxiv

This paper presents SiliconHealth, a comprehensive blockchain-based healthcare infrastructure designed for resource-constrained regions, particularly sub-Saharan Africa. We demonstrate that obsolete Bitcoin mining Applic…

Semantic Retrieval

Uralic Language Identification (ULI) 2020 shared task dataset and the Wanca 2017 corpus

2020-08-27 · Tommi Jauhiainen, Heidi Jauhiainen, Niko Partanen, Krister Lindén

This article introduces the Wanca 2017 corpus of texts crawled from the internet from which the sentences in rare Uralic languages for the use of the Uralic Language Identification (ULI) 2020 shared task were collected. …

Language Identification

Building Large Monolingual Dictionaries at the Leipzig Corpora Collection: From 100 to 200 Languages

2012-05-01 · LREC 2012 5 · Dirk Goldhahn, Thomas Eckart, Uwe Quasthoff

The Leipzig Corpora Collection offers free online access to 136 monolingual dictionaries enriched with statistical information. In this paper we describe current advances of the project in collecting and processing text …

Lemmatization