paper-with-me

홈 › Papers

RKorAPClient: An R Package for Accessing the German Reference Corpus DeReKo via KorAP

2020-05-01 · LREC 2020 5 · Marc Kupietz, Nils Diewald, Eliza Margaretha

Making corpora accessible and usable for linguistic research is a huge challenge in view of (too) big data, legal issues and a rapidly evolving methodology. This does not only affect the design of user-friendly graphical interfaces to corpus analysis tools, but also the availability of programming interfaces supporting access to the functionality of these tools from various analysis and development environments. RKorAPClient is a new research tool in the form of an R package that interacts with the Web API of the corpus analysis platform KorAP, which provides access to large annotated corpora, including the German reference corpus DeReKo with 45 billion tokens.In addition to optionally authenticated KorAP API access, RKorAPClient provides further processing and visualization features to simplify common corpus analysis tasks. This paper introduces the basic functionality of RKorAPClient and exemplifies various analysis tasks based on DeReKo, that are bundled within the R package and can serve as a basic framework for advanced analysis and visualization approaches.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bundesrecht: An Open Library and Corpus for German Statutory Reference Processing

2026-05-29 · Harshil Darji, Martin Heckelmann, Christina Kratsch, Gerard de Melo arxiv

Statutory references are central to legal language understanding, but are difficult to process automatically, as they appear in compact and variable surface forms, may combine multiple targets, use special abbreviations,…

Information Extraction

Automatic Annotation and Manual Evaluation of the Diachronic German Corpus T\"uBa-D/DC

2012-05-01 · LREC 2012 5 · Erhard Hinrichs, Thomas Zastrow

This paper presents the Tu{\`I}ˆbingen Baumbank des Deutschen Diachron (Tu{\`I}ˆBa-D/DC), a linguistically annotated corpus of selected diachronic materials from the German Gutenberg Project. It was automatically annotat…

A Tidy Data Model for Natural Language Processing using cleanNLP

2017-03-27 · Taylor Arnold

The package cleanNLP provides a set of fast tools for converting a textual corpus into a set of normalized tables. The underlying natural language processing pipeline utilizes Stanford's CoreNLP library, exposing a numbe…

coreference-resolutionCoreference ResolutionDependency ParsingEntity Linking+5

OpusTools and Parallel Corpus Diagnostics

2020-05-01 · LREC 2020 5 · Mikko Aulamo, Umut Sulubacak, Sami Virpioja, J{\"o}rg Tiedemann

This paper introduces OpusTools, a package for downloading and processing parallel corpora included in the OPUS corpus collection. The package implements tools for accessing compressed data in their archived release form…

Language Identification

SwissDial: Parallel Multidialectal Corpus of Spoken Swiss German

2021-03-21 · Pelin Dogan-Schönberger, Julian Mäder, Thomas Hofmann

Swiss German is a dialect continuum whose natively acquired dialects significantly differ from the formal variety of the language. These dialects are mostly used for verbal communication and do not have standard orthogra…

Speech Synthesis