paper-with-me

홈 › Papers

Chunking Historical German

2021-05-01 · NoDaLiDa 2021 5 · Katrin Ortmann

Quantitative studies of historical syntax require large amounts of syntactically annotated data, which are rarely available. The application of NLP methods could reduce manual annotation effort, provided that they achieve sufficient levels of accuracy. The present study investigates the automatic identification of chunks in historical German texts. Because no training data exists for this task, chunks are extracted from modern and historical constituency treebanks and used to train a CRF-based neural sequence labeling tool. The evaluation shows that the neural chunker outperforms an unlexicalized baseline and achieves overall F-scores between 90% and 94% for different historical data sets when POS tags are used as feature. The conducted experiments demonstrate the usefulness of including historical training data while also highlighting the importance of reducing boundary errors to improve annotation precision.

📄 PDF Abstract BibTeX

Code (1)

rubcompling/nodalida2021 공식 구현

Tasks

ChunkingPOS

Similar Papers 제목 키워드 기반

Chunking German Legal Code

2026-05-19 · Max Prior, Natalia Milanova, Andreas Schultz arxiv

This paper investigates chunking strategies for retrieval-augmented generation on German statutory law, using the German Civil Code as a structured benchmark corpus. We implement and compare a range of segmentation appro…

Computational EfficiencyInformation Retrieval

Sequence Labeling: A Practical Approach

2018-08-12 · Adnan Akhundov, Dietrich Trautmann, Georg Groh

We take a practical approach to solving sequence labeling problem assuming unavailability of domain expertise and scarcity of informational and computational resources. To this end, we utilize a universal end-to-end Bi-L…

ChunkingNERPOSPOS Tagging

Diachronic Analysis of German Parliamentary Proceedings: Ideological Shifts through the Lens of Political Biases

2021-08-13 · Tobias Walter, Celina Kirschner, Steffen Eger, Goran Glavaš 외

We analyze bias in historical corpora as encoded in diachronic distributional semantic models by focusing on two specific forms of bias, namely a political (i.e., anti-communism) and racist (i.e., antisemitism) one. For …

Diachronic Word EmbeddingsWord Embeddings

GATEtoGerManC: A GATE-based Annotation Pipeline for Historical German

2012-05-01 · LREC 2012 5 · Silke Scheible, Richard J. Whitt, Martin Durrell, Paul Bennett

We describe a new GATE-based linguistic annotation pipeline for Early Modern German, which can be used to annotate historical texts with word tokens, sentence boundaries, lemmas, and POS tags. The pipeline is based on a …

POSPOS TaggingSentence

Emotion Classification in German Plays with Transformer-based Language Models Pretrained on Historical and Contemporary Language

2021-11-01 · EMNLP (LaTeCHCLfL, CLFL, LaTeCH) 2021 11 · Thomas Schmidt, Katrin Dennerlein, Christian Wolff

We present results of a project on emotion classification on historical German plays of Enlightenment, Storm and Stress, and German Classicism. We have developed a hierarchical annotation scheme consisting of 13 sub-emot…

ClassificationEmotion ClassificationEmotion Classification in German