Chunking Historical German
Quantitative studies of historical syntax require large amounts of syntactically annotated data, which are rarely available. The application of NLP methods could reduce manual annotation effort, provided that they achieve sufficient levels of accuracy. The present study investigates the automatic identification of chunks in historical German texts. Because no training data exists for this task, chunks are extracted from modern and historical constituency treebanks and used to train a CRF-based neural sequence labeling tool. The evaluation shows that the neural chunker outperforms an unlexicalized baseline and achieves overall F-scores between 90% and 94% for different historical data sets when POS tags are used as feature. The conducted experiments demonstrate the usefulness of including historical training data while also highlighting the importance of reducing boundary errors to improve annotation precision.
Code (1)
Tasks
ChunkingPOSSimilar Papers 제목 키워드 기반
Chunking German Legal Code
This paper investigates chunking strategies for retrieval-augmented generation on German statutory law, using the German Civil Code as a structured benchmark corpus. We implement and compare a range of segmentation appro…
Computational EfficiencyInformation RetrievalSequence Labeling: A Practical Approach
We take a practical approach to solving sequence labeling problem assuming unavailability of domain expertise and scarcity of informational and computational resources. To this end, we utilize a universal end-to-end Bi-L…
ChunkingNERPOSPOS TaggingDiachronic Analysis of German Parliamentary Proceedings: Ideological Shifts through the Lens of Political Biases
We analyze bias in historical corpora as encoded in diachronic distributional semantic models by focusing on two specific forms of bias, namely a political (i.e., anti-communism) and racist (i.e., antisemitism) one. For …
Diachronic Word EmbeddingsWord EmbeddingsGATEtoGerManC: A GATE-based Annotation Pipeline for Historical German
We describe a new GATE-based linguistic annotation pipeline for Early Modern German, which can be used to annotate historical texts with word tokens, sentence boundaries, lemmas, and POS tags. The pipeline is based on a …
POSPOS TaggingSentenceEmotion Classification in German Plays with Transformer-based Language Models Pretrained on Historical and Contemporary Language
We present results of a project on emotion classification on historical German plays of Enlightenment, Storm and Stress, and German Classicism. We have developed a hierarchical annotation scheme consisting of 13 sub-emot…
ClassificationEmotion ClassificationEmotion Classification in German