paper-with-me

Papers

Automatic Orality Identification in Historical Texts

2020-05-01 · LREC 2020 5 · Katrin Ortmann, Stefanie Dipper

Independently of the medial representation (written/spoken), language can exhibit characteristics of conceptual orality or literacy, which mainly manifest themselves on the lexical or syntactic level. In this paper we aim at automatically identifying conceptually-oral historical texts, with the ultimate goal of gaining knowledge about spoken data of historical time stages. We apply a set of general linguistic features that have been proven to be effective for the classification of modern language data to historical German texts from various registers. Many of the features turn out to be equally useful in determining the conceptuality of historical data as they are for modern data, especially the frequency of different types of pronouns and the ratio of verbs to nouns. Other features like sentence length, particles or interjections point to peculiarities of the historical data and reveal problems with the adoption of a feature set that was developed on modern language data.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

A Study of Temporal Fusion Strategies for Named Entity Recognition in Historical Texts

2026-06-26 · Emanuela Boros arxiv

Temporal variation poses a unique challenge for named entity recognition (NER) in historical texts, where entities drift in surface form and salience across time. While language models (LMs) have made progress in various…

Variation between Different Discourse Types: Literate vs. Oral

2019-06-01 · WS 2019 6 · Katrin Ortmann, Stefanie Dipper

This paper deals with the automatic identification of literate and oral discourse in German texts. A range of linguistic features is selected and their role in distinguishing between literate- and oral-oriented registers…

Sentence

Towards Few-Shot Identification of Morality Frames using In-Context Learning

2023-02-03 · Shamik Roy, Nishanth Sridhar Nakshatri, Dan Goldwasser

Data scarcity is a common problem in NLP, especially when the annotation pertains to nuanced socio-linguistic concepts that require specialized knowledge. As a result, few-shot identification of these concepts is desirab…

In-Context Learning

Automatic Topological Field Identification in (Historical) German Texts

2020-12-01 · COLING (LaTeCHCLfL, CLFL, LaTeCH) 2020 12 · Katrin Ortmann

For the study of certain linguistic phenomena and their development over time, large amounts of textual data must be enriched with relevant annotations. Since the manual creation of such annotations requires a lot of eff…

Sentence

Chunking Historical German

2021-05-01 · NoDaLiDa 2021 5 · Katrin Ortmann

Quantitative studies of historical syntax require large amounts of syntactically annotated data, which are rarely available. The application of NLP methods could reduce manual annotation effort, provided that they achiev…

ChunkingPOS