A Neural Model for Part-of-Speech Tagging in Historical Texts
Historical texts are challenging for natural language processing because they differ linguistically from modern texts and because of their lack of orthographical and grammatical standardisation. We use a character-level neural network to build a part-of-speech (POS) tagger that can process historical data directly without requiring a separate spelling normalisation stage. Its performance in a Swedish verb identification and a German POS tagging task is similar to that of a two-stage model. We analyse the performance of this tagger and a more traditional baseline system, discuss some of the remaining problems for tagging historical data and suggest how the flexibility of our neural tagger could be exploited to address diachronic divergences in morphology and syntax in early modern Swedish with the help of data from closely related languages.
Code (0)
등록된 구현이 없습니다.
Tasks
Part-Of-Speech TaggingPOSPOS TaggingSimilar Papers 제목 키워드 기반
Part-of-Speech Tagging for Historical English
As more historical texts are digitized, there is interest in applying natural language processing tools to these archives. However, the performance of these tools is often unsatisfactory, due to language change and genre…
Domain AdaptationPart-Of-Speech TaggingUnsupervised Domain AdaptationWord EmbeddingsA Comparative Analysis of Word Segmentation, Part-of-Speech Tagging, and Named Entity Recognition for Historical Chinese Sources, 1900-1950
This paper compares large language models (LLMs) and traditional natural language processing (NLP) tools for performing word segmentation, part-of-speech (POS) tagging, and named entity recognition (NER) on Chinese texts…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+3ParsiPy: NLP Toolkit for Historical Persian Texts in Python
The study of historical languages presents unique challenges due to their complex orthographic systems, fragmentary textual evidence, and the absence of standardized digital representations of text in those languages. Ta…
LemmatizationPart-Of-Speech TaggingTransliterationAdapting a part-of-speech tagset to non-standard text: The case of STTS
The Stuttgart-T{\"u}bingen TagSet (STTS) is a de-facto standard for the part-of-speech tagging of German texts. Since its first publication in 1995, STTS has been used in a variety of annotation projects, some of which h…
Domain AdaptationMachine TranslationOpinion MiningPart-Of-Speech Tagging+2Reliable Part-of-Speech Tagging of Historical Corpora through Set-Valued Prediction
Syntactic annotation of corpora in the form of part-of-speech (POS) tags is a key requirement for both linguistic research and subsequent automated natural language processing (NLP) tasks. This problem is commonly tackle…
Part-Of-Speech TaggingPOSPOS TaggingTAG