Towards JointUD: Part-of-speech Tagging and Lemmatization using Recurrent Neural Networks
This paper describes our submission to CoNLL 2018 UD Shared Task. We have extended an LSTM-based neural network designed for sequence tagging to additionally generate character-level sequences. The network was jointly trained to produce lemmas, part-of-speech tags and morphological features. Sentence segmentation, tokenization and dependency parsing were handled by UDPipe 1.2 baseline. The results demonstrate the viability of the proposed multitask architecture, although its performance still remains far from state-of-the-art.
Code (1)
Tasks
Dependency ParsingLemmatizationPart-Of-Speech TaggingSentenceSentence segmentationSimilar Papers 제목 키워드 기반
Transformer-based Part-of-Speech Tagging and Lemmatization for Latin
The paper presents a submission to the EvaLatin 2022 shared task. Our system places first for lemmatization, part-of-speech and morphological tagging in both closed and open modalities. The results for cross-genre and cr…
LemmatizationMorphological TaggingPart-Of-Speech TaggingeFontes. Part of Speech Tagging and Lemmatization of Medieval Latin Texts.A Cross-Genre Survey
This study introduces the eFontes models for automatic linguistic annotation of Medieval Latin texts, focusing on lemmatization, part-of-speech tagging, and morphological feature determination. Using the Transformers lib…
Lemmatizationnamed-entity-recognitionNamed Entity RecognitionPart-Of-Speech TaggingMorphological Tagging and Lemmatization of Albanian: A Manually Annotated Corpus and Neural Models
In this paper, we present the first publicly available part-of-speech and morphologically tagged corpus for the Albanian language, as well as a neural morphological tagger and lemmatizer trained on it. There is currently…
LemmatizationMorphological TaggingPart-Of-Speech TaggingCzech Text Processing with Contextual Embeddings: POS Tagging, Lemmatization, Parsing and NER
Contextualized embeddings, which capture appropriate word meaning depending on context, have recently been proposed. We evaluate two meth ods for precomputing such embeddings, BERT and Flair, on four Czech text processin…
Dependency ParsingLemmatizationnamed-entity-recognitionNamed Entity Recognition+5Under-resourced studies of under-resourced languages: lemmatization and POS-tagging with LLM annotators for historical Armenian, Georgian, Greek and Syriac
Low-resource languages pose persistent challenges for Natural Language Processing tasks such as lemmatization and part-of-speech (POS) tagging. This paper investigates the capacity of recent large language models (LLMs),…