Character-level Supervision for Low-resource POS Tagging
Neural part-of-speech (POS) taggers are known to not perform well with little training data. As a step towards overcoming this problem, we present an architecture for learning more robust neural POS taggers by jointly training a hierarchical, recurrent model and a recurrent character-based sequence-to-sequence network supervised using an auxiliary objective. This way, we introduce stronger character-level supervision into the model, which enables better generalization to unseen words and provides regularization, making our encoding less prone to overfitting. We experiment with three auxiliary tasks: lemmatization, character-based word autoencoding, and character-based random string autoencoding. Experiments with minimal amounts of labeled data on 34 languages show that our new architecture outperforms a single-task baseline and, surprisingly, that, on average, raw text autoencoding can be as beneficial for low-resource POS tagging as using lemma information. Our neural POS tagger closes the gap to a state-of-the-art POS tagger (MarMoT) for low-resource scenarios by 43{\%}, even outperforming it on languages with templatic morphology, e.g., Arabic, Hebrew, and Turkish, by some margin.
Code (0)
등록된 구현이 없습니다.
Tasks
Feature EngineeringLEMMALemmatizationMulti-Task LearningPOSPOS TaggingSimilar Papers 제목 키워드 기반
Cross-lingual Character-Level Neural Morphological Tagging
Even for common NLP tasks, sufficient supervision is not available in many languages {--} morphological tagging is no exception. In the work presented here, we explore a transfer learning scheme, whereby we train charact…
Language ModelingLanguage ModellingMorphological TaggingPart-Of-Speech Tagging+1Cross-lingual, Character-Level Neural Morphological Tagging
Even for common NLP tasks, sufficient supervision is not available in many languages -- morphological tagging is no exception. In the work presented here, we explore a transfer learning scheme, whereby we train character…
Morphological TaggingTransfer LearningSiPOS: A Benchmark Dataset for Sindhi Part-of-Speech Tagging
In this paper, we introduce the SiPOS dataset for part-of-speech tagging in the low-resource Sindhi language with quality baselines. The dataset consists of more than 293K tokens annotated with sixteen universal part-of-…
Part-Of-Speech Taggingtext annotationHeidelberg-Boston @ SIGTYP 2024 Shared Task: Enhancing Low-Resource Language Analysis With Character-Aware Hierarchical Transformers
Historical languages present unique challenges to the NLP community, with one prominent hurdle being the limited resources available in their closed corpora. This work describes our submission to the constrained subtask …
LemmatizationMorphological TaggingPOSPOS TaggingWeakly Supervised POS Taggers Perform Poorly on Truly Low-Resource Languages
Part-of-speech (POS) taggers for low-resource languages which are exclusively based on various forms of weak supervision - e.g., cross-lingual transfer, type-level supervision, or a combination thereof - have been report…
Cross-Lingual TransferPOSPOS Tagging