Cross-lingual and Supervised Models for Morphosyntactic Annotation: a Comparison on Romanian
Because of the small size of Romanian corpora, the performance of a PoS tagger or a dependency parser trained with the standard supervised methods fall far short from the performance achieved in most languages. That is why, we apply state-of-the-art methods for cross-lingual transfer on Romanian tagging and parsing, from English and several Romance languages. We compare the performance with monolingual systems trained with sets of different sizes and establish that training on a few sentences in target language yields better results than transferring from large datasets in other languages.
Code (0)
등록된 구현이 없습니다.
Tasks
Cross-Lingual TransferPOSSimilar Papers 제목 키워드 기반
CorefUD 1.0: Coreference Meets Universal Dependencies
Recent advances in standardization for annotated language resources have led to successful large scale efforts, such as the Universal Dependencies (UD) project for multilingual syntactically annotated data. By comparison…
coreference-resolutionCoreference Resolutionnamed-entity-recognitionNamed Entity Recognition+1Investigating Multilingual Coreference Resolution by Universal Annotations
Multilingual coreference resolution (MCR) has been a long-standing and challenging task. With the newly proposed multilingual coreference dataset, CorefUD (Nedoluzhko et al., 2022), we conduct an investigation into the t…
coreference-resolutionCoreference ResolutionEnhancing the PARSEME Turkish Corpus of Verbal Multiword Expressions
The PARSEME (Parsing and Multiword Expressions) project proposes multilingual corpora annotated for multiword expressions (MWEs). In this case study, we focus on the Turkish corpus of PARSEME. Turkish is an agglutinative…
A Joint Matrix Factorization Analysis of Multilingual Representations
We present an analysis tool based on joint matrix factorization for comparing latent representations of multilingual and monolingual models. An alternative to probing, this tool allows us to analyze multiple sets of repr…
LLM Probe: Evaluating LLMs for Low-Resource Languages
Despite rapid advances in large language models (LLMs), their linguistic abilities in low-resource and morphologically rich languages are still not well understood due to limited annotated resources and the absence of st…
Speech Recognition