paper-with-me

Papers

Cross-lingual and Supervised Models for Morphosyntactic Annotation: a Comparison on Romanian

2016-05-01 · LREC 2016 5 · Lauriane Aufrant, Guillaume Wisniewski, Fran{\c{c}}ois Yvon

Because of the small size of Romanian corpora, the performance of a PoS tagger or a dependency parser trained with the standard supervised methods fall far short from the performance achieved in most languages. That is why, we apply state-of-the-art methods for cross-lingual transfer on Romanian tagging and parsing, from English and several Romance languages. We compare the performance with monolingual systems trained with sets of different sizes and establish that training on a few sentences in target language yields better results than transferring from large datasets in other languages.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Lingual TransferPOS

Similar Papers 제목 키워드 기반

CorefUD 1.0: Coreference Meets Universal Dependencies

2022-06-01 · LREC 2022 6 · Anna Nedoluzhko, Michal Novák, Martin Popel, Zdeněk Žabokrtský 외

Recent advances in standardization for annotated language resources have led to successful large scale efforts, such as the Universal Dependencies (UD) project for multilingual syntactically annotated data. By comparison…

coreference-resolutionCoreference Resolutionnamed-entity-recognitionNamed Entity Recognition+1

Investigating Multilingual Coreference Resolution by Universal Annotations

2023-10-26 · Haixia Chai, Michael Strube

Multilingual coreference resolution (MCR) has been a long-standing and challenging task. With the newly proposed multilingual coreference dataset, CorefUD (Nedoluzhko et al., 2022), we conduct an investigation into the t…

coreference-resolutionCoreference Resolution

Enhancing the PARSEME Turkish Corpus of Verbal Multiword Expressions

2022-06-01 · LREC (MWE) 2022 6 · Yagmur Ozturk, Najet Hadj Mohamed, Adam Lion-Bouton, Agata Savary

The PARSEME (Parsing and Multiword Expressions) project proposes multilingual corpora annotated for multiword expressions (MWEs). In this case study, we focus on the Turkish corpus of PARSEME. Turkish is an agglutinative…

A Joint Matrix Factorization Analysis of Multilingual Representations

2023-10-24 · Zheng Zhao, Yftah Ziser, Bonnie Webber, Shay B. Cohen

We present an analysis tool based on joint matrix factorization for comparing latent representations of multilingual and monolingual models. An alternative to probing, this tool allows us to analyze multiple sets of repr…

LLM Probe: Evaluating LLMs for Low-Resource Languages

2026-03-31 · Hailay Kidu Teklehaymanot, Gebrearegawi Gebremariam, Wolfgang Nejdl arxiv

Despite rapid advances in large language models (LLMs), their linguistic abilities in low-resource and morphologically rich languages are still not well understood due to limited annotated resources and the absence of st…

Speech Recognition