paper-with-me

홈 › Papers

Exploiting Cross-Dialectal Gold Syntax for Low-Resource Historical Languages: Towards a Generic Parser for Pre-Modern Slavic

2020-11-12 · Nilo Pedrazzini

This paper explores the possibility of improving the performance of specialized parsers for pre-modern Slavic by training them on data from different related varieties. Because of their linguistic heterogeneity, pre-modern Slavic varieties are treated as low-resource historical languages, whereby cross-dialectal treebank data may be exploited to overcome data scarcity and attempt the training of a variety-agnostic parser. Previous experiments on early Slavic dependency parsing are discussed, particularly with regard to their ability to tackle different orthographic, regional and stylistic features. A generic pre-modern Slavic parser and two specialized parsers -- one for East Slavic and one for South Slavic -- are trained using jPTDP (Nguyen & Verspoor 2018), a neural network model for joint part-of-speech (POS) tagging and dependency parsing which had shown promising results on a number of Universal Dependency (UD) treebanks, including Old Church Slavonic (OCS). With these experiments, a new state of the art is obtained for both OCS (83.79\% unlabelled attachment score (UAS) and 78.43\% labelled attachement score (LAS)) and Old East Slavic (OES) (85.7\% UAS and 80.16\% LAS).

📄 PDF Abstract BibTeX arXiv:2011.06467

Code (1)

npedrazzini/jPTDP-Early-Slavic 공식 구현

Tasks

Dependency ParsingPart-Of-Speech TaggingPOSPOS Tagging

Similar Papers 제목 키워드 기반

Multi-VALUE: A Framework for Cross-Dialectal English NLP

2022-12-15 · Caleb Ziems, William Held, Jingfeng Yang, Jwala Dhamala 외

Dialect differences caused by regional, social, and economic factors cause performance discrepancies for many groups of language technology users. Inclusive and equitable language technology must critically be dialect in…

Data AugmentationMachine TranslationQuestion AnsweringSemantic Parsing+1

A Catalog of Basque Dialectal Resources: Online Collections and Standard-to-Dialectal Adaptations

2026-03-26 · Jaione Bengoetxea, Itziar Gonzalez-Dios, Rodrigo Agerri arxiv

Recent research on dialectal NLP has identified data scarcity as a primary limitation. To address this limitation, this paper presents a catalog of contemporary Basque dialectal data and resources, offering a systematic …

Natural Language Inference

Linear Semantic Segmentation for Low-Resource Spoken Dialects

2026-05-07 · Kirill Chirkunov, Younes Samih, Abed Alhakim Freihat, Hanan Aldarmaki arxiv

Semantic segmentation is a core component of discourse analysis, yet existing models are primarily developed and evaluated on high-resource written text, limiting their effectiveness on low-resource spoken varieties. In …

Semantic Segmentation

Benchmarking Bengali Dialectal Bias: A Multi-Stage Framework Integrating RAG-Based Translation and Human-Augmented RLAIF

2026-03-22 · K. M. Jubair Sami, Dipto Sumit, Ariyan Hossain, Farig Sadeque arxiv

Large language models (LLMs) frequently exhibit performance biases against regional dialects of low-resource languages. However, frameworks to quantify these disparities remain scarce. We propose a two-phase framework to…

Content-Localization based Neural Machine Translation for Informal Dialectal Arabic: Spanish/French to Levantine/Gulf Arabic

2023-12-12 · Fatimah Alzamzami, Abdulmotaleb El Saddik

Resources in high-resource languages have not been efficiently exploited in low-resource languages to solve language-dependent research problems. Spanish and French are considered high resource languages in which an adeq…

Machine TranslationTranslation