paper-with-me

홈 › Papers

UD-KSL Treebank v1.3: A semi-automated framework for aligning XPOS-extracted units with UPOS tags

2025-06-10 · Hakyung Sung, Gyu-Ho Shin, Chanyoung Lee, You Kyung Sung, Boo Kyung Jung

The present study extends recent work on Universal Dependencies annotations for second-language (L2) Korean by introducing a semi-automated framework that identifies morphosyntactic constructions from XPOS sequences and aligns those constructions with corresponding UPOS categories. We also broaden the existing L2-Korean corpus by annotating 2,998 new sentences from argumentative essays. To evaluate the impact of XPOS-UPOS alignments, we fine-tune L2-Korean morphosyntactic analysis models on datasets both with and without these alignments, using two NLP toolkits. Our results indicate that the aligned dataset not only improves consistency across annotation layers but also enhances morphosyntactic tagging and dependency-parsing accuracy, particularly in cases of limited annotated data.

📄 PDF Abstract BibTeX arXiv:2506.09009

Code (0)

등록된 구현이 없습니다.

Tasks

Dependency Parsing

Similar Papers 제목 키워드 기반

A Universal Dependencies Treebank of Ancient Hebrew

2022-06-01 · LREC 2022 6 · Daniel Swanson, Francis Tyers

In this paper we present the initial construction of a Universal Dependencies treebank with morphological annotations of Ancient Hebrew containing portions of the Hebrew Scriptures (1579 sentences, 27K tokens) for use in…

Aligning the Romanian Reference Treebank and the Valence Lexicon of Romanian Verbs

2022-06-01 · LREC 2022 6 · Ana-Maria Barbu, Verginica Barbu Mititelu, Cătălin Mititelu

We present here the efforts of aligning two language resources for Romanian: the Romanian Reference Treebank and the Valence Lexicon of Romanian Verbs: for each occurrence of those verbs in the treebank that were include…

Bengali and Magahi PUD Treebank and Parser

2022-06-01 · WILDRE (LREC) 2022 6 · Pritha Majumdar, Deepak Alok, Akanksha Bansal, Atul Kr. Ojha 외

This paper presents the development of the Parallel Universal Dependency (PUD) Treebank for two Indo-Aryan languages: Bengali and Magahi. A treebank of 1,000 sentences has been created using a parallel corpus of English …

A New Latin Treebank for Universal Dependencies: Charters between Ancient Latin and Romance Languages

2020-05-01 · LREC 2020 5 · Flavio Massimiliano Cecchini, Timo Korkiakangas, Marco Passarotti

The present work introduces a new Latin treebank that follows the Universal Dependencies (UD) annotation standard. The treebank is obtained from the automated conversion of the Late Latin Charter Treebank 2 (LLCT2), orig…

Aligning the Norwegian UD Treebank with Entity and Coreference Information

2023-05-22 · Tollef Emil Jørgensen, Andre Kåsen

This paper presents a merged collection of entity and coreference annotated data grounded in the Universal Dependencies (UD) treebanks for the two written forms of Norwegian: Bokm{\aa}l and Nynorsk. The aligned and conve…