paper-with-me

홈 › Papers

Second language Korean Universal Dependency treebank v1.2: Focus on data augmentation and annotation scheme refinement

2025-03-18 · Hakyung Sung, Gyu-Ho Shin

We expand the second language (L2) Korean Universal Dependencies (UD) treebank with 5,454 manually annotated sentences. The annotation guidelines are also revised to better align with the UD framework. Using this enhanced treebank, we fine-tune three Korean language models and evaluate their performance on in-domain and out-of-domain L2-Korean datasets. The results show that fine-tuning significantly improves their performance across various metrics, thus highlighting the importance of using well-tailored L2 datasets for fine-tuning first-language-based, general-purpose language models for the morphosyntactic analysis of L2 data.

📄 PDF Abstract BibTeX arXiv:2503.14718

Code (1)

nlpxl2korean/ud-ksl 공식 구현

Tasks

Data Augmentation

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Building Universal Dependency Treebanks in Korean

2018-05-01 · LREC 2018 5 · Jayeol Chun, Na-Rae Han, Jena D. Hwang, Jinho D. Choi
Dependency Parsing

Yet Another Format of Universal Dependencies for Korean

2022-09-20 · COLING 2022 10 · Yige Chen, Eunkyul Leah Jo, Yundong Yao, Kyungtae Lim 외

In this study, we propose a morpheme-based scheme for Korean dependency parsing and adopt the proposed scheme to Universal Dependencies. We present the linguistic rationale that illustrates the motivation and the necessi…

Dependency Parsing

Analysis of the Penn Korean Universal Dependency Treebank (PKT-UD): Manual Revision to Build Robust Parsing Model in Korean

2020-05-26 · WS 2020 7 · Tae Hwan Oh, Ji Yoon Han, Hyonsu Choe, Seokwon Park 외

In this paper, we first open on important issues regarding the Penn Korean Universal Treebank (PKT-UD) and address these issues by revising the entire corpus manually with the aim of producing cleaner UD annotations that…

Annotation Issues in Universal Dependencies for Korean and Japanese

2020-12-01 · UDW (COLING) 2020 12 · Ji Yoon Han, Tae Hwan Oh, Lee Jin, Hansaem Kim

To investigate issues that arise in the process of developing a Universal Dependency (UD) treebank for Korean and Japanese, we begin by addressing the typological characteristics of Korean and Japanese. Both Korean and J…

UD-KSL Treebank v1.3: A semi-automated framework for aligning XPOS-extracted units with UPOS tags

2025-06-10 · Hakyung Sung, Gyu-Ho Shin, Chanyoung Lee, You Kyung Sung 외

The present study extends recent work on Universal Dependencies annotations for second-language (L2) Korean by introducing a semi-automated framework that identifies morphosyntactic constructions from XPOS sequences and …

Dependency Parsing