paper-with-me

홈 › Papers

Analysis of the Penn Korean Universal Dependency Treebank (PKT-UD): Manual Revision to Build Robust Parsing Model in Korean

2020-05-26 · WS 2020 7 · Tae Hwan Oh, Ji Yoon Han, Hyonsu Choe, Seokwon Park, Han He, Jinho D. Choi, Na-Rae Han, Jena D. Hwang, Hansaem Kim

In this paper, we first open on important issues regarding the Penn Korean Universal Treebank (PKT-UD) and address these issues by revising the entire corpus manually with the aim of producing cleaner UD annotations that are more faithful to Korean grammar. For compatibility to the rest of UD corpora, we follow the UDv2 guidelines, and extensively revise the part-of-speech tags and the dependency relations to reflect morphological features and flexible word-order aspects in Korean. The original and the revised versions of PKT-UD are experimented with transformer-based parsing models using biaffine attention. The parsing model trained on the revised corpus shows a significant improvement of 3.0% in labeled attachment score over the model trained on the previous corpus. Our error analysis demonstrates that this revision allows the parsing model to learn relations more robustly, reducing several critical errors that used to be made by the previous model.

📄 PDF Abstract BibTeX arXiv:2005.12898

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Constituency Structure over Eojeol in Korean Treebanks

2025-12-27 · Jungyeul Park, Chulwoo Park arxiv

The design of Korean constituency treebanks raises a central representational question concerning the choice of terminal units. Although Korean words are morphologically complex, treating morphemes as constituency termin…

Constituency Parsing

Preparing Korean Data for the Shared Task on Parsing Morphologically Rich Languages

2013-09-06 · Jinho D. Choi

This document gives a brief description of Korean data prepared for the SPMRL 2013 shared task. A total of 27,363 sentences with 350,090 tokens are used for the shared task. All constituent trees are collected from the K…

Morphological Analysis

Second language Korean Universal Dependency treebank v1.2: Focus on data augmentation and annotation scheme refinement

2025-03-18 · Hakyung Sung, Gyu-Ho Shin

We expand the second language (L2) Korean Universal Dependencies (UD) treebank with 5,454 manually annotated sentences. The annotation guidelines are also revised to better align with the UD framework. Using this enhance…

Data Augmentation

Building Universal Dependency Treebanks in Korean

2018-05-01 · LREC 2018 5 · Jayeol Chun, Na-Rae Han, Jena D. Hwang, Jinho D. Choi
Dependency Parsing

Universal Dependencies for Arabic

2017-04-01 · WS 2017 4 · Dima Taji, Nizar Habash, Daniel Zeman

We describe the process of creating NUDAR, a Universal Dependency treebank for Arabic. We present the conversion from the Penn Arabic Treebank to the Universal Dependency syntactic representation through an intermediate …

Machine TranslationQuestion Answering