paper-with-me

홈 › Papers

A Treebank for the Healthcare Domain

2018-08-01 · COLING 2018 8 · Nganthoibi Oinam, Diwakar Mishra, Pinal Patel, Narayan Choudhary, Hitesh Desai

This paper presents a treebank for the healthcare domain developed at ezDI. The treebank is created from a wide array of clinical health record documents across hospitals. The data has been de-identified and annotated for constituent syntactic structure. The treebank contains a total of 52053 sentences that have been sampled for subdomains as well as linguistic variations. The paper outlines the sampling process followed to ensure a better domain representation in the corpus, the annotation process and challenges, and corpus statistics. The Penn Treebank tagset and guidelines were largely followed, but there were many syntactic contexts that warranted adaptation of the guidelines. The treebank created was used to re-train the Berkeley parser and the Stanford parser. These parsers were also trained with the GENIA treebank for comparative quality assessment. Our treebank yielded great-er accuracy on both parsers. Berkeley parser performed better on our treebank with an average F1 measure of 91 across 5-folds. This was a significant jump from the out-of-the-box F1 score of 70 on Berkeley parser{'}s default grammar.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SCTB: A Chinese Treebank in Scientific Domain

2016-12-01 · WS 2016 12 · Chenhui Chu, Toshiaki Nakazawa, Daisuke Kawahara, Sadao Kurohashi

Treebanks are curial for natural language processing (NLP). In this paper, we present our work for annotating a Chinese treebank in scientific domain (SCTB), to address the problem of the lack of Chinese treebanks in thi…

Chinese Word SegmentationMachine TranslationTranslation

An Arabic Dependency Treebank in the Travel Domain

2019-01-29 · Dima Taji, Jamila El Gizuli, Nizar Habash

In this paper we present a dependency treebank of travel domain sentences in Modern Standard Arabic. The text comes from a translation of the English equivalent sentences in the Basic Traveling Expressions Corpus. The tr…

Translation

Bidirectional Domain Adaptation Using Weighted Multi-Task Learning

2021-08-01 · ACL (IWPT) 2021 8 · Daniel Dakota, Zeeshan Ali Sayyed, Sandra Kübler

Domain adaption in syntactic parsing is still a significant challenge. We address the issue of data imbalance between the in-domain and out-of-domain treebank typically used for the problem. We define domain adaptation a…

Domain AdaptationMulti-Task Learning

Czech Legal Text Treebank 1.0

2016-05-01 · LREC 2016 5 · Vincent Kr{\'\i}{\v{z}}, Barbora Hladk{\'a}, Zde{\v{n}}ka Ure{\v{s}}ov{\'a}

We introduce a new member of the family of Prague dependency treebanks. The Czech Legal Text Treebank 1.0 is a morphologically and syntactically annotated corpus of 1,128 sentences. The treebank contains texts from the l…

Improving Domain Independent Question Parsing with Synthetic Treebanks

2018-08-01 · COLING 2018 8 · Halim-Antoine Boukaram, Nizar Habash, Micheline Ziadee, Majd Sakr

Automatic syntactic parsing for question constructions is a challenging task due to the paucity of training examples in most treebanks. The near absence of question constructions is due to the dominance of the news domai…