paper-with-me

홈 › Papers

The IULA Treebank

2012-05-01 · LREC 2012 5 · Montserrat Marimon, Beatriz Fisas, N{\'u}ria Bel, Jorge Vivaldi, Sergi Torner, Merc{\`e} Lorente, Silvia V{\'a}zquez, Marta Villegas

This paper describes on-going work for the construction of a new treebank for Spanish, The IULA Treebank. This new resource will contain about 60,000 richly annotated sentences as an extension of the already existing IULA Technical Corpus which is only PoS tagged. In this paper we have focused on describing the work done for defining the annotation process and the treebank design principles. We report on how the used framework, the DELPH-IN processing framework, has been crucial in the design principles and in the bootstrapping strategy followed, especially in what refers to the use of stochastic modules for reducing parsing overgeneration. We also report on the different evaluation experiments carried out to guarantee the quality of the already available results.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

POS

Similar Papers 제목 키워드 기반

The IULA Spanish LSP Treebank

2014-05-01 · LREC 2014 5 · Montserrat Marimon, N{\'u}ria Bel, Beatriz Fisas, Blanca Arias 외

This paper presents the IULA Spanish LSP Treebank, a dependency treebank of over 41,000 sentences of different domains (Law, Economy, Computing Science, Environment, and Medicine), developed in the framework of the Europ…

Annotation of negation in the IULA Spanish Clinical Record Corpus

2017-04-01 · WS 2017 4 · Montserrat Marimon, Jorge Vivaldi, N{\'u}ria Bel

This paper presents the IULA Spanish Clinical Record Corpus, a corpus of 3,194 sentences extracted from anonymized clinical records and manually annotated with negation markers and their scope. The corpus was conceived a…

Medical DiagnosisNegationNegation DetectionTerm Extraction

Iula2Standoff: a tool for creating standoff documents for the IULACT

2012-05-01 · LREC 2012 5 · Carlos Morell, Jorge Vivaldi, N{\'u}ria Bel

Due to the increase in the number and depth of analyses required over the text, like entity recognition, POS tagging, syntactic analysis, etc. the annotation in-line has become unpractical. In Natural Language Processing…

LemmatizationPOSPOS Tagging

BKTreebank: Building a Vietnamese Dependency Treebank

2017-10-16 · LREC 2018 5 · Kiem-Hieu Nguyen

Dependency treebank is an important resource in any language. In this paper, we present our work on building BKTreebank, a dependency treebank for Vietnamese. Important points on designing POS tagset, dependency relation…

Dependency ParsingPOSPOS Tagging

SCTB: A Chinese Treebank in Scientific Domain

2016-12-01 · WS 2016 12 · Chenhui Chu, Toshiaki Nakazawa, Daisuke Kawahara, Sadao Kurohashi

Treebanks are curial for natural language processing (NLP). In this paper, we present our work for annotating a Chinese treebank in scientific domain (SCTB), to address the problem of the lack of Chinese treebanks in thi…

Chinese Word SegmentationMachine TranslationTranslation