A Gold Standard Dependency Treebank for Turkish
We introduce TWT; a new treebank for Turkish which consists of web and Wikipedia sentences that are annotated for segmentation, morphology, part-of-speech and dependency relations. To date, it is the largest publicly available human-annotated morpho-syntactic Turkish treebank in terms of the annotated word count. It is also the first large Turkish dependency treebank that has a dedicated Wikipedia section. We present the tagsets and the methodology that are used in annotating the treebank and also the results of the baseline experiments on Turkish dependency parsing with this treebank.
Code (0)
등록된 구현이 없습니다.
Tasks
Dependency ParsingSimilar Papers 제목 키워드 기반
Turkish Treebank as a Gold Standard for Morphological Disambiguation and Its Influence on Parsing
So far predicted scenarios for Turkish dependency parsing have used a morphological disambiguator that is trained on the data distributed with the tool(Sak et al., 2008). Although models trained on this data have high ac…
Dependency ParsingMorphological AnalysisMorphological DisambiguationThe Impact of Automatic Morphological Analysis \& Disambiguation on Dependency Parsing of Turkish
The studies on dependency parsing of Turkish so far gave their results on the Turkish Dependency Treebank. This treebank consists of sentences where gold standard part-of-speech tags are manually assigned to each word an…
Dependency ParsingInformation RetrievalMorphological AnalysisA Learning-Based Dependency to Constituency Conversion Algorithm for the Turkish Language
This study aims to create the very first dependency-to-constituency conversion algorithm optimised for Turkish language. For this purpose, a state-of-the-art morphologic analyser and a feature-based machine learning mode…
From Constituency to UD-Style Dependency: Building the First Conversion Tool of Turkish
This paper deliberates on the process of building the first constituency-to-dependency conversion tool of Turkish. The starting point of this work is a previous study in which 10,000 phrase structure trees were manually …
BIG-bench Machine LearningResources for Turkish Dependency Parsing: Introducing the BOUN Treebank and the BoAT Annotation Tool
In this paper, we introduce the resources that we developed for Turkish dependency parsing, which include a novel manually annotated treebank (BOUN Treebank), along with the guidelines we adopted, and a new annotation to…
ArticlesCultural Vocal Bursts Intensity PredictionDependency Parsing