paper-with-me

Papers

Multi-Layer Discourse Annotation of a Dutch Text Corpus

2012-05-01 · LREC 2012 5 · Gisela Redeker, Ildik{\'o} Berzl{\'a}novich, Nynke van der Vliet, Gosse Bouma, Markus Egg

We have compiled a corpus of 80 Dutch texts from expository and persuasive genres, which we annotated for rhetorical and genre-specific discourse structure, and lexical cohesion with the goal of creating a gold standard for further research. The annota{\^A}{\neg}tions are based on a segmentation of the text in elementary discourse units that takes into account cues from syntax and punctuation. During the labor-intensive discourse-structure annotation (RST analysis), we took great care to thoroughly reconcile the initial analyses. That process and the availability of two independent initial analyses for each text allows us to analyze our disagreements and to assess the confusability of RST relations, and thereby improve the annotation guidelines and gather evidence for the classification of these relations into larger groups. We are using this resource for corpus-based studies of discourse relations, discourse markers, cohesion, and genre differences, e.g., the question of how discourse structure and lexical cohesion interact for different genres in the overall organization of texts. We are also exploring automatic text segmentation and semi-automatic discourse annotation.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

SegmentationText Segmentation

Similar Papers 제목 키워드 기반

Toward Dialogue Modeling: A Semantic Annotation Scheme for Questions and Answers

2019-08-23 · WS 2019 8 · Maria-Andrea Cruz-Blandón, Gosse Minnema, Aria Nourbakhsh, Maria Boritchev 외

The present study proposes an annotation scheme for classifying the content and discourse contribution of question-answer pairs. We propose detailed guidelines for using the scheme and apply them to dialogues in English,…

BIG-bench Machine Learning

DALC: the Dutch Abusive Language Corpus

2021-08-01 · ACL (WOAH) 2021 8 · Tommaso Caselli, Arjan Schelhaas, Marieke Weultjes, Folkert Leistra 외

As socially unacceptable language become pervasive in social media platforms, the need for automatic content moderation become more pressing. This contribution introduces the Dutch Abusive Language Corpus (DALC v1.0), a …

Abusive LanguageBinary ClassificationClassification

The Automatic Identification of Discourse Units in Dutch Text

2013-05-01 · WS 2013 5 · Nynke van der Vliet, Gosse Bouma, Gisela Redeker

Introducing Frege to Fillmore: A FrameNet Dataset that Captures both Sense and Reference

2022-06-01 · LREC 2022 6 · Levi Remijnse, Piek Vossen, Antske Fokkens, Sam Titarsolej

This article presents the first output of the Dutch FrameNet annotation tool, which facilitates both referential- and frame annotations of language-independent corpora. On the referential level, the tool links in-text me…

Sentence

Annotation and Analysis of Discourse Relations, Temporal Relations and Multi-Layered Situational Relations in Japanese Texts

2016-12-01 · WS 2016 12 · Kimi Kaneko, Saku Sugawara, Koji Mineshima, Daisuke Bekki

This paper proposes a methodology for building a specialized Japanese data set for recognizing temporal relations and discourse relations. In addition to temporal and discourse relations, multi-layered situational relati…

ArticlesNatural Language InferenceQuestion AnsweringText Summarization