paper-with-me

Papers

ISO-based Annotated Multilingual Parallel Corpus for Discourse Markers

2022-06-01 · LREC 2022 6 · Purificação Silvano, Mariana Damova, Giedrė Valūnaitė Oleškevičienė, Chaya Liebeskind, Christian Chiarcos, Dimitar Trajanov, Ciprian-Octavian Truică, Elena-Simona Apostol, Anna Baczkowska

Discourse markers carry information about the discourse structure and organization, and also signal local dependencies or epistemological stance of speaker. They provide instructions on how to interpret the discourse, and their study is paramount to understand the mechanism underlying discourse organization. This paper presents a new language resource, an ISO-based annotated multilingual parallel corpus for discourse markers. The corpus comprises nine languages, Bulgarian, Lithuanian, German, European Portuguese, Hebrew, Romanian, Polish, and Macedonian, with English as a pivot language. In order to represent the meaning of the discourse markers, we propose an annotation scheme of discourse relations from ISO 24617-8 with a plug-in to ISO 24617-2 for communicative functions. We describe an experiment in which we applied the annotation scheme to assess its validity. The results reveal that, although some extensions are required to cover all the multilingual data, it provides a proper representation of discourse markers value. Additionally, we report some relevant contrastive phenomena concerning discourse markers interpretation and role in discourse. This first step will allow us to develop deep learning methods to identify and extract discourse relations and communicative functions, and to represent that information as Linguistic Linked Open Data (LLOD).

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Using Machine Translation Techniques to Induce Multilingual Lexica of Discourse Markers

2015-03-31 · António Lopes, David Martins de Matos, Vera Cabarrão, Ricardo Ribeiro 외

Discourse markers are universal linguistic events subject to language variation. Although an extensive literature has already reported language specific traits of these events, little has been said on their cross-languag…

Machine TranslationSentenceTranslation

Discourse-level Annotation over Europarl for Machine Translation: Connectives and Pronouns

2012-05-01 · LREC 2012 5 · Andrei Popescu-Belis, Thomas Meyer, Jeevanthi Liyanapathirana, Bruno Cartoni 외

This paper describes methods and results for the annotation of two discourse-level phenomena, connectives and pronouns, over a multilingual parallel corpus. Excerpts from Europarl in English and French have been annotate…

Machine TranslationTranslation

The RST Spanish-Chinese Treebank

2018-08-01 · COLING 2018 8 · Shuyuan Cao, Iria da Cunha, Mikel Iruskieta

Discourse analysis is necessary for different tasks of Natural Language Processing (NLP). As two of the most spoken languages in the world, discourse analysis between Spanish and Chinese is important for NLP research. Th…

The Parallel Meaning Bank: Towards a Multilingual Corpus of Translations Annotated with Compositional Meaning Representations

2017-02-13 · EACL 2017 4 · Lasha Abzianidze, Johannes Bjerva, Kilian Evang, Hessel Haagsma 외

The Parallel Meaning Bank is a corpus of translations annotated with shared, formal meaning representations comprising over 11 million words divided over four languages (English, German, Italian, and Dutch). Our approach…

Zero-shot transfer for implicit discourse relation classification

2019-07-30 · Murathan Kurfali, Robert Östling

Automatically classifying the relation between sentences in a discourse is a challenging task, in particular when there is no overt expression of the relation. It becomes even more challenging by the fact that annotated …

ClassificationGeneral ClassificationImplicit Discourse Relation ClassificationRelation+2