KOJAK: A New Corpus for Studying German Discourse Particle ja
In German, ja can be used as a discourse particle to indicate that a proposition, according to the speaker, is believed by both the speaker and audience. We use this observation to create KoJaK, a distantly-labeled English dataset derived from Europarl for studying when a speaker believes a statement to be common ground. This corpus is then analyzed to identify lexical choices in English that correspond with German ja. Finally, we perform experiments on the dataset to predict if an English clause corresponds to a German clause containing ja and achieve an F-measure of 75.3% on a balanced test corpus.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Swiss-AL: A Multilingual Swiss Web Corpus for Applied Linguistics
The Swiss Web Corpus for Applied Linguistics (Swiss-AL) is a multilingual (German, French, Italian) collection of texts from selected web sources. Unlike most other web corpora it is not intended for NLP purposes, but ra…
The Potsdam Commentary Corpus 2.2: Extending Annotations for Shallow Discourse Parsing
We present the Potsdam Commentary Corpus 2.2, a German corpus of news editorials annotated on several different levels. New in the 2.2 version of the corpus are two additional annotation layers for coherence relations fo…
Discourse ParsingRelationIdentifying Explicit Discourse Connectives in German
We are working on an end-to-end Shallow Discourse Parsing system for German and in this paper focus on the first subtask: the identification of explicit connectives. Starting with the feature set from an English system a…
Discourse ParsingDiscovery of Discourse-Related Language Contrasts through Alignment Discrepancies in English-German Translation
In this paper, we analyse alignment discrepancies for discourse structures in English-German parallel data {--} sentence pairs, in which discourse structures in target or source texts have no alignment in the correspondi…
Machine TranslationSentenceTranslationWord Alignment“The word expired when that world awoke.” New Challenges for Research with Large Text Corpora and Corpus-Based Discourse Studies in Totalitarian Times
In the following poster proposal a report will be given on the prospects of a promising corpus project initiated by one of the large digital text corpora hosted by the Austrian Academy of Sciences. First, the resources o…