Same domain different discourse style - A case study on Language Resources for data-driven Machine Translation
Data-driven machine translation (MT) approaches became very popular during last years, especially for language pairs for which it is difficult to find specialists to develop transfer rules. Statistical (SMT) or example-based (EBMT) systems can provide reasonable translation quality for assimilation purposes, as long as a large amount of training data is available. Especially SMT systems rely on parallel aligned corpora which have to be statistical relevant for the given language pair. The construction of large domain specific parallel corpora is time- and cost-consuming; the current practice relies on one or two big such corpora per language pair. Recent developed strategies ensure certain portability to other domains through specialized lexicons or small domain specific corpora. In this paper we discuss the influence of different discourse styles on statistical machine translation systems. We investigate how a pure SMT performs when training and test data belong to same domain but the discourse style varies.
Code (0)
등록된 구현이 없습니다.
Tasks
Information RetrievalMachine TranslationTranslationSimilar Papers 제목 키워드 기반
Predicting Discourse Structure using Distant Supervision from Sentiment
Discourse parsing could not yet take full advantage of the neural NLP revolution, mostly due to the lack of annotated datasets. We propose a novel approach that uses distant supervision on an auxiliary task (sentiment cl…
Discourse ParsingMultiple Instance LearningPredictionSentiment Analysis+1Using Referring Expression Generation to Model Literary Style
Novels and short stories are not just remarkable because of what events they represent. The narrative style they employ is significant. To understand the specific contributions of different aspects of this style, it is p…
modelReferring ExpressionReferring expression generationToward Cross-theory Discourse Relation Annotation
In this exploratory study, we attempt to automatically induce PDTB-style relations from RST trees. We work with a German corpus of news commentary articles, annotated for RST trees and explicit PDTB-style relations and w…
ArticlesImplicit RelationsRelationPredicting Discourse Trees from Transformer-based Neural Summarizers
Previous work indicates that discourse information benefits summarization. In this paper, we explore whether this synergy between discourse and summarization is bidirectional, by inferring document-level discourse trees …
Discourse Parsing