Creating a Corpus for Russian Data-to-Text Generation Using Neural Machine Translation and Post-Editing
In this paper, we propose an approach for semi-automatically creating a data-to-text (D2T) corpus for Russian that can be used to learn a D2T natural language generation model. An error analysis of the output of an English-to-Russian neural machine translation system shows that 80{\%} of the automatically translated sentences contain an error and that 53{\%} of all translation errors bear on named entities (NE). We therefore focus on named entities and introduce two post-editing techniques for correcting wrongly translated NEs.
Code (1)
Tasks
Data-to-Text GenerationMachine TranslationText GenerationTranslationSimilar Papers 제목 키워드 기반
Russian Texts Detoxification with Levenshtein Editing
Text detoxification is a style transfer task of creating neutral versions of toxic texts. In this paper, we use the concept of text editing to build a two-step tagging-based detoxification model using a parallel corpus o…
Style TransferAutomatically Ranked Russian Paraphrase Corpus for Text Generation
The article is focused on automatic development and ranking of a large corpus for Russian paraphrase generation which proves to be the first corpus of such type in Russian computational linguistics. Existing manually ann…
Paraphrase GenerationSentenceSentence SimilarityText GenerationThe Algorithmic Inflection of Russian and Generation of Grammatically Correct Text
We present a deterministic algorithm for Russian inflection. This algorithm is implemented in a publicly available web-service www.passare.ru which provides functions for inflection of single words, word matching and syn…
Sense-Annotated Corpus for Russian
We present a sense-annotated corpus for Russian. The resource was obtained my manually annotating texts from the OpenCorpora corpus, an open corpus for the Russian language, by senses of Russian wordnet RuWordNet. The an…
Word Sense DisambiguationCollocation2Text: Controllable Text Generation from Guide Phrases in Russian
Large pre-trained language models are capable of generating varied and fluent texts. Starting from the prompt, these models generate a narrative that can develop unpredictably. The existing methods of controllable text g…
ArticlesText Generation