paper-with-me

Papers

Automatic Corpus Extension for Data-driven Natural Language Generation

2016-05-01 · LREC 2016 5 · Elena Manishina, Bassam Jabaian, St{\'e}phane Huet, Fabrice Lef{\`e}vre

As data-driven approaches started to make their way into the Natural Language Generation (NLG) domain, the need for automation of corpus building and extension became apparent. Corpus creation and extension in data-driven NLG domain traditionally involved manual paraphrasing performed by either a group of experts or with resort to crowd-sourcing. Building the training corpora manually is a costly enterprise which requires a lot of time and human resources. We propose to automate the process of corpus extension by integrating automatically obtained synonyms and paraphrases. Our methodology allowed us to significantly increase the size of the training corpus and its level of variability (the number of distinct tokens and specific syntactic structures). Our extension solutions are fully automatic and require only some initial validation. The human evaluation results confirm that in many cases native speakers favor the outputs of the model built on the extended corpus.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

An Extension of the Slovak Broadcast News Corpus based on Semi-Automatic Annotation

2016-05-01 · LREC 2016 5 · Peter Viszlay, J{\'a}n Sta{\v{s}}, Tom{\'a}{\v{s}} Koct{\'u}r, Martin Lojka 외

In this paper, we introduce an extension of our previously released TUKE-BNews-SK corpus based on a semi-automatic annotation scheme. It firstly relies on the automatic transcription of the BN data performed by our Slova…

speech-recognitionSpeech Recognition

An extension of ISO-Space for annotating object direction

2016-12-01 · WS 2016 12 · Daiki Gotou, Hitoshi Nishikawa, Takenobu Tokunaga

In this paper, we extend an existing annotation scheme ISO-Space for annotating necessary spatial information for the task placing an specified object at a specified location with a specified direction according to a nat…

ObjectTAG

A Corpus for Automatic Readability Assessment and Text Simplification of German

2019-09-19 · LREC 2020 5 · Alessia Battisti, Sarah Ebling

In this paper, we present a corpus for use in automatic readability assessment and automatic text simplification of German. The corpus is compiled from web sources and consists of approximately 211,000 sentences. As a no…

BIG-bench Machine LearningText Simplification

Convex Polytope Modelling for Unsupervised Derivation of Semantic Structure for Data-efficient Natural Language Understanding

2022-01-25 · Jingyan Zhou, Xiaohan Feng, King Keung Wu, Helen Meng

Popular approaches for Natural Language Understanding (NLU) usually rely on a huge amount of annotated data or handcrafted rules, which is laborious and not adaptive to domain extension. We recently proposed a Convex-Pol…

Natural Language Understanding

More or less controlled elicitation of argumentative text: Enlarging a microtext corpus via crowdsourcing

2018-11-01 · WS 2018 11 · Maria Skeppstedt, Andreas Peldszus, Manfred Stede

We present an extension of an annotated corpus of short argumentative texts that had originally been built in a controlled text production experiment. Our extension more than doubles the size of the corpus by means of cr…

Argument Mining