paper-with-me

홈 › Papers

TNT-NLG, System 1: Using a statistical NLG to massively augment crowd-sourced data for neural generation

2018-04-26 · E2E NLG Challenge System Descriptions 2018 4 · Shereen Oraby, Lena Reed, Shubhangi Tandon, Stephanie Lukin, Marilyn A. Walker

Ever since the successful application of sequence to sequence learning for neural machine translation systems (Sutskever et al., 2014), interest has surged in its applicability towards language generation in other problem domains. In the area of natural language generation (NLG), there has been a great deal of interest in end-to-end (E2E) neural models that learn and generate natural language sentence realizations in one step. In this paper, we present TNT-NLG System 1, our first system submission to the E2E NLG Challenge, where we generate natural language (NL) realizations from meaning representations (MRs) in the restaurant domain by massively expanding the training dataset. We develop two models for this system, based on Dusek et al.’s (2016a) open source baseline model and context-aware neural language generator. Starting with the MR and NL pairs from the E2E generation challenge dataset, we explode the size of the training set using PERSONAGE (Mairesse and Walker, 2010), a statistical generator able to produce varied realizations from MRs, and use our expanded data as contextual input into our models. We present evaluation results using automated and human evaluation metrics, and describe directions for future work.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Data-to-Text GenerationMachine TranslationSentenceText GenerationTranslation

Similar Papers 제목 키워드 기반

Probabilistic Multigraph Modeling for Improving the Quality of Crowdsourced Affective Data

2017-01-04 · Jianbo Ye, Jia Li, Michelle G. Newman, Reginald B. Adams, Jr. 외

We proposed a probabilistic approach to joint modeling of participants' reliability and humans' regularity in crowdsourced affective studies. Reliability measures how likely a subject will respond to a question seriously…

Improve Learning from Crowds via Generative Augmentation

2021-07-22 · Zhendong Chu, Hongning Wang

Crowdsourcing provides an efficient label collection schema for supervised machine learning. However, to control annotation cost, each instance in the crowdsourced data is typically annotated by a small number of annotat…

BIG-bench Machine LearningData Augmentation

Beyond the Crowd: LLM-Augmented Community Notes for Governing Health Misinformation

2025-10-13 · Jiaying Wu, Zihang Fu, Haonan Wang, Fanxiao Li 외 arxiv

Community Notes, the crowd-sourced misinformation governance system on X (formerly Twitter), allows users to flag misleading posts, attach contextual notes, and rate the notes' helpfulness. However, our empirical analysi…

Letting the Data Speak: Extracting Keywords from Crowdsourced Collections with AI

2026-07-10 · Miguel Arana-Catania, Catherine Conisbee, Matthew Kidd arxiv

Identifying and assigning keywords at scale is a technical, practical, and ethical challenge for crowdsourced collections. This article reports the findings of the "Extracting Keywords from Crowdsourced Collections" proj…

Keyword Extraction

Detecting gender differences in perception of emotion in crowdsourced data

2019-10-24 · Shahan Ali Memon, Hira Dhamyal, Oren Wright, Daniel Justice 외

Do men and women perceive emotions differently? Popular convictions place women as more emotionally perceptive than men. Empirical findings, however, remain inconclusive. Most prior studies focus on visual modalities. In…