paper-with-me

홈 › Papers

Evaluation Guidelines to Deal with Implicit Phenomena to Assess Factuality in Data-to-Text Generation

2021-08-01 · ACL (unimplicit) 2021 8 · Roy Eisenstadt, Michael Elhadad

Data-to-text generation systems are trained on large datasets, such as WebNLG, Ro-toWire, E2E or DART. Beyond traditional token-overlap evaluation metrics (BLEU or METEOR), a key concern faced by recent generators is to control the factuality of the generated text with respect to the input data specification. We report on our experience when developing an automatic factuality evaluation system for data-to-text generation that we are testing on WebNLG and E2E data. We aim to prepare gold data annotated manually to identify cases where the text communicates more information than is warranted based on the in-put data (extra) or fails to communicate data that is part of the input (missing). While analyzing reference (data, text) samples, we encountered a range of systematic uncertainties that are related to cases on implicit phenomena in text, and the nature of non-linguistic knowledge we expect to be involved when assessing factuality. We derive from our experience a set of evaluation guidelines to reach high inter-annotator agreement on such cases.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Data-to-Text GenerationText Generation

Similar Papers 제목 키워드 기반

Iterative Refinement and Quality Checking of Annotation Guidelines --- How to Deal Effectively with Semantically Sloppy Named Entity Types, such as Pathological Phenomena

2012-05-01 · LREC 2012 5 · Udo Hahn, Elena Beisswanger, Ekaterina Buyko, Erik Faessler 외

We here discuss a methodology for dealing with the annotation of semantically hard to delineate, i.e., sloppy, named entity types. To illustrate sloppiness of entities, we treat an example from the medical domain, namely…

DescriptiveNamed Entity Recognition (NER)

I Feel Offended, Don't Be Abusive! Implicit/Explicit Messages in Offensive and Abusive Language

2020-05-01 · LREC 2020 5 · Tommaso Caselli, Valerio Basile, Jelena Mitrovi{\'c}, Inga Kartoziya 외

Abusive language detection is an unsolved and challenging problem for the NLP community. Recent literature suggests various approaches to distinguish between different language phenomena (e.g., hate speech vs. cyberbully…

Abusive Language

Assisted Counterspeech Writing at the Crossroads of Hate Speech and Misinformation

2026-05-21 · Genoveffa Martone, Helena Bonaldi, Marco Guerini arxiv

Hate speech and misinformation frequently co-occur online, amplifying prejudice and polarization. Given their scale, using Large Language Models (LLMs) to assist expert counterspeech (CS) writing has gained interest, yet…

Defining and Detecting Vulnerability in Human Evaluation Guidelines: A Preliminary Study Towards Reliable NLG Evaluation

2024-06-12 · Jie Ruan, Wenqing Wang, Xiaojun Wan

Human evaluation serves as the gold standard for assessing the quality of Natural Language Generation (NLG) systems. Nevertheless, the evaluation guideline, as a pivotal element ensuring reliable and reproducible human a…

nlg evaluationText GenerationVulnerability Detection

Implicit Phenomena in Short-answer Scoring Data

2021-08-01 · ACL (unimplicit) 2021 8 · Marie Bexte, Andrea Horbach, Torsten Zesch

Short-answer scoring is the task of assessing the correctness of a short text given as response to a question that can come from a variety of educational scenarios. As only content, not form, is important, the exact word…

Word Embeddings