paper-with-me

Papers

Shared Task on Evaluating Accuracy

2020-12-01 · INLG (ACL) 2020 12 · Ehud Reiter, Craig Thomson

We propose a shared task on methodologies and algorithms for evaluating the accuracy of generated texts, specifically summaries of basketball games produced from basketball box score and other game data. We welcome submissions based on protocols for human evaluation, automatic metrics, as well as combinations of human evaluations and metrics.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Shared Task on Evaluating Accuracy in Natural Language Generation

2020-06-22 · Ehud Reiter, Craig Thomson

We propose a shared task on methodologies and algorithms for evaluating the accuracy of generated texts. Participants will measure the accuracy of basketball game summaries produced by NLG systems from basketball box sco…

Text Generation

Generation Challenges: Results of the Accuracy Evaluation Shared Task

2021-08-12 · INLG (ACL) 2021 8 · Craig Thomson, Ehud Reiter

The Shared Task on Evaluating Accuracy focused on techniques (both manual and automatic) for evaluating the factual accuracy of texts produced by neural NLG systems, in a sports-reporting domain. Four teams submitted eva…

Shared Task in Evaluating Accuracy: Leveraging Pre-Annotations in the Validation Process

2021-08-01 · INLG (ACL) 2021 8 · Nicolas Garneau, Luc Lamontagne

We hereby present our submission to the Shared Task in Evaluating Accuracy at the INLG 2021 Conference. Our evaluation protocol relies on three main components; rules and text classifiers that pre-annotate the dataset, a…

A Course Shared Task on Evaluating LLM Output for Clinical Questions

2024-07-31 · Yufang Hou, Thy Thy Tran, Doan Nam Long Vu, Yiwen Cao 외

This paper presents a shared task that we organized at the Foundations of Language Technology (FoLT) course in 2023/2024 at the Technical University of Darmstadt, which focuses on evaluating the output of Large Language …

ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation

2026-08-31 · Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti, Farina Amir 외 arxiv

We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, covering spoken visual question answering and image-grounded hallucination …

Visual Question AnsweringText-to-Image Generation