paper-with-me

Papers

Shared Task on Evaluating Accuracy in Natural Language Generation

2020-06-22 · Ehud Reiter, Craig Thomson

We propose a shared task on methodologies and algorithms for evaluating the accuracy of generated texts. Participants will measure the accuracy of basketball game summaries produced by NLG systems from basketball box score data.

📄 PDF Abstract BibTeX arXiv:2006.12234

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

A Course Shared Task on Evaluating LLM Output for Clinical Questions

2024-07-31 · Yufang Hou, Thy Thy Tran, Doan Nam Long Vu, Yiwen Cao 외

This paper presents a shared task that we organized at the Foundations of Language Technology (FoLT) course in 2023/2024 at the Technical University of Darmstadt, which focuses on evaluating the output of Large Language …

Beyond Accuracy: Evaluating Self-Consistency of Code Large Language Models with IdentityChain

2023-10-21 · Marcus J. Min, Yangruibo Ding, Luca Buratti, Saurabh Pujar 외

Code Large Language Models (Code LLMs) are being increasingly employed in real-life applications, so evaluating them is critical. While the conventional accuracy evaluates the performance of Code LLMs on a set of individ…

Code GenerationCode Summarization

You Cannot Feed Two Birds with One Score: the Accuracy-Naturalness Tradeoff in Translation

2025-03-31 · Gergely Flamich, David Vilar, Jan-Thorsten Peter, Markus Freitag

The goal of translation, be it by human or by machine, is, given some text in a source language, to produce text in a target language that simultaneously 1) preserves the meaning of the source text and 2) achieves natura…

Machine TranslationTranslation

TextGraphs 2022 Shared Task on Natural Language Premise Selection

2022-10-01 · COLING (TextGraphs) 2022 10 · Marco Valentino, Deborah Ferreira, Mokanarangan Thayaparan, André Freitas 외

The Shared Task on Natural Language Premise Selection (NLPS) asks participants to retrieve the set of premises that are most likely to be useful for proving a given mathematical statement from a supporting knowledge base…

Recurrent Neural Network-Based Sentence Encoder with Gated Attention for Natural Language Inference

2017-08-04 · WS 2017 9 · Qian Chen, Xiaodan Zhu, Zhen-Hua Ling, Si Wei 외

The RepEval 2017 Shared Task aims to evaluate natural language understanding models for sentence representation, in which a sentence is represented as a fixed-length vector with neural networks and the quality of the rep…

Natural Language InferenceNatural Language UnderstandingSentence