paper-with-me

홈 › Papers

HEDS 3.0: The Human Evaluation Data Sheet Version 3.0

2024-12-10 · Anya Belz, Craig Thomson

This paper presents version 3.0 of the Human Evaluation Datasheet (HEDS). This update is the result of our experience using HEDS in the context of numerous recent human evaluation experiments, including reproduction studies, and of feedback received. Our main overall goal was to improve clarity, and to enable users to complete the datasheet more consistently and comparably. The HEDS 3.0 package consists of the digital data sheet, documentation, and code for exporting completed data sheets as latex files, all available from the HEDS GitHub.

📄 PDF Abstract BibTeX arXiv:2412.07940

Code (1)

DCU-NLG/HEDS-3.0 공식 구현

Similar Papers 제목 키워드 기반

The Human Evaluation Datasheet: A Template for Recording Details of Human Evaluation Experiments in NLP

2022-05-01 · HumEval (ACL) 2022 5 · Anastasia Shimorina, Anya Belz

This paper presents the Human Evaluation Datasheet (HEDS), a template for recording the details of individual human evaluation experiments in Natural Language Processing (NLP), and reports on first experience of research…

Unsupervised Lead Sheet Generation via Semantic Compression

2023-10-16 · Zachary Novack, Nikita Srivatsan, Taylor Berg-Kirkpatrick, Julian McAuley

Lead sheets have become commonplace in generative music research, being used as an initial compressed representation for downstream tasks like multitrack music generation and automatic arrangement. Despite this, research…

Music CompressionMusic GenerationSemantic Compression

The Human Evaluation Datasheet 1.0: A Template for Recording Details of Human Evaluation Experiments in NLP

2021-03-17 · Anastasia Shimorina, Anya Belz

This paper introduces the Human Evaluation Datasheet, a template for recording the details of individual human evaluation experiments in Natural Language Processing (NLP). Originally taking inspiration from seminal paper…

SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation

2024-06-21 · Zeyao Ma, Bohan Zhang, Jing Zhang, Jifan Yu 외

We introduce SpreadsheetBench, a challenging spreadsheet manipulation benchmark exclusively derived from real-world scenarios, designed to immerse current large language models (LLMs) in the actual workflow of spreadshee…

Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows

2025-12-15 · Haoyu Dong, Pengkun Zhang, Yan Gao, Xuanyu Dong 외 arxiv

We introduce FinWorkBench (a.k.a. Finch) for evaluating AI agents on real-world, enterprise-grade finance and accounting workflows that interleave data entry, structuring, formatting, web search, cross-file retrieval, ca…