paper-with-me

Papers

Wikibench: Community-Driven Data Curation for AI Evaluation on Wikipedia

2024-02-21 · Tzu-Sheng Kuo, Aaron Halfaker, Zirui Cheng, Jiwoo Kim, Meng-Hsin Wu, Tongshuang Wu, Kenneth Holstein, Haiyi Zhu

AI tools are increasingly deployed in community contexts. However, datasets used to evaluate AI are typically created by developers and annotators outside a given community, which can yield misleading conclusions about AI performance. How might we empower communities to drive the intentional design and curation of evaluation datasets for AI that impacts them? We investigate this question on Wikipedia, an online community with multiple AI-based content moderation tools deployed. We introduce Wikibench, a system that enables communities to collaboratively curate AI evaluation datasets, while navigating ambiguities and differences in perspective through discussion. A field study on Wikipedia shows that datasets curated using Wikibench can effectively capture community consensus, disagreement, and uncertainty. Furthermore, study participants used Wikibench to shape the overall data curation process, including refining label definitions, determining data inclusion criteria, and authoring data statements. Based on our findings, we propose future directions for systems that support community-driven data curation.

📄 PDF Abstract BibTeX arXiv:2402.14147

Code (1)

tskuo/wikibench 공식 구현

Similar Papers 제목 키워드 기반

The Esethu Framework: Reimagining Sustainable Dataset Governance and Curation for Low-Resource Languages

2025-02-21 · Jenalea Rajab, Anuoluwapo Aremu, Everlyn Asiko Chimoto, Dale Dunbar 외

This paper presents the Esethu Framework, a sustainable data curation framework specifically designed to empower local communities and ensure equitable benefit-sharing from their linguistic resources. This framework is s…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

A quick guide for student-driven community genome annotation

2018-10-16

High quality gene models are necessary to expand the molecular and genetic tools available for a target organism, but these are available for only a handful of model organisms that have undergone extensive curation and e…

WelQrate: Defining the Gold Standard in Small Molecule Drug Discovery Benchmarking

2024-11-14 · Yunchao, Liu, Ha Dong, Xin Wang 외

While deep learning has revolutionized computer-aided drug discovery, the AI community has predominantly focused on model innovation and placed less emphasis on establishing best benchmarking practices. We posit that wit…

BenchmarkingDrug Discovery

Atla Selene Mini: A General Purpose Evaluation Model

2025-01-27 · Andrei Alexandru, Antonia Calvi, Henry Broomfield, Jackson Golden 외

We introduce Atla Selene Mini, a state-of-the-art small language model-as-a-judge (SLMJ). Selene Mini is a general-purpose evaluator that outperforms the best SLMJs and GPT-4o-mini on overall performance across 11 out-of…

Language ModelingLanguage ModellingmodelSmall Language Model

NeurIPS 2023 LLM Efficiency Fine-tuning Competition

2025-03-13 · Mark Saroufim, Yotam Perlitz, Leshem Choshen, Luca Antiga 외

Our analysis of the NeurIPS 2023 large language model (LLM) fine-tuning competition revealed the following trend: top-performing models exhibit significant overfitting on benchmark datasets, mirroring the broader issue o…

Language ModelingLanguage ModellingLarge Language Model