paper-with-me

홈 › Papers

Raison d’être of the benchmark dataset: A Survey of Current Practices of Benchmark Dataset Sharing Platforms

2022-05-01 · nlppower (ACL) 2022 5 · Jaihyun Park, Sullam Jeoung

This paper critically examines the current practices of benchmark dataset sharing in NLP and suggests a better way to inform reusers of the benchmark dataset. As the dataset sharing platform plays a key role not only in distributing the dataset but also in informing the potential reusers about the dataset, we believe data-sharing platforms should provide a comprehensive context of the datasets. We survey four benchmark dataset sharing platforms: HuggingFace, PaperswithCode, Tensorflow, and Pytorch to diagnose the current practices of how the dataset is shared which metadata is shared and omitted. To be specific, drawing on the concept of data curation which considers the future reuse when the data is made public, we advance the direction that benchmark dataset sharing platforms should take into consideration. We identify that four benchmark platforms have different practices of using metadata and there is a lack of consensus on what social impact metadata is. We believe the problem of missing a discussion around social impact in the dataset sharing platforms has to do with the failed agreement on who should be in charge. We propose that the benchmark dataset should develop social impact metadata and data curator should take a role in managing the social impact metadata.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Extraction et analyse automatique des comparaisons et des pseudo-comparaisons pour la d\'etection des comparaisons figuratives

2015-06-01 · JEPTALNRECITAL 2015 6 · Suzanne Mpouli, Jean-Gabriel Ganascia

Le pr{\'e}sent article s{'}int{\'e}resse {\`a} la d{\'e}tection et {\`a} la d{\'e}sambigu{\"\i}sation des comparaisons figuratives. Il d{\'e}crit un algorithme qui utilise un analyseur syntaxique de surface (chunker) et …

The Scientific Method in the Science of Machine Learning

2019-04-24 · Jessica Zosa Forde, Michela Paganini

In the quest to align deep learning with the sciences to address calls for rigor, safety, and interpretability in machine learning systems, this contribution identifies key missing pieces: the stages of hypothesis formul…

BIG-bench Machine Learning

Data Management For Training Large Language Models: A Survey

2023-12-04 · Zige Wang, Wanjun Zhong, YuFei Wang, Qi Zhu 외

Data plays a fundamental role in training Large Language Models (LLMs). Efficient data management, particularly in formulating a well-suited training dataset, is significant for enhancing model performance and improving …

ManagementSurvey

About Face: A Survey of Facial Recognition Evaluation

2021-02-01 · Inioluwa Deborah Raji, Genevieve Fried

We survey over 100 face datasets constructed between 1976 to 2019 of 145 million images of over 17 million subjects from a range of sources, demographics and conditions. Our historical survey reveals that these datasets …

Survey

Naming the Pain in Machine Learning-Enabled Systems Engineering

2024-05-20 · Marcos Kalinowski, Daniel Mendez, Görkem Giray, Antonio Pedro Santos Alves 외

Context: Machine learning (ML)-enabled systems are being increasingly adopted by companies aiming to enhance their products and operational processes. Objective: This paper aims to deliver a comprehensive overview of the…

Survey