paper-with-me

홈 › Papers

A Framework for Deprecating Datasets: Standardizing Documentation, Identification, and Communication

2021-10-18 · Alexandra Sasha Luccioni, Frances Corry, Hamsini Sridharan, Mike Ananny, Jason Schultz, Kate Crawford

Datasets are central to training machine learning (ML) models. The ML community has recently made significant improvements to data stewardship and documentation practices across the model development life cycle. However, the act of deprecating, or deleting, datasets has been largely overlooked, and there are currently no standardized approaches for structuring this stage of the dataset life cycle. In this paper, we study the practice of dataset deprecation in ML, identify several cases of datasets that continued to circulate despite having been deprecated, and describe the different technical, legal, ethical, and organizational issues raised by such continuations. We then propose a Dataset Deprecation Framework that includes considerations of risk, mitigation of impact, appeal mechanisms, timeline, post-deprecation protocols, and publication checks that can be adapted and implemented by the ML community. Finally, we propose creating a centralized, sustainable repository system for archiving datasets, tracking dataset modifications or deprecations, and facilitating practices of care and stewardship that can be integrated into research and publication processes.

📄 PDF Abstract BibTeX arXiv:2111.04424

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An LSTM-Based Deep Learning Approach for Detecting Self-Deprecating Sarcasm in Textual Data

2019-12-01 · ICON 2019 12 · Ashraf Kamal, Muhammad Abulaish

Self-deprecating sarcasm is a special category of sarcasm, which is nowadays popular and useful for many real-life applications, such as brand endorsement, product campaign, digital marketing, and advertisement. The self…

Deep LearningMarketingSarcasm Detection

Datasheets for AI and medical datasets (DAIMS): a data validation and documentation framework before machine learning analysis in medical research

2025-01-23 · Ramtin Zargari Marandi, Anne Svane Frahm, Maja Milojevic

Despite progresses in data engineering, there are areas with limited consistencies across data validation and documentation procedures causing confusions and technical problems in research involving machine learning. The…

Deprecating Benchmarks: Criteria and Framework

2025-07-08 · Ayrton San Joaquin, Rokas Gipiškis, Leon Staufer, Ariel Gil arxiv

As frontier artificial intelligence (AI) models rapidly advance, benchmarks are integral to comparing different models and measuring their progress in different task-specific domains. However, there is a lack of guidance…

ATAG: AI-Agent Application Threat Assessment with Attack Graphs

2025-06-03 · Parth Atulbhai Gandhi, Akansha Shukla, David Tayouri, Beni Ifland 외

Evaluating the security of multi-agent systems (MASs) powered by large language models (LLMs) is challenging, primarily because of the systems' complex internal dynamics and the evolving nature of LLM vulnerabilities. Tr…

AI Agent

Reusable Templates and Guides For Documenting Datasets and Models for Natural Language Processing and Generation: A Case Study of the HuggingFace and GEM Data and Model Cards

2021-08-16 · ACL (GEM) 2021 8 · Angelina McMillan-Major, Salomey Osei, Juan Diego Rodriguez, Pawan Sasanka Ammanamanchi 외

Developing documentation guidelines and easy-to-use templates for datasets and models is a challenging task, especially given the variety of backgrounds, skills, and incentives of the people involved in the building of n…

Text Generation