paper-with-me

홈 › Papers

Baselines and a datasheet for the Cerema AWP dataset

2018-06-11 · Ismaïla Seck, Khouloud Dahmane, Pierre Duthon, Gaëlle Loosli

This paper presents the recently published Cerema AWP (Adverse Weather Pedestrian) dataset for various machine learning tasks and its exports in machine learning friendly format. We explain why this dataset can be interesting (mainly because it is a greatly controlled and fully annotated image dataset) and present baseline results for various tasks. Moreover, we decided to follow the very recent suggestions of datasheets for dataset, trying to standardize all the available information of the dataset, with a transparency objective.

📄 PDF Abstract BibTeX arXiv:1806.04016

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine Learning

Similar Papers 제목 키워드 기반

Datasheets for Datasets

2018-03-23 · Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan 외

The machine learning community currently has no standardized process for documenting datasets, which can lead to severe consequences in high-stakes domains. To address this gap, we propose datasheets for datasets. In the…

BIG-bench Machine Learning

Healthsheet: Development of a Transparency Artifact for Health Datasets

2022-02-26 · Negar Rostamzadeh, Diana Mincu, Subhrajit Roy, Andrew Smart 외

Machine learning (ML) approaches have demonstrated promising results in a wide range of healthcare applications. Data plays a crucial role in developing ML-based healthcare systems that directly affect people's lives. Ma…

Diagnostic

MT-Adapted Datasheets for Datasets: Template and Repository

2020-05-27 · Marta R. Costa-jussà, Roger Creus, Oriol Domingo, Albert Domínguez 외

In this report we are taking the standardized model proposed by Gebru et al. (2018) for documenting the popular machine translation datasets of the EuroParl (Koehn, 2005) and News-Commentary (Barrault et al., 2019). With…

Machine TranslationTranslation

Datasheet for the Pile

2022-01-13 · Stella Biderman, Kieran Bicheno, Leo Gao

This datasheet describes the Pile, a 825 GiB dataset of human-authored text compiled by EleutherAI for use in large-scale language modeling. The Pile is comprised of 22 different text sources, ranging from original scrap…

Language ModelingLanguage Modelling

Causal datasheet: An approximate guide to practically assess Bayesian networks in the real world

2020-03-12 · Bradley Butcher, Vincent S. Huang, Jeremy Reffin, Sema K. Sgaier 외

In solving real-world problems like changing healthcare-seeking behaviors, designing interventions to improve downstream outcomes requires an understanding of the causal links within the system. Causal Bayesian Networks …