paper-with-me

홈 › Papers

Causal datasheet: An approximate guide to practically assess Bayesian networks in the real world

2020-03-12 · Bradley Butcher, Vincent S. Huang, Jeremy Reffin, Sema K. Sgaier, Grace Charles, Novi Quadrianto

In solving real-world problems like changing healthcare-seeking behaviors, designing interventions to improve downstream outcomes requires an understanding of the causal links within the system. Causal Bayesian Networks (BN) have been proposed as one such powerful method. In real-world applications, however, confidence in the results of BNs are often moderate at best. This is due in part to the inability to validate against some ground truth, as the DAG is not available. This is especially problematic if the learned DAG conflicts with pre-existing domain doctrine. At the policy level, one must justify insights generated by such analysis, preferably accompanying them with uncertainty estimation. Here we propose a causal extension to the datasheet concept proposed by Gebru et al (2018) to include approximate BN performance expectations for any given dataset. To generate the results for a prototype Causal Datasheet, we constructed over 30,000 synthetic datasets with properties mirroring characteristics of real data. We then recorded the results given by state-of-the-art structure learning algorithms. These results were used to populate the Causal Datasheet, and recommendations were automatically generated dependent on expected performance. As a proof of concept, we used our Causal Datasheet Generation Tool (CDG-T) to assign expected performance expectations to a maternal health survey we conducted in Uttar Pradesh, India.

📄 PDF Abstract BibTeX arXiv:2003.07182

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Healthsheet: Development of a Transparency Artifact for Health Datasets

2022-02-26 · Negar Rostamzadeh, Diana Mincu, Subhrajit Roy, Andrew Smart 외

Machine learning (ML) approaches have demonstrated promising results in a wide range of healthcare applications. Data plays a crucial role in developing ML-based healthcare systems that directly affect people's lives. Ma…

Diagnostic

The Human Evaluation Datasheet: A Template for Recording Details of Human Evaluation Experiments in NLP

2022-05-01 · HumEval (ACL) 2022 5 · Anastasia Shimorina, Anya Belz

This paper presents the Human Evaluation Datasheet (HEDS), a template for recording the details of individual human evaluation experiments in Natural Language Processing (NLP), and reports on first experience of research…

Datasheets for Datasets

2018-03-23 · Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan 외

The machine learning community currently has no standardized process for documenting datasets, which can lead to severe consequences in high-stakes domains. To address this gap, we propose datasheets for datasets. In the…

BIG-bench Machine Learning

SpecMap: Hierarchical LLM Agent for Datasheet-to-Code Traceability Link Recovery in Systems Engineering

2026-01-16 · Vedant Nipane, Pulkit Agrawal, Amit Singh arxiv

Establishing precise traceability between embedded systems datasheets and their corresponding code implementations remains a fundamental challenge in systems engineering, particularly for low-level software where manual …

Information Retrieval

MT-Adapted Datasheets for Datasets: Template and Repository

2020-05-27 · Marta R. Costa-jussà, Roger Creus, Oriol Domingo, Albert Domínguez 외

In this report we are taking the standardized model proposed by Gebru et al. (2018) for documenting the popular machine translation datasets of the EuroParl (Koehn, 2005) and News-Commentary (Barrault et al., 2019). With…

Machine TranslationTranslation