paper-with-me

Papers

Veridical Data Science

2019-01-23 · Bin Yu, Karl Kumbier

Building and expanding on principles of statistics, machine learning, and scientific inquiry, we propose the predictability, computability, and stability (PCS) framework for veridical data science. Our framework, comprised of both a workflow and documentation, aims to provide responsible, reliable, reproducible, and transparent results across the entire data science life cycle. The PCS workflow uses predictability as a reality check and considers the importance of computation in data collection/storage and algorithm design. It augments predictability and computability with an overarching stability principle for the data science life cycle. Stability expands on statistical uncertainty considerations to assess how human judgment calls impact data results through data and model/algorithm perturbations. Moreover, we develop inference procedures that build on PCS, namely PCS perturbation intervals and PCS hypothesis testing, to investigate the stability of data results relative to problem formulation, data cleaning, modeling decisions, and interpretations. We illustrate PCS inference through neuroscience and genomics projects of our own and others and compare it to existing methods in high dimensional, sparse linear model simulations. Over a wide range of misspecified simulation models, PCS inference demonstrates favorable performance in terms of ROC curves. Finally, we propose PCS documentation based on R Markdown or Jupyter Notebook, with publicly available, reproducible codes and narratives to back up human choices made throughout an analysis. The PCS workflow and documentation are demonstrated in a genomics case study available on Zenodo.

📄 PDF Abstract BibTeX arXiv:1901.08152

Code (0)

등록된 구현이 없습니다.

Tasks

Two-sample testing

Similar Papers 제목 키워드 기반

VDSAgents: A PCS-Guided Multi-Agent System for Veridical Data Science Automation

2025-10-28 · Yunxuan Jiang, Silan Hu, Xiaoning Wang, Yuanyuan Zhang 외 arxiv

Large language models (LLMs) become increasingly integrated into data science workflows for automated system design. However, these LLM-driven data science systems rely solely on the internal reasoning of LLMs, lacking g…

Feature Engineering

Veridical Data Science for Medical Foundation Models

2024-09-15 · Ahmed Alaa, Bin Yu

The advent of foundation models (FMs) such as large language models (LLMs) has led to a cultural shift in data science, both in medicine and beyond. This shift involves moving away from specialized predictive models trai…

Decision Making

Next Waves in Veridical Network Embedding

2020-07-10 · Owen G. Ward, Zhen Huang, Andrew Davison, Tian Zheng

Embedding nodes of a large network into a metric (e.g., Euclidean) space has become an area of active research in statistical machine learning, which has found applications in natural and social sciences. Generally, a re…

Community DetectionLink PredictionNetwork EmbeddingNode Classification

How well do NLI models capture verb veridicality?

2019-11-01 · IJCNLP 2019 11 · Alexis Ross, Ellie Pavlick

In natural language inference (NLI), contexts are considered veridical if they allow us to infer that their underlying propositions make true claims about the real world. We investigate whether a state-of-the-art natural…

Natural Language InferenceNegationSentence

Do You Believe It Happened? Assessing Chinese Readers' Veridicality Judgments

2020-05-01 · LREC 2020 5 · Yu-Yun Chang, Shu-Kai Hsieh

This work collects and studies Chinese readers{'} veridicality judgments to news events (whether an event is viewed as happening or not). For instance, in {``}The FBI alleged in court documents that Zazi had admitted hav…