Skip to main page content
U.S. flag

An official website of the United States government

Dot gov

The .gov means it’s official.
Federal government websites often end in .gov or .mil. Before sharing sensitive information, make sure you’re on a federal government site.

Https

The site is secure.
The https:// ensures that you are connecting to the official website and that any information you provide is encrypted and transmitted securely.

Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
. 2020 Feb 25;117(8):3920-3929.
doi: 10.1073/pnas.1901326117. Epub 2020 Feb 13.

Veridical data science

Affiliations

Veridical data science

Bin Yu et al. Proc Natl Acad Sci U S A. .

Abstract

Building and expanding on principles of statistics, machine learning, and scientific inquiry, we propose the predictability, computability, and stability (PCS) framework for veridical data science. Our framework, composed of both a workflow and documentation, aims to provide responsible, reliable, reproducible, and transparent results across the data science life cycle. The PCS workflow uses predictability as a reality check and considers the importance of computation in data collection/storage and algorithm design. It augments predictability and computability with an overarching stability principle. Stability expands on statistical uncertainty considerations to assess how human judgment calls impact data results through data and model/algorithm perturbations. As part of the PCS workflow, we develop PCS inference procedures, namely PCS perturbation intervals and PCS hypothesis testing, to investigate the stability of data results relative to problem formulation, data cleaning, modeling decisions, and interpretations. We illustrate PCS inference through neuroscience and genomics projects of our own and others. Moreover, we demonstrate its favorable performance over existing methods in terms of receiver operating characteristic (ROC) curves in high-dimensional, sparse linear model simulations, including a wide range of misspecified models. Finally, we propose PCS documentation based on R Markdown or Jupyter Notebook, with publicly available, reproducible codes and narratives to back up human choices made throughout an analysis. The PCS workflow and documentation are demonstrated in a genomics case study available on Zenodo.

Keywords: computation; data science; prediction; stability.

PubMed Disclaimer

Conflict of interest statement

The authors declare no competing interest.

Figures

Fig. 1.
Fig. 1.
The data science life cycle. The PCS workflow considers predictability, computability, and stability at every step, with a strong emphasis on stability.
Fig. 2.
Fig. 2.
ROC curves showing true positive rate (TPR) and false positive rate (FPR) for feature selection in a linear model setting with n=250 observations. Each plot corresponds to a different generative model. Color corresponds to method of inference: red, PCS; blue, selective inference; green, linear model asymptotic normality. Error bars give the 10th and 90th percentiles over replicates.
Fig. 3.
Fig. 3.
Assumptions made throughout the DSLC allow researchers to use models such as decision trees, neural networks, or probability distributions as an approximation of reality, which may include physical, chemical, or biological laws. Narratives provided in PCS documentation can help justify assumptions to connect these two worlds.

References

    1. Murdoch W. J., Singh C., Kumbier K., Abbasi-Asl R., Yu B., Definitions, methods, and applications in interpretable machine learning. Proc. Natl. Acad. Sci. U.S.A. 116, 22071–22080 (2019). - PMC - PubMed
    1. Stark P. B., Saltelli A., Cargo-cult statistics and scientific crisis. Significance 15, 40–43 (2018).
    1. Ioannidis J. P. A., Why most published research findings are false. PLoS Med. 2, e124 (2005). - PMC - PubMed
    1. Popper K. R., The Logic of Scientific Discovery (University Press, 1959).
    1. Breiman L., et al. , Statistical modeling: The two cultures (with comments and a rejoinder by the author). Stat. Sci. 16, 199–231 (2001).

Publication types