Multiple imputation and test-wise deletion for causal discovery with incomplete cohort data.

Causal discovery algorithms estimate causal graphs from observational data. This can provide a valuable complement to analyses focusing on the causal relation between individual treatment-outcome pairs. Constraint-based causal discovery algorithms rely on conditional independence testing when buildi...

Full description

Bibliographic Details
Published in:Statistics in Medicine Vol. 41; no. 23; pp. 4716 - 4744
Main Authors: Witte, Janine, Foraita, Ronja, Didelez, Vanessa
Format: equations & formulas research tables/charts Journal Article
Published: Wiley-Blackwell Oct2022
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=159193679&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 159193679
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        02776715
        2DZ
      jtl: Statistics in Medicine
      issn: 02776715
      maglogo: Y
    pubinfo:
      dt: Oct2022
      vid: 41
      iid: 23
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        159193679
        158256703
        159193679
        NLM35908775
        159193679
        10.1002/sim.9535
        NLM35908775
        159193679
      ppf: 4716
      ppct: 28
      formats:
      tig:
        atl: Multiple imputation and test-wise deletion for causal discovery with incomplete cohort data.
      aug:
        au:
          Witte, Janine
          Foraita, Ronja
          Didelez, Vanessa
        affil: Leibniz Institute for Prevention Research and Epidemiology – BIPS, Bremen, Germany
      sug:
        subj:
          Study Design
          Algorithms
          Causal Attribution
          Child
          Prospective Studies
          Funding Source
          Human
          Child: 6-12 years
      ab: Causal discovery algorithms estimate causal graphs from observational data. This can provide a valuable complement to analyses focusing on the causal relation between individual treatment-outcome pairs. Constraint-based causal discovery algorithms rely on conditional independence testing when building the graph. Until recently, these algorithms have been unable to handle missing values. In this article, we investigate two alternative solutions: test-wise deletion and multiple imputation. We establish necessary and sufficient conditions for the recoverability of causal structures under test-wise deletion, and argue that multiple imputation is more challenging in the context of causal discovery than for estimation. We conduct an extensive comparison by simulating from benchmark causal graphs: as one might expect, we find that test-wise deletion and multiple imputation both clearly outperform list-wise deletion and single imputation. Crucially, our results further suggest that multiple imputation is especially useful in settings with a small number of either Gaussian or discrete variables, but when the dataset contains a mix of both neither method is uniformly best. The methods we compare include random forest imputation and a hybrid procedure combining test-wise deletion and multiple imputation. An application to data from the IDEFICS cohort study on diet- and lifestyle-related diseases in European children serves as an illustrating example.
      pubtype: Academic Journal
      doctype:
        equations & formulas
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N