Causal discoveries for high dimensional mixed data.

Causal relationships are of crucial importance for biological and medical research. Algorithms have been proposed for causal structure learning with graphical visualizations. While much of the literature focuses on biological studies where data often follow the same distribution, for example, the no...

Full description

Bibliographic Details
Published in:Statistics in Medicine Vol. 41; no. 24; pp. 4924 - 4941
Main Authors: Cai, Zhanrui, Xi, Dong, Zhu, Xuan, Li, Runze
Format: equations & formulas research tables/charts Journal Article
Published: Wiley-Blackwell 10/30/2022
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=159688165&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 159688165
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        02776715
        2DZ
      jtl: Statistics in Medicine
      issn: 02776715
      maglogo: Y
    pubinfo:
      dt: 10/30/2022
      vid: 41
      iid: 24
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        159688165
        158475360
        159688165
        NLM35968913
        159688165
        10.1002/sim.9544
        NLM35968913
        159688165
      ppf: 4924
      ppct: 17
      formats:
      tig:
        atl: Causal discoveries for high dimensional mixed data.
      aug:
        au:
          Cai, Zhanrui
          Xi, Dong
          Zhu, Xuan
          Li, Runze
        affil: Department of Statistics, Iowa State University, Ames Iowa,, USA
      sug:
        subj:
          Algorithms
          Causal Attribution
          Statistics
          Computer Simulation
          Human
      ab: Causal relationships are of crucial importance for biological and medical research. Algorithms have been proposed for causal structure learning with graphical visualizations. While much of the literature focuses on biological studies where data often follow the same distribution, for example, the normal distribution for all variables, challenges emerge from epidemiological and clinical studies where data are often mixed with continuous, binary, and ordinal variables. We propose to use a mixed latent Gaussian copula model to estimate the underlying correlation structure via the rank correlation for mixed data. This correlation structure is then incorporated into a popular causal discovery algorithm, the PC algorithm, to identify causal structures. The proposed algorithm, called the latent-PC algorithm, is able to discover the true causal structure consistently under mild conditions in high dimensional settings. From simulation studies, the latent-PC algorithm delivers a competitive performance in terms of a similar or higher true positive rate and a similar or lower false positive rate, compared with other variants of the PC algorithm. In the high dimensional settings where the number of variables is more than the number of observations, the causal graphs identified by the latent-PC algorithm are closer to the true causal structures, compared to other competing algorithms. Further, we demonstrate the utility of the latent-PC algorithm in a real dataset for hepatocellular carcinoma. Causal structures for patient survival are visualized and connected with clinical interpretations in the literature.
      pubtype: Academic Journal
      doctype:
        equations & formulas
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N