Missing data imputation and synthetic data simulation through modeling graphical probabilistic dependencies between variables (ModGraProDep): An application to breast cancer survival.

Background: Two common issues may arise in certain population-based breast cancer (BC) survival studies: I) missing values in a survivals' predictive variable, such as "Stage" at diagnosis, and II) small sample size due to "imbalance class problem" in certain subsets of patients, demanding data mode...

Descripción completa

Detalles Bibliográficos
Publicado en:Artificial Intelligence in Medicine Vol. 107
Autores principales: Vilardell, Mireia, Buxó, Maria, Clèries, Ramon, Martínez, José Miguel, Garcia, Gemma, Ameijide, Alberto, Font, Rebeca, Civit, Sergi, Marcos-Gragera, Rafael, Vilardell, Maria Loreto, Carulla, Marià, Espinàs, Josep Alfons, Galceran, Jaume, Izquierdo, Angel, Borràs, Josep Ma
Formato: research tables/charts Journal Article
Publicado: Elsevier B.V. Jul2020
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=145474617&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 145474617
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        09333657
        3HY
      jtl: Artificial Intelligence in Medicine
      issn: 09333657
      maglogo: N
    pubinfo:
      dt: Jul2020
      vid: 107
      pid: 1004
      pub: Elsevier B.V.
    artinfo:
      ui:
        145474617
        145474617
        NLM32828436
        145474617
        10.1016/j.artmed.2020.101875
        NLM32828436
        145474617
      ppct: 1
      formats:
      tig:
        atl: Missing data imputation and synthetic data simulation through modeling graphical probabilistic dependencies between variables (ModGraProDep): An application to breast cancer survival.
      aug:
        au:
          Vilardell, Mireia
          Buxó, Maria
          Clèries, Ramon
          Martínez, José Miguel
          Garcia, Gemma
          Ameijide, Alberto
          Font, Rebeca
          Civit, Sergi
          Marcos-Gragera, Rafael
          Vilardell, Maria Loreto
          Carulla, Marià
          Espinàs, Josep Alfons
          Galceran, Jaume
          Izquierdo, Angel
          Borràs, Josep Ma
        affil: School of Medicine, University of Girona (UdG), Girona, Spain
      sug:
        subj:
          Breast Neoplasms
          Survival Analysis
          Algorithms
          Female
          Computer Simulation
          Data Collection
          Human
          Comparative Studies
          Multicenter Studies
          Evaluation Research
          Validation Studies
          Female
      ab: Background: Two common issues may arise in certain population-based breast cancer (BC) survival studies: I) missing values in a survivals' predictive variable, such as "Stage" at diagnosis, and II) small sample size due to "imbalance class problem" in certain subsets of patients, demanding data modeling/simulation methods.Methods: We present a procedure, ModGraProDep, based on graphical modeling (GM) of a dataset to overcome these two issues. The performance of the models derived from ModGraProDep is compared with a set of frequently used classification and machine learning algorithms (Missing Data Problem) and with oversampling algorithms (Synthetic Data Simulation). For the Missing Data Problem we assessed two scenarios: missing completely at random (MCAR) and missing not at random (MNAR). Two validated BC datasets provided by the cancer registries of Girona and Tarragona (northeastern Spain) were used.Results: In both MCAR and MNAR scenarios all models showed poorer prediction performance compared to three GM models: the saturated one (GM.SAT) and two with penalty factors on the partial likelihood (GM.K1 and GM.TEST). However, GM.SAT predictions could lead to non-reliable conclusions in BC survival analysis. Simulation of a "synthetic" dataset derived from GM.SAT could be the worst strategy, but the use of the remaining GMs models could be better than oversampling.Conclusion: Our results suggest the use of the GM-procedure presented for one-variable imputation/prediction of missing data and for simulating "synthetic" BC survival datasets. The "synthetic" datasets derived from GMs could be also used in clinical applications of cancer survival data such as predictive risk analysis.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N