Empirically-derived synthetic populations to mitigate small sample sizes.

Limited sample sizes can lead to spurious modeling findings in biomedical research. The objective of this work is to present a new method to generate synthetic populations (SPs) from limited samples using matched case-control data (n = 180 pairs), considered as two separate limited samples. SPs were...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Biomedical Informatics Vol. 105
Autores principales: Fowler, Erin E., Berglund, Anders, Schell, Michael J., Sellers, Thomas A., Eschrich, Steven, Heine, John
Formato: research tables/charts Journal Article
Publicado: Academic Press Inc. May2020
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=143235051&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 143235051
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        15320464
        OMB
      jtl: Journal of Biomedical Informatics
      issn: 15320464
      maglogo: N
    pubinfo:
      dt: May2020
      vid: 105
      pid: 735
      pub: Academic Press Inc.
      place: Burlington, Massachusetts
    artinfo:
      ui:
        143235051
        143235051
        NLM32173502
        143235051
        10.1016/j.jbi.2020.103408
        NLM32173502
        143235051
      ppct: 1
      formats:
      tig:
        atl: Empirically-derived synthetic populations to mitigate small sample sizes.
      aug:
        au:
          Fowler, Erin E.
          Berglund, Anders
          Schell, Michael J.
          Sellers, Thomas A.
          Eschrich, Steven
          Heine, John
        affil: Cancer Epidemiology Department, MCC, Moffitt Cancer Center & Research Institute, 12901 Bruce B. Downs Blvd, Tampa, FL 33612, United States
      sug:
        subj:
          Study Design
          Sample Size
          Case Control Studies
          Factor Analysis
          Human
          Comparative Studies
          Multicenter Studies
          Evaluation Research
          Validation Studies
          Funding Source
      ab: Limited sample sizes can lead to spurious modeling findings in biomedical research. The objective of this work is to present a new method to generate synthetic populations (SPs) from limited samples using matched case-control data (n = 180 pairs), considered as two separate limited samples. SPs were generated with multivariate kernel density estimations (KDEs) with unconstrained bandwidth matrices. We included four continuous variables and one categorical variable for each individual. Bandwidth matrices were determined with Differential Evolution (DE) optimization by covariance comparisons. Four synthetic samples (n = 180) were derived from their respective SPs. Similarity between observed samples with synthetic samples was compared assuming their empirical probability density functions (EPDFs) were similar. EPDFs were compared with the maximum mean discrepancy (MMD) test statistic based on the Kernel Two-Sample Test. To evaluate similarity within a modeling context, EPDFs derived from the Principal Component Analysis (PCA) scores and residuals were summarized with the distance to the model in X-space (DModX) as additional comparisons. Four SPs were generated from each sample. The probability of selecting a replicate when randomly constructing synthetic samples (n = 180) was infinitesimally small. MMD tests indicated that the observed sample EPDFs were similar to the respective synthetic EPDFs. For the samples, PCA scores and residuals did not deviate significantly when compared with their respective synthetic samples. The feasibility of this approach was demonstrated by producing synthetic data at the individual level, statistically similar to the observed samples. The methodology coupled KDE with DE optimization and deployed novel similarity metrics derived from PCA. This approach could be used to generate larger-sized synthetic samples. To develop this approach into a research tool for data exploration purposes, additional evaluation with increased dimensionality is required. Moreover, given a fully specified population, the degree to which individuals can be discarded while synthesizing the respective population accurately will be investigated. When these objectives are addressed, comparisons with other techniques such as bootstrapping will be required for a complete evaluation.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N