Optimizing the synthesis of clinical trial data using sequential trees.

Objective: With the growing demand for sharing clinical trial data, scalable methods to enable privacy protective access to high-utility data are needed. Data synthesis is one such method. Sequential trees are commonly used to synthesize health data. It is hypothesized that the utility of the genera...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of the American Medical Informatics Association Vol. 28; no. 1; pp. 3 - 14
Autores principales: Emam, Khaled El, Mosquera, Lucy, Zheng, Chaoyi
Formato: research Journal Article
Publicado: Oxford University Press / USA Jan2021
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=148191059&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 148191059
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        10675027
        FZ9
      jtl: Journal of the American Medical Informatics Association
      issn: 10675027
      maglogo: N
    pubinfo:
      dt: Jan2021
      vid: 28
      iid: 1
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        148191059
        148191059
        NLM33186440
        148191059
        10.1093/jamia/ocaa249
        NLM33186440
        148191059
      ppf: 3
      ppct: 11
      formats:
      tig:
        atl: Optimizing the synthesis of clinical trial data using sequential trees.
      aug:
        au:
          Emam, Khaled El
          Mosquera, Lucy
          Zheng, Chaoyi
        affil: School of Epidemiology and Public Health, Faculty of Medicine, University of Ottawa , Ottawa, Ontario, Canada
      sug:
        subj:
          Data Collection
          Clinical Trials
          Communication Methods
          Analysis of Variance
          Algorithms
          Human
          Privacy and Confidentiality
          Comparative Studies
          Multicenter Studies
          Evaluation Research
          Validation Studies
          Impact of Events Scale
          Scales
      ab: Objective: With the growing demand for sharing clinical trial data, scalable methods to enable privacy protective access to high-utility data are needed. Data synthesis is one such method. Sequential trees are commonly used to synthesize health data. It is hypothesized that the utility of the generated data is dependent on the variable order. No assessments of the impact of variable order on synthesized clinical trial data have been performed thus far. Through simulation, we aim to evaluate the variability in the utility of synthetic clinical trial data as variable order is randomly shuffled and implement an optimization algorithm to find a good order if variability is too high.Materials and Methods: Six oncology clinical trial datasets were evaluated in a simulation. Three utility metrics were computed comparing real and synthetic data: univariate similarity, similarity in multivariate prediction accuracy, and a distinguishability metric. Particle swarm was implemented to optimize variable order, and was compared with a curriculum learning approach to ordering variables.Results: As the number of variables in a clinical trial dataset increases, there is a pattern of a marked increase in variability of data utility with order. Particle swarm with a distinguishability hinge loss ensured adequate utility across all 6 datasets. The hinge threshold was selected to avoid overfitting which can create a privacy problem. This was superior to curriculum learning in terms of utility.Conclusions: The optimization approach presented in this study gives a reliable way to synthesize high-utility clinical trial datasets.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N