Federated learning of predictive models from federated Electronic Health Records.

Background: In an era of "big data," computationally efficient and privacy-aware solutions for large-scale machine learning problems become crucial, especially in the healthcare domain, where large amounts of data are stored in different locations and owned by different entities. Past research has b...

Descripción completa

Detalles Bibliográficos
Publicado en:International Journal of Medical Informatics Vol. 112; pp. 59 - 68
Autores principales: Brisimi, Theodora S., Chen, Ruidi, Mela, Theofanie, Olshevsky, Alex, Paschalidis, Ioannis Ch., Shi, Wei
Formato: research Journal Article
Publicado: Elsevier B.V. Apr2018
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=128227021&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 128227021
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        13865056
        JR4
      jtl: International Journal of Medical Informatics
      issn: 13865056
      maglogo: N
    pubinfo:
      dt: Apr2018
      vid: 112
      pid: 467
      pub: Elsevier B.V.
      place: New York, New York
    artinfo:
      ui:
        128227021
        128227021
        NLM29500022
        128227021
        10.1016/j.ijmedinf.2018.01.007
        NLM29500022
        128227021
      ppf: 59
      ppct: 9
      formats:
      tig:
        atl: Federated learning of predictive models from federated Electronic Health Records.
      aug:
        au:
          Brisimi, Theodora S.
          Chen, Ruidi
          Mela, Theofanie
          Olshevsky, Alex
          Paschalidis, Ioannis Ch.
          Shi, Wei
        affil: Department of Electrical & Computer Engineering, and Division of Systems Engineering, Boston University, 8 Saint Mary’s St., Boston, MA 02215, United States
      sug:
        subj:
          Algorithms
          Hospitalization Statistics and Numerical Data
          ROC Curve
          Resource Databases
          Human
          Validation Studies
          Comparative Studies
          Evaluation Research
          Multicenter Studies
          Questionnaires
      ab: Background: In an era of "big data," computationally efficient and privacy-aware solutions for large-scale machine learning problems become crucial, especially in the healthcare domain, where large amounts of data are stored in different locations and owned by different entities. Past research has been focused on centralized algorithms, which assume the existence of a central data repository (database) which stores and can process the data from all participants. Such an architecture, however, can be impractical when data are not centrally located, it does not scale well to very large datasets, and introduces single-point of failure risks which could compromise the integrity and privacy of the data. Given scores of data widely spread across hospitals/individuals, a decentralized computationally scalable methodology is very much in need.Objective: We aim at solving a binary supervised classification problem to predict hospitalizations for cardiac events using a distributed algorithm. We seek to develop a general decentralized optimization framework enabling multiple data holders to collaborate and converge to a common predictive model, without explicitly exchanging raw data.Methods: We focus on the soft-margin l1-regularized sparse Support Vector Machine (sSVM) classifier. We develop an iterative cluster Primal Dual Splitting (cPDS) algorithm for solving the large-scale sSVM problem in a decentralized fashion. Such a distributed learning scheme is relevant for multi-institutional collaborations or peer-to-peer applications, allowing the data holders to collaborate, while keeping every participant's data private.Results: We test cPDS on the problem of predicting hospitalizations due to heart diseases within a calendar year based on information in the patients Electronic Health Records prior to that year. cPDS converges faster than centralized methods at the cost of some communication between agents. It also converges faster and with less communication overhead compared to an alternative distributed algorithm. In both cases, it achieves similar prediction accuracy measured by the Area Under the Receiver Operating Characteristic Curve (AUC) of the classifier. We extract important features discovered by the algorithm that are predictive of future hospitalizations, thus providing a way to interpret the classification results and inform prevention efforts.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N