High-throughput phenotyping with temporal sequences.
Objective: High-throughput electronic phenotyping algorithms can accelerate translational research using data from electronic health record (EHR) systems. The temporal information buried in EHRs is often underutilized in developing computational phenotypic definitions. This study aims to develop a h...
| Publicado en: | Journal of the American Medical Informatics Association Vol. 28; no. 4; pp. 772 - 782 |
|---|---|
| Autores principales: | , , |
| Formato: | Journal Article |
| Publicado: |
Oxford University Press / USA
Apr2021
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=149414957&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 149414957 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 10675027 FZ9 jtl: Journal of the American Medical Informatics Association issn: 10675027 maglogo: N pubinfo: dt: Apr2021 vid: 28 iid: 4 pid: 622 pub: Oxford University Press / USA artinfo: ui: 149414957 149414957 NLM33313899 10.1093/jamia/ocaa288 NLM33313899 149414957 ppf: 772 ppct: 10 formats: tig: atl: High-throughput phenotyping with temporal sequences. aug: au: Estiri, Hossein Strasser, Zachary H Murphy, Shawn N affil: Harvard Medical School , Boston, Massachusetts, USA sug: subj: Algorithms Data Mining Methods Diagnosis Drug Therapy Time Factors Arthritis Impact Measurement Scales Questionnaires ab: Objective: High-throughput electronic phenotyping algorithms can accelerate translational research using data from electronic health record (EHR) systems. The temporal information buried in EHRs is often underutilized in developing computational phenotypic definitions. This study aims to develop a high-throughput phenotyping method, leveraging temporal sequential patterns from EHRs.Materials and Methods: We develop a representation mining algorithm to extract 5 classes of representations from EHR diagnosis and medication records: the aggregated vector of the records (aggregated vector representation), the standard sequential patterns (sequential pattern mining), the transitive sequential patterns (transitive sequential pattern mining), and 2 hybrid classes. Using EHR data on 10 phenotypes from the Mass General Brigham Biobank, we train and validate phenotyping algorithms.Results: Phenotyping with temporal sequences resulted in a superior classification performance across all 10 phenotypes compared with the standard representations in electronic phenotyping. The high-throughput algorithm's classification performance was superior or similar to the performance of previously published electronic phenotyping algorithms. We characterize and evaluate the top transitive sequences of diagnosis records paired with the records of risk factors, symptoms, complications, medications, or vaccinations.Discussion: The proposed high-throughput phenotyping approach enables seamless discovery of sequential record combinations that may be difficult to assume from raw EHR data. Transitive sequences offer more accurate characterization of the phenotype, compared with its individual components, and reflect the actual lived experiences of the patients with that particular disease.Conclusion: Sequential data representations provide a precise mechanism for incorporating raw EHR records into downstream machine learning. Our approach starts with user interpretability and works backward to the technology. pubtype: Academic Journal doctype: Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|