Development and Validation of an Algorithm to Identify Nonalcoholic Fatty Liver Disease in the Electronic Medical Record.

Background and Aims: Nonalcoholic fatty liver disease (NAFLD) is the most common cause of chronic liver disease worldwide. Risk factors for NAFLD disease progression and liver-related outcomes remain incompletely understood due to the lack of computational identification methods. The present study s...

Descripción completa

Detalles Bibliográficos
Publicado en:Digestive Diseases & Sciences Vol. 61; no. 3; pp. 913 - 920
Autores principales: Corey, Kathleen, Kartoun, Uri, Zheng, Hui, Shaw, Stanley, Corey, Kathleen E, Shaw, Stanley Y
Formato: research Journal Article
Publicado: Springer Nature Mar2016
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=113205114&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 113205114
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        01632116
        1VP
      jtl: Digestive Diseases & Sciences
      issn: 01632116
      maglogo: N
    pubinfo:
      dt: Mar2016
      vid: 61
      iid: 3
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        113205114
        113205114
        NLM26537487
        113205114
        10.1007/s10620-015-3952-x
        NLM26537487
        PMC4761309 [Available on 03/01/17]
        113205114
      ppf: 913
      ppct: 7
      formats:
      tig:
        atl: Development and Validation of an Algorithm to Identify Nonalcoholic Fatty Liver Disease in the Electronic Medical Record.
      aug:
        au:
          Corey, Kathleen
          Kartoun, Uri
          Zheng, Hui
          Shaw, Stanley
          Corey, Kathleen E
          Shaw, Stanley Y
        affil: Biostatistics Center, Massachusetts General Hospital, Boston USA
      sug:
        subj:
          Natural Language Processing
          Nonalcoholic Fatty Liver Disease Blood
          Electronic Health Records
          Algorithms
          Nonalcoholic Fatty Liver Disease Epidemiology
          Triglycerides Blood
          Alanine Aminotransferase Blood
          Male
          Data Collection
          Adult
          Diabetes Mellitus Epidemiology
          International Classification of Diseases
          Prevalence
          Female
          Aged
          Prospective Studies
          Biopsy
          Aspartate Aminotransferase Blood
          United States
          Human
          Middle Age
          Logistic Regression
          Validation Studies
          Comparative Studies
          Evaluation Research
          Multicenter Studies
          Arthritis Impact Measurement Scales
          Scales
          Adult: 19-44 years
          Aged: 65+ years
          Middle Aged: 45-64 years
          Male
          Female
      ab: Background and Aims: Nonalcoholic fatty liver disease (NAFLD) is the most common cause of chronic liver disease worldwide. Risk factors for NAFLD disease progression and liver-related outcomes remain incompletely understood due to the lack of computational identification methods. The present study sought to design a classification algorithm for NAFLD within the electronic medical record (EMR) for the development of large-scale longitudinal cohorts.Methods: We implemented feature selection using logistic regression with adaptive LASSO. A training set of 620 patients was randomly selected from the Research Patient Data Registry at Partners Healthcare. To assess a true diagnosis for NAFLD we performed chart reviews and considered either a documentation of a biopsy or a clinical diagnosis of NAFLD. We included in our model variables laboratory measurements, diagnosis codes, and concepts extracted from medical notes. Variables with P < 0.05 were included in the multivariable analysis.Results: The NAFLD classification algorithm included number of natural language mentions of NAFLD in the EMR, lifetime number of ICD-9 codes for NAFLD, and triglyceride level. This classification algorithm was superior to an algorithm using ICD-9 data alone with AUC of 0.85 versus 0.75 (P < 0.0001) and leads to the creation of a new independent cohort of 8458 individuals with a high probability for NAFLD.Conclusions: The NAFLD classification algorithm is superior to ICD-9 billing data alone. This approach is simple to develop, deploy, and can be applied across different institutions to create EMR-based cohorts of individuals with NAFLD.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N