Develop and validate a computable phenotype for identifying alcohol-use disorder patients using structure and unstructured EHR data.

Background Alcohol Use Disorder (AUD) drives significant morbidity through alcohol-related liver disease. Accurate AUD identification in electronic health records is critical for research and care delivery, yet International Classification of Diseases (ICD) code-based algorithms miss many cases whil...

Descripción completa

Detalles Bibliográficos
Publicado en:Alcohol & Alcoholism Vol. 61; no. 1; pp. 1 - 10
Autores principales: Dai, Hao, Tapper, Elliot B, Zhao, Lili, Scheiffele, Grant D, Li, Xiaohan, Ramon, Ronald, He, Xing, Lee, Yao An, Guo, Jingchuan, Bian, Jiang
Formato: research tables/charts Journal Article
Publicado: Oxford University Press / USA Jan2026
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=190916540&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 190916540
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        07350414
        FA3
      jtl: Alcohol & Alcoholism
      issn: 07350414
      maglogo: N
    pubinfo:
      dt: Jan2026
      vid: 61
      iid: 1
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        190916540
        190916540
        190916540
        10.1093/alcalc/agaf086
        190916540
      ppf: 1
      ppct: 9
      formats:
      tig:
        atl: Develop and validate a computable phenotype for identifying alcohol-use disorder patients using structure and unstructured EHR data.
      aug:
        au:
          Dai, Hao
          Tapper, Elliot B
          Zhao, Lili
          Scheiffele, Grant D
          Li, Xiaohan
          Ramon, Ronald
          He, Xing
          Lee, Yao An
          Guo, Jingchuan
          Bian, Jiang
        affil: Department of Biostatistics & Health Data Science, Indiana University School of Medicine, 410 W 10th St, Indianapolis, IN 46202, United States
      sug:
        subj:
          Alcoholism Diagnosis
          Persons with Alcoholism Psychosocial Factors
          Patient Identification Methods
          Phenotype
          Electronic Health Records
          Natural Language Processing
          Human
          Female
          Male
          Adult
          Middle Age
          Aged
          Validation Studies
          Prospective Studies
          Sensitivity and Specificity
          Predictive Value of Tests
          International Classification of Diseases
          Algorithms
          Data Mining
          Medical Informatics
          Registries, Disease
          Record Review
          Random Sample
          Descriptive Statistics
          Confidence Intervals
          Funding Source
          Adult: 19-44 years
          Middle Aged: 45-64 years
          Aged: 65+ years
          Female
          Male
      ab: Background Alcohol Use Disorder (AUD) drives significant morbidity through alcohol-related liver disease. Accurate AUD identification in electronic health records is critical for research and care delivery, yet International Classification of Diseases (ICD) code-based algorithms miss many cases while manual review is impractical at scale. Computable phenotypes (CPs) integrating structured and unstructured EHR data offer a scalable solution. Methods Using University of Florida Health's Integrated Data Repository covering two million patients, we developed AUD CPs through a two-step process. First, candidate cohorts were identified using AUD-related ICD codes, medications, and keyword searches across structured and unstructured data. Second, rule-based combinations were iteratively refined through manual chart review. Final algorithms were evaluated against gold-standard chart review, measuring sensitivity, positive predictive value (PPV), and F1-score, then validated in an independent testing set and an external dataset. Results The F1-optimized CP achieved an F1-score of.87 (sensitivity:.98, PPV:.78) in the testing set, while the precision-optimized CP achieved PPV of.9 (sensitivity:.68, F1-score:.77). Minimal performance attenuation between training and testing sets demonstrated robustness and generalizability. Both CPs substantially outperformed restricted AUD-specific ICD code-based approaches. Conclusions CPs integrating structured and unstructured EHR data enable accurate, reproducible AUD identification, surpassing traditional AUD-specific ICD-based methods. This approach facilitates efficient cohort construction for clinical research, public health surveillance, and quality improvement initiatives targeting AUD and its consequences, addressing a critical gap in identifying patients who may benefit from screening and intervention.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N