A study of machine-learning-based approaches to extract clinical entities and their assertions from discharge summaries.

Objective: The authors' goal was to develop and evaluate machine-learning-based approaches to extracting clinical entities-including medical problems, tests, and treatments, as well as their asserted status-from hospital discharge summaries written using natural language. This project was part of th...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of the American Medical Informatics Association Vol. 18; no. 5; pp. 601 - 607
Autores principales: Jiang M, Chen Y, Liu M, Rosenbloom ST, Mani S, Denny JC, Xu H, Jiang, Min, Chen, Yukun, Liu, Mei, Rosenbloom, S Trent, Mani, Subramani, Denny, Joshua C, Xu, Hua
Formato: research Journal Article
Publicado: Oxford University Press / USA Sep2011
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=104577234&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 104577234
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        10675027
        FZ9
      jtl: Journal of the American Medical Informatics Association
      issn: 10675027
      maglogo: N
    pubinfo:
      dt: Sep2011
      vid: 18
      iid: 5
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        104577234
        NLM21508414
        2011241432
        10.1136/amiajnl-2011-000163
        NLM21508414
        PMC3168315
        104577234
      ppf: 601
      ppct: 6
      formats:
      tig:
        atl: A study of machine-learning-based approaches to extract clinical entities and their assertions from discharge summaries.
      aug:
        au:
          Jiang M
          Chen Y
          Liu M
          Rosenbloom ST
          Mani S
          Denny JC
          Xu H
          Jiang, Min
          Chen, Yukun
          Liu, Mei
          Rosenbloom, S Trent
          Mani, Subramani
          Denny, Joshua C
          Xu, Hua
        affil: Department of Biomedical Informatics, Vanderbilt University, School of Medicine, Nashville, Tennessee 37232, USA
      sug:
        subj:
          Data Mining Classification
          Decision Support Systems, Clinical Classification
          Electronic Health Records Classification
          Natural Language Processing
          Patient Discharge
          Information Science
          Artificial Intelligence
          Semantics
          Vocabulary, Controlled
      ab: Objective: The authors' goal was to develop and evaluate machine-learning-based approaches to extracting clinical entities-including medical problems, tests, and treatments, as well as their asserted status-from hospital discharge summaries written using natural language. This project was part of the 2010 Center of Informatics for Integrating Biology and the Bedside/Veterans Affairs (VA) natural-language-processing challenge.Design: The authors implemented a machine-learning-based named entity recognition system for clinical text and systematically evaluated the contributions of different types of features and ML algorithms, using a training corpus of 349 annotated notes. Based on the results from training data, the authors developed a novel hybrid clinical entity extraction system, which integrated heuristic rule-based modules with the ML-base named entity recognition module. The authors applied the hybrid system to the concept extraction and assertion classification tasks in the challenge and evaluated its performance using a test data set with 477 annotated notes.Measurements: Standard measures including precision, recall, and F-measure were calculated using the evaluation script provided by the Center of Informatics for Integrating Biology and the Bedside/VA challenge organizers. The overall performance for all three types of clinical entities and all six types of assertions across 477 annotated notes were considered as the primary metric in the challenge.Results and Discussion: Systematic evaluation on the training set showed that Conditional Random Fields outperformed Support Vector Machines, and semantic information from existing natural-language-processing systems largely improved performance, although contributions from different types of features varied. The authors' hybrid entity extraction system achieved a maximum overall F-score of 0.8391 for concept extraction (ranked second) and 0.9313 for assertion classification (ranked fourth, but not statistically different than the first three systems) on the test data set in the challenge.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N