Measuring the effect of different types of unsupervised word representations on Medical Named Entity Recognition.

Background: This work deals with Natural Language Processing applied to the clinical domain. Specifically, the work deals with a Medical Entity Recognition (MER) on Electronic Health Records (EHRs). Developing a MER system entailed heavy data preprocessing and feature engineering until Deep Neural N...

Full description

Bibliographic Details
Published in:International Journal of Medical Informatics Vol. 129; pp. 100 - 107
Main Authors: Casillas, Arantza, Ezeiza, Nerea, Goenaga, Iakes, Pérez, Alicia, Soto, Xabier
Format: research Journal Article
Published: Elsevier B.V. Sep2019
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=138293945&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 138293945
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        13865056
        JR4
      jtl: International Journal of Medical Informatics
      issn: 13865056
      maglogo: N
    pubinfo:
      dt: Sep2019
      vid: 129
      pid: 467
      pub: Elsevier B.V.
      place: New York, New York
    artinfo:
      ui:
        138293945
        138293945
        NLM31445243
        138293945
        10.1016/j.ijmedinf.2019.05.022
        NLM31445243
        138293945
      ppf: 100
      ppct: 7
      formats:
      tig:
        atl: Measuring the effect of different types of unsupervised word representations on Medical Named Entity Recognition.
      aug:
        au:
          Casillas, Arantza
          Ezeiza, Nerea
          Goenaga, Iakes
          Pérez, Alicia
          Soto, Xabier
        affil: IXA Group, University of the Basque Country (UPV-EHU), Manuel Lardizabal 1, 20080 Donostia, Spain
      sug:
        subj:
          Natural Language Processing
          Algorithms
          Neural Networks (Computer)
          Subject Headings
          Semantics
          Validation Studies
          Comparative Studies
          Evaluation Research
          Multicenter Studies
          Ferrans and Powers Quality of Life Index
          Impact of Events Scale
          Questionnaires
          Scales
          Short Portable Mental Status Questionnaire
      ab: Background: This work deals with Natural Language Processing applied to the clinical domain. Specifically, the work deals with a Medical Entity Recognition (MER) on Electronic Health Records (EHRs). Developing a MER system entailed heavy data preprocessing and feature engineering until Deep Neural Networks (DNNs) emerged. However, the quality of the word representations in terms of embedded layers is still an important issue for the inference of the DNNs.Goal: The main goal of this work is to develop a robust MER system adapting general-purpose DNNs to cope with the high lexical variability shown in EHRs. In addition, given that EHRs tend to be scarce when there are out-domain corpora available, the aim is to assess the impact of the word representations on the performance of the MER as we move to other domains. In this line, exhaustive experimentation varying information generation methods and network parameters are crucial.Methods: We adapted a general purpose sequential tagger based on Bidirectional Long-Short Term Memory cells and Conditional Random Fields (CRFs) in order to make it tolerant to high lexical variability and a limited amount of corpora. To this end, we incorporated part of speech (POS) and semantic-tag embedding layers to the word representations.Results: One of the strengths of this work is the exhaustive evaluation of dense word representations obtained varying not only the domain and genre but also the learning algorithms and their parameter settings. With the proposed method, we attained an error reduction of 1.71 (5.7%) compared to the state-of-the-art even that no preprocessing or feature engineering was used.Conclusions: Our results indicate that dense representations built taking word order into account leverage the entity extraction system. Besides, we found that using a medical corpus (not necessarily EHRs) to infer the representations improves the performance, even if it does not correspond to the same genre.
      pubtype: Academic Journal
      doctype:
        research
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N