Disease named entity recognition using semisupervised learning and conditional random fields.

Information extraction is an important text-mining task that aims at extracting prespecified types of information from large text collections and making them available in structured representations such as databases. In the biomedical domain, information extraction can be applied to help biologists...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of the American Society for Information Science & Technology Vol. 62; no. 4; pp. 727 - 738
Autores principales: Suakkaphong, Nichalin, Zhang, Zhu, Chen, Hsinchun
Formato: algorithm research tables/charts Journal Article
Publicado: Wiley-Blackwell Apr2011
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=104838808&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 104838808
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        15322882
        IGD
      jtl: Journal of the American Society for Information Science & Technology
      issn: 15322882
      maglogo: Y
    pubinfo:
      dt: Apr2011
      vid: 62
      iid: 4
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        104838808
        59161845
        10.1002/asi.21488
        104838808
      ppf: 727
      ppct: 11
      formats:
      tig:
        atl: Disease named entity recognition using semisupervised learning and conditional random fields.
      aug:
        au:
          Suakkaphong, Nichalin
          Zhang, Zhu
          Chen, Hsinchun
        affil: Management Information Systems Department, Eller College of Management, University of Arizona, Tucson, AZ 85721
      sug:
        subj:
          Medical Literature
          Information Retrieval Methods
          Disease
          Human
          Algorithms
          Knowledge
          Classification
          Medline
          Data Mining
      ab: Information extraction is an important text-mining task that aims at extracting prespecified types of information from large text collections and making them available in structured representations such as databases. In the biomedical domain, information extraction can be applied to help biologists make the most use of their digital-literature archives. Currently, there are large amounts of biomedical literature that contain rich information about biomedical substances. Extracting such knowledge requires a good named entity recognition technique. In this article, we combine conditional random fields (CRFs), a state-of-the-art sequence-labeling algorithm, with two semisupervised learning techniques, bootstrapping and feature sampling, to recognize disease names from biomedical literature. Two data-processing strategies for each technique also were analyzed: one sequentially processing unlabeled data partitions and another one processing unlabeled data partitions in a round-robin fashion. The experimental results showed the advantage of semisupervised learning techniques given limited labeled training data. Specifically, CRFs with bootstrapping implemented in sequential fashion outperformed strictly supervised CRFs for disease name recognition. The project was supported by NIH/NLM Grant R33 LM07299-01, 2002-2005.
      pubtype: Academic Journal
      doctype:
        algorithm
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N