Disease named entity recognition using semisupervised learning and conditional random fields.
Information extraction is an important text-mining task that aims at extracting prespecified types of information from large text collections and making them available in structured representations such as databases. In the biomedical domain, information extraction can be applied to help biologists...
| Publicado en: | Journal of the American Society for Information Science & Technology Vol. 62; no. 4; pp. 727 - 738 |
|---|---|
| Autores principales: | , , |
| Formato: | algorithm research tables/charts Journal Article |
| Publicado: |
Wiley-Blackwell
Apr2011
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=104838808&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 104838808 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 15322882 IGD jtl: Journal of the American Society for Information Science & Technology issn: 15322882 maglogo: Y pubinfo: dt: Apr2011 vid: 62 iid: 4 pid: 480 pub: Wiley-Blackwell place: Malden, Massachusetts artinfo: ui: 104838808 59161845 10.1002/asi.21488 104838808 ppf: 727 ppct: 11 formats: tig: atl: Disease named entity recognition using semisupervised learning and conditional random fields. aug: au: Suakkaphong, Nichalin Zhang, Zhu Chen, Hsinchun affil: Management Information Systems Department, Eller College of Management, University of Arizona, Tucson, AZ 85721 sug: subj: Medical Literature Information Retrieval Methods Disease Human Algorithms Knowledge Classification Medline Data Mining ab: Information extraction is an important text-mining task that aims at extracting prespecified types of information from large text collections and making them available in structured representations such as databases. In the biomedical domain, information extraction can be applied to help biologists make the most use of their digital-literature archives. Currently, there are large amounts of biomedical literature that contain rich information about biomedical substances. Extracting such knowledge requires a good named entity recognition technique. In this article, we combine conditional random fields (CRFs), a state-of-the-art sequence-labeling algorithm, with two semisupervised learning techniques, bootstrapping and feature sampling, to recognize disease names from biomedical literature. Two data-processing strategies for each technique also were analyzed: one sequentially processing unlabeled data partitions and another one processing unlabeled data partitions in a round-robin fashion. The experimental results showed the advantage of semisupervised learning techniques given limited labeled training data. Specifically, CRFs with bootstrapping implemented in sequential fashion outperformed strictly supervised CRFs for disease name recognition. The project was supported by NIH/NLM Grant R33 LM07299-01, 2002-2005. pubtype: Academic Journal doctype: algorithm research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|