Identifying bacterial biotope entities using sequence labeling: Performance and feature analysis.

Habitat information is important to biodiversity conservation and research. Extracting bacterial biotope entities from scientific publications is important to large scale study of the relationships between bacteria and their living environments. To facilitate the further development of robust habita...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of the Association for Information Science & Technology Vol. 69; no. 9; pp. 1134 - 1148
Autores principales: Mao, Jin, Cui, Hong
Formato: equations & formulas research tables/charts Journal Article
Publicado: Wiley-Blackwell Sep2018
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=131499571&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 131499571
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        23301635
        H6JN
      jtl: Journal of the Association for Information Science & Technology
      issn: 23301635
      maglogo: N
    pubinfo:
      dt: Sep2018
      vid: 69
      iid: 9
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        131499571
        131499571
        131499571
        10.1002/asi.24032
        131499571
      ppf: 1134
      ppct: 14
      formats:
      tig:
        atl: Identifying bacterial biotope entities using sequence labeling: Performance and feature analysis.
      aug:
        au:
          Mao, Jin
          Cui, Hong
        affil: Center for Studies of Information Resources, Wuhan University, 299 Bayi St., Wuhan, Hubei Province, 430072, China
      sug:
        subj:
          Bacteria
          Data Mining
          Machine Learning
          Human
          Experimental Studies
          Funding Source
      ab: Habitat information is important to biodiversity conservation and research. Extracting bacterial biotope entities from scientific publications is important to large scale study of the relationships between bacteria and their living environments. To facilitate the further development of robust habitat text mining systems for biodiversity, following the BioNLP task framework, three sequence labeling techniques, CRFs (Conditional Random Fields), MEMM (Maximum Entropy Markov Model) and SVMhmm (Support Vector Machine) and one classifier, SVMmulticlass, are compared on their performance in identifying three types of bacterial biotope entities: bacteria, habitats and geographical locations. The effectiveness of a variety of basic word formation features, syntactic features, and semantic features are exploited and compared for the three sequence labeling methods. Experiments on two publicly available BioNLP collections show that, in addition to a WordNet feature, word embedding featured clusters (although not trained with the task‐specific corpus) consistently improve the performance for all methods on all entity types in both collections. Other features produce various results. Our results also show that when trained on limited corpora, Brown clusters resulted in better performance than word embedding clusters did. Further analysis suggests that the entity recognition performance can be greatly boosted through improving the accuracy of entity boundary identification.
      pubtype: Academic Journal
      doctype:
        equations & formulas
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N