ChemTok: A New Rule Based Tokenizer for Chemical Named Entity Recognition.

Named Entity Recognition (NER) from text constitutes the first step in many text mining applications. The most important preliminary step for NER systems using machine learning approaches is tokenization where raw text is segmented into tokens. This study proposes an enhanced rule based tokenizer, C...

Descripción completa

Detalles Bibliográficos
Publicado en:BioMed Research International Vol. 2016; pp. 1 - 10
Autores principales: Akkasi, Abbas, Varoğlu, Ekrem, Dimililer, Nazife
Formato: algorithm research tables/charts Journal Article
Publicado: Wiley-Blackwell 1/28/2016
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=113630382&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 113630382
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        23146133
        FT2T
      jtl: BioMed Research International
      issn: 23146133
      maglogo: N
    pubinfo:
      dt: 1/28/2016
      vid: 2016
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        113630382
        113630382
        113630382
        10.1155/2016/4248026
        113630382
      ppf: 1
      ppct: 9
      formats:
        fmt:
          @attributes:
            type: P
      tig:
        atl: ChemTok: A New Rule Based Tokenizer for Chemical Named Entity Recognition.
      aug:
        au:
          Akkasi, Abbas
          Varoğlu, Ekrem
          Dimililer, Nazife
        affil: Computer Engineering Department, Eastern Mediterranean University, Famagusta, Northern Cyprus, Mersin 10, Turkey
      sug:
        subj:
          Artificial Intelligence
          Data Mining
          Algorithms
          Software
          Nomenclature
      ab: Named Entity Recognition (NER) from text constitutes the first step in many text mining applications. The most important preliminary step for NER systems using machine learning approaches is tokenization where raw text is segmented into tokens. This study proposes an enhanced rule based tokenizer, ChemTok, which utilizes rules extracted mainly from the train data set. The main novelty of ChemTok is the use of the extracted rules in order to merge the tokens split in the previous steps, thus producing longer and more discriminative tokens. ChemTok is compared to the tokenization methods utilized by ChemSpot and tmChem. Support Vector Machines and Conditional Random Fields are employed as the learning algorithms. The experimental results show that the classifiers trained on the output of ChemTok outperforms all classifiers trained on the output of the other two tokenizers in terms of classification performance, and the number of incorrectly segmented entities.
      pubtype: Academic Journal
      doctype:
        algorithm
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N