Improving Classification of Protein Interaction Articles Using Context Similarity-Based Feature Selection.

Protein interaction article classification is a text classification task in the biological domain to determine which articles describe protein-protein interactions. Since the feature space in text classification is high-dimensional, feature selection is widely used for reducing the dimensionality of...

Descripción completa

Detalles Bibliográficos
Publicado en:BioMed Research International Vol. 2015; pp. 1 - 11
Autores principales: Chen, Yifei, Sun, Yuxing, Han, Bing-Qing
Formato: pictorial research tables/charts Journal Article
Publicado: Wiley-Blackwell 8/3/2015
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=109031012&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 109031012
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        23146133
        FT2T
      jtl: BioMed Research International
      issn: 23146133
      maglogo: N
    pubinfo:
      dt: 8/3/2015
      vid: 2015
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        109031012
        109031012
        109031012
        10.1155/2015/751646
        109031012
      ppf: 1
      ppct: 10
      formats:
        fmt:
          @attributes:
            type: P
      tig:
        atl: Improving Classification of Protein Interaction Articles Using Context Similarity-Based Feature Selection.
      aug:
        au:
          Chen, Yifei
          Sun, Yuxing
          Han, Bing-Qing
        affil: School of Technology, Nanjing Audit University, 86 W. Yushan Road, Nanjing 211815, China
      sug:
        subj:
          Proteins Physiology
          Writing for Publication Classification
          Data Mining Methods
          Models, Statistical
      ab: Protein interaction article classification is a text classification task in the biological domain to determine which articles describe protein-protein interactions. Since the feature space in text classification is high-dimensional, feature selection is widely used for reducing the dimensionality of features to speed up computation without sacrificing classification performance. Many existing feature selection methods are based on the statistical measure of document frequency and term frequency. One potential drawback of these methods is that they treat features separately. Hence, first we design a similarity measure between the context information to take word cooccurrences and phrase chunks around the features into account. Then we introduce the similarity of context information to the importance measure of the features to substitute the document and term frequency. Hence we propose new context similarity-based feature selection methods. Their performance is evaluated on two protein interaction article collections and compared against the frequency-based methods. The experimental results reveal that the context similarity-based methods perform better in terms of the F1 measure and the dimension reduction rate. Benefiting from the context information surrounding the features, the proposed methods can select distinctive features effectively for protein interaction article classification.
      pubtype: Academic Journal
      doctype:
        pictorial
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N