Comparing multiple categories of feature selection methods for text classification.

Selecting effective features from data sets is a particularly important part in text classification, data mining, pattern recognition, and artificial intelligence. Feature selection (FS) is capable of excluding irrelevant features for the classification task and reducing the dimensionality of data s...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 35; no. 1; pp. 208 - 225
Autores principales: Zheng, Wanwan, Jin, Mingzhe
Formato: Artículo
Publicado: Oxford University Press / USA Apr2020
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=142636788&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 142636788
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Apr2020
      vid: 35
      iid: 1
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        142636788
        10.1093/llc/fqz003
      ppf: 208
      ppct: 17
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 517KB
      tig:
        atl: Comparing multiple categories of feature selection methods for text classification.
      aug:
        au:
          Zheng, Wanwan
          Jin, Mingzhe
        affil: Graduate School of Culture and Information Science, Doshisha University, Japan
      su:
        Feature selection
        Pattern recognition systems
        Support vector machines
        Data mining
        Text processing (Computer science)
        Machine performance
      sug:
        subj:
          Feature selection
          Pattern recognition systems
          Support vector machines
          Data mining
          Text processing (Computer science)
          Machine performance
      ab: Selecting effective features from data sets is a particularly important part in text classification, data mining, pattern recognition, and artificial intelligence. Feature selection (FS) is capable of excluding irrelevant features for the classification task and reducing the dimensionality of data sets, which help us better understand data. Through FS selection, the performance of machine learning techniques is improved, and computation requirement is minimized. Thus far, a large number of FS methods have been proposed, whereas the most practically effective one has not been found. Although it is conceivable that different categories of FS methods follow different criteria for evaluating variables, rare studies have focused on evaluating various categories of FS methods. This article first lists thirteen superior FS methods under five different categories and focuses on evaluating and comparing the effectiveness and general versatility of these methods. The thirteen FS methods were ranked using rank aggregation method. Subsequently, the best five FS methods were elected to perform multi-class classifications. Support vector machine served as the classifier. Different languages, different numbers of selected features, and different performance measures were employed to measure the effectiveness and general versatility of these methods together. The analysis results suggest that Mahalanobis distance is the best method on the whole.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2020
    holdings:
      @attributes:
        islocal: N