Comparing multiple categories of feature selection methods for text classification.
Selecting effective features from data sets is a particularly important part in text classification, data mining, pattern recognition, and artificial intelligence. Feature selection (FS) is capable of excluding irrelevant features for the classification task and reducing the dimensionality of data s...
| Publicado en: | Digital Scholarship in the Humanities Vol. 35; no. 1; pp. 208 - 225 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Apr2020
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=142636788&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 142636788 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Apr2020 vid: 35 iid: 1 pid: 622 pub: Oxford University Press / USA artinfo: ui: 142636788 10.1093/llc/fqz003 ppf: 208 ppct: 17 formats: fmt: – @attributes: type: T – @attributes: type: P size: 517KB tig: atl: Comparing multiple categories of feature selection methods for text classification. aug: au: Zheng, Wanwan Jin, Mingzhe affil: Graduate School of Culture and Information Science, Doshisha University, Japan su: Feature selection Pattern recognition systems Support vector machines Data mining Text processing (Computer science) Machine performance sug: subj: Feature selection Pattern recognition systems Support vector machines Data mining Text processing (Computer science) Machine performance ab: Selecting effective features from data sets is a particularly important part in text classification, data mining, pattern recognition, and artificial intelligence. Feature selection (FS) is capable of excluding irrelevant features for the classification task and reducing the dimensionality of data sets, which help us better understand data. Through FS selection, the performance of machine learning techniques is improved, and computation requirement is minimized. Thus far, a large number of FS methods have been proposed, whereas the most practically effective one has not been found. Although it is conceivable that different categories of FS methods follow different criteria for evaluating variables, rare studies have focused on evaluating various categories of FS methods. This article first lists thirteen superior FS methods under five different categories and focuses on evaluating and comparing the effectiveness and general versatility of these methods. The thirteen FS methods were ranked using rank aggregation method. Subsequently, the best five FS methods were elected to perform multi-class classifications. Support vector machine served as the classifier. Different languages, different numbers of selected features, and different performance measures were employed to measure the effectiveness and general versatility of these methods together. The analysis results suggest that Mahalanobis distance is the best method on the whole. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2020 holdings: @attributes: islocal: N |
|---|