Comparative evaluation of text classification techniques using a large diverse Arabic dataset.

A vast amount of valuable human knowledge is recorded in documents. The rapid growth in the number of machine-readable documents for public or private access necessitates the use of automatic text classification. While a lot of effort has been put into Western languages-mostly English-minimal experi...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 47; no. 2; pp. 513 - 539
Main Authors: Khorsheed, Mohammad, Al-Thubaity, Abdulmohsen
Format: Article
Published: Springer Nature Jun2013
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=87846092&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 87846092
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Jun2013
      vid: 47
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        87846092
        10.1007/s10579-013-9221-8
      ppf: 513
      ppct: 26
      formats:
        fmt:
          @attributes:
            type: P
            size: 413KB
      tig:
        atl: Comparative evaluation of text classification techniques using a large diverse Arabic dataset.
      aug:
        au:
          Khorsheed, Mohammad
          Al-Thubaity, Abdulmohsen
        affil: King Abdulaziz City for Science & Technology, Riyadh 11442 Saudi Arabia
      su:
        Machine learning
        Arabic language
        Language & languages
        Algorithms
        Decision trees
        Comparative studies
      sug:
        subj:
          Machine learning
          Arabic language
          Language & languages
          Algorithms
          Decision trees
          Comparative studies
      keyword:
        Arabic text categorization
        Arabic text classification
      ab: A vast amount of valuable human knowledge is recorded in documents. The rapid growth in the number of machine-readable documents for public or private access necessitates the use of automatic text classification. While a lot of effort has been put into Western languages-mostly English-minimal experimentation has been done with Arabic. This paper presents, first, an up-to-date review of the work done in the field of Arabic text classification and, second, a large and diverse dataset that can be used for benchmarking Arabic text classification algorithms. The different techniques derived from the literature review are illustrated by their application to the proposed dataset. The results of various feature selections, weighting methods, and classification algorithms show, on average, the superiority of support vector machine, followed by the decision tree algorithm (C4.5) and Naïve Bayes. The best classification accuracy was 97 % for the Islamic Topics dataset, and the least accurate was 61 % for the Arabic Poems dataset.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2013. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2013
    holdings:
      @attributes:
        islocal: N