Comparative evaluation of text classification techniques using a large diverse Arabic dataset.
A vast amount of valuable human knowledge is recorded in documents. The rapid growth in the number of machine-readable documents for public or private access necessitates the use of automatic text classification. While a lot of effort has been put into Western languages-mostly English-minimal experi...
| Published in: | Language Resources & Evaluation Vol. 47; no. 2; pp. 513 - 539 |
|---|---|
| Main Authors: | , |
| Format: | Article |
| Published: |
Springer Nature
Jun2013
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=87846092&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 87846092 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2013 vid: 47 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 87846092 10.1007/s10579-013-9221-8 ppf: 513 ppct: 26 formats: fmt: @attributes: type: P size: 413KB tig: atl: Comparative evaluation of text classification techniques using a large diverse Arabic dataset. aug: au: Khorsheed, Mohammad Al-Thubaity, Abdulmohsen affil: King Abdulaziz City for Science & Technology, Riyadh 11442 Saudi Arabia su: Machine learning Arabic language Language & languages Algorithms Decision trees Comparative studies sug: subj: Machine learning Arabic language Language & languages Algorithms Decision trees Comparative studies keyword: Arabic text categorization Arabic text classification ab: A vast amount of valuable human knowledge is recorded in documents. The rapid growth in the number of machine-readable documents for public or private access necessitates the use of automatic text classification. While a lot of effort has been put into Western languages-mostly English-minimal experimentation has been done with Arabic. This paper presents, first, an up-to-date review of the work done in the field of Arabic text classification and, second, a large and diverse dataset that can be used for benchmarking Arabic text classification algorithms. The different techniques derived from the literature review are illustrated by their application to the proposed dataset. The results of various feature selections, weighting methods, and classification algorithms show, on average, the superiority of support vector machine, followed by the decision tree algorithm (C4.5) and Naïve Bayes. The best classification accuracy was 97 % for the Islamic Topics dataset, and the least accurate was 61 % for the Arabic Poems dataset. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2013. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2013 holdings: @attributes: islocal: N |
|---|