A comparison of text-classification techniques applied to Arabic text.

Many algorithms have been implemented for the problem of text classification. Most of the work in this area was carried out for English text. Very little research has been carried out on Arabic text. The nature of Arabic text is different than that of English text, and preprocessing of Arabic text i...

Full description

Bibliographic Details
Published in:Journal of the American Society for Information Science & Technology Vol. 60; no. 9; pp. 1836 - 1845
Main Authors: Kanaan G, Al-Shalabi R, Ghwanmeh S, Al-Ma'adeed H
Format: algorithm equations & formulas research tables/charts Journal Article
Published: Wiley-Blackwell Sep2009
Online Access:View this record in EBSCOhost
Description
Summary:Many algorithms have been implemented for the problem of text classification. Most of the work in this area was carried out for English text. Very little research has been carried out on Arabic text. The nature of Arabic text is different than that of English text, and preprocessing of Arabic text is more challenging. This paper presents an implementation of three automatic text-classification techniques for Arabic text. A corpus of 1445 Arabic text documents belonging to nine categories has been automatically classified using the kNN, Rocchio, and naïve Bayes algorithms. The research results reveal that Naïve Bayes was the best performer, followed by kNN and Rocchio.