A new hybrid stemming method for Persian language.

One of the important issues in natural language processing and information retrieval is the automatic extraction of the word 's stem. Both statistical and rule-based approaches for stemming have their own advantages and limitations. The statistical stemmers are not accurate and fail to take advantag...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 32; no. 1; pp. 209 - 222
Autores principales: Taghi-Zadeh, Hossein, Sadreddini, Mohammad Hadi, Diyanati, Mohammad Hasan, Rasekh, Amir Hossein
Formato: Artículo
Publicado: Oxford University Press / USA Apr2017
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:One of the important issues in natural language processing and information retrieval is the automatic extraction of the word 's stem. Both statistical and rule-based approaches for stemming have their own advantages and limitations. The statistical stemmers are not accurate and fail to take advantage of some language phenomenon which can be easily expressed by simple rules. On the other hand, handcrafting the stemming rules in the rule-based stemmers is a time-consuming, tedious, and impractical task. In this regard, we propose a new hybrid stemming method based on a combination of affix stripping and statistical techniques for Persian language. The proposed method combines cues from the orthography, word frequency, and syntactic distributions to induce the stemming rules. In general, the proposed method is divided into two main parts. In the first part, all words of the annotated text corpus are used to automatically induce the stemming rules; while in the second part, the rule-based stemmer uses the induced stemming rules to discover the word's stem. We test the performance of the proposed scheme on two different data sets. The encouraging results indicate the superior performance of the proposed method compared with its counterparts.