A new hybrid stemming method for Persian language.
One of the important issues in natural language processing and information retrieval is the automatic extraction of the word 's stem. Both statistical and rule-based approaches for stemming have their own advantages and limitations. The statistical stemmers are not accurate and fail to take advantag...
| Publicado en: | Digital Scholarship in the Humanities Vol. 32; no. 1; pp. 209 - 222 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Apr2017
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=122282417&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 122282417 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Apr2017 vid: 32 iid: 1 pid: 622 pub: Oxford University Press / USA artinfo: ui: 122282417 10.1093/llc/fqv053 ppf: 209 ppct: 13 formats: fmt: @attributes: type: P size: 1.2MB tig: atl: A new hybrid stemming method for Persian language. aug: au: Taghi-Zadeh, Hossein Sadreddini, Mohammad Hadi Diyanati, Mohammad Hasan Rasekh, Amir Hossein affil: School of Computer and Electrical Engineering, Shiraz University, Iran su: Persian language Computational linguistics Information retrieval Combination (Linguistics) Orthography & spelling sug: subj: Persian language Computational linguistics Information retrieval Combination (Linguistics) Orthography & spelling ab: One of the important issues in natural language processing and information retrieval is the automatic extraction of the word 's stem. Both statistical and rule-based approaches for stemming have their own advantages and limitations. The statistical stemmers are not accurate and fail to take advantage of some language phenomenon which can be easily expressed by simple rules. On the other hand, handcrafting the stemming rules in the rule-based stemmers is a time-consuming, tedious, and impractical task. In this regard, we propose a new hybrid stemming method based on a combination of affix stripping and statistical techniques for Persian language. The proposed method combines cues from the orthography, word frequency, and syntactic distributions to induce the stemming rules. In general, the proposed method is divided into two main parts. In the first part, all words of the annotated text corpus are used to automatically induce the stemming rules; while in the second part, the rule-based stemmer uses the induced stemming rules to discover the word's stem. We test the performance of the proposed scheme on two different data sets. The encouraging results indicate the superior performance of the proposed method compared with its counterparts. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2017 holdings: @attributes: islocal: N |
|---|