A new hybrid stemming method for Persian language.

One of the important issues in natural language processing and information retrieval is the automatic extraction of the word 's stem. Both statistical and rule-based approaches for stemming have their own advantages and limitations. The statistical stemmers are not accurate and fail to take advantag...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 32; no. 1; pp. 209 - 222
Autores principales: Taghi-Zadeh, Hossein, Sadreddini, Mohammad Hadi, Diyanati, Mohammad Hasan, Rasekh, Amir Hossein
Formato: Artículo
Publicado: Oxford University Press / USA Apr2017
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=122282417&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 122282417
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Apr2017
      vid: 32
      iid: 1
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        122282417
        10.1093/llc/fqv053
      ppf: 209
      ppct: 13
      formats:
        fmt:
          @attributes:
            type: P
            size: 1.2MB
      tig:
        atl: A new hybrid stemming method for Persian language.
      aug:
        au:
          Taghi-Zadeh, Hossein
          Sadreddini, Mohammad Hadi
          Diyanati, Mohammad Hasan
          Rasekh, Amir Hossein
        affil: School of Computer and Electrical Engineering, Shiraz University, Iran
      su:
        Persian language
        Computational linguistics
        Information retrieval
        Combination (Linguistics)
        Orthography & spelling
      sug:
        subj:
          Persian language
          Computational linguistics
          Information retrieval
          Combination (Linguistics)
          Orthography & spelling
      ab: One of the important issues in natural language processing and information retrieval is the automatic extraction of the word 's stem. Both statistical and rule-based approaches for stemming have their own advantages and limitations. The statistical stemmers are not accurate and fail to take advantage of some language phenomenon which can be easily expressed by simple rules. On the other hand, handcrafting the stemming rules in the rule-based stemmers is a time-consuming, tedious, and impractical task. In this regard, we propose a new hybrid stemming method based on a combination of affix stripping and statistical techniques for Persian language. The proposed method combines cues from the orthography, word frequency, and syntactic distributions to induce the stemming rules. In general, the proposed method is divided into two main parts. In the first part, all words of the annotated text corpus are used to automatically induce the stemming rules; while in the second part, the rule-based stemmer uses the induced stemming rules to discover the word's stem. We test the performance of the proposed scheme on two different data sets. The encouraging results indicate the superior performance of the proposed method compared with its counterparts.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2017
    holdings:
      @attributes:
        islocal: N