Generation, implementation, and appraisal of an N-gram-based stemming algorithm.

A language-independent stemmer has always been looked for. Single N -gram tokenization technique works well; however, it often generates stems that start with intermediate characters, rather than initial ones. We present a novel technique that takes the concept of N -gram stemming one step ahead and...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 34; no. 3; pp. 558 - 569
Autores principales: Pande, Bhagwati P, Tamta, Pawan, Dhami, Hoshiyar S
Formato: Artículo
Publicado: Oxford University Press / USA Sep2019
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=138342368&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 138342368
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Sep2019
      vid: 34
      iid: 3
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        138342368
        10.1093/llc/fqy053
      ppf: 558
      ppct: 11
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 193KB
      tig:
        atl: Generation, implementation, and appraisal of an N-gram-based stemming algorithm.
      aug:
        au:
          Pande, Bhagwati P
          Tamta, Pawan
          Dhami, Hoshiyar S
        affil:
          Department of Computer Science, SSJ Campus, Kumaun University, Almora, Uttarakhand, India
          Government Post Graduate College, Manila, Almora, Uttarakhand, India
          Uttarakhand Residential University, Almora, Uttarakhand, India
      su:
        Portuguese language
        Algorithms
        Reproduction
      sug:
        subj:
          Portuguese language
          Algorithms
          Reproduction
      ab: A language-independent stemmer has always been looked for. Single N -gram tokenization technique works well; however, it often generates stems that start with intermediate characters, rather than initial ones. We present a novel technique that takes the concept of N -gram stemming one step ahead and compare our method with an established algorithm in the field, say, Porter's stemmer for English, Spanish, and Portuguese languages. Results indicate that our N -gram stemmer is comparable with the Porter's linguistic stemmer.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2019
    holdings:
      @attributes:
        islocal: N