Generation, implementation, and appraisal of an N-gram-based stemming algorithm.

A language-independent stemmer has always been looked for. Single N -gram tokenization technique works well; however, it often generates stems that start with intermediate characters, rather than initial ones. We present a novel technique that takes the concept of N -gram stemming one step ahead and...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 34; no. 3; pp. 558 - 569
Autores principales: Pande, Bhagwati P, Tamta, Pawan, Dhami, Hoshiyar S
Formato: Artículo
Publicado: Oxford University Press / USA Sep2019
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:A language-independent stemmer has always been looked for. Single N -gram tokenization technique works well; however, it often generates stems that start with intermediate characters, rather than initial ones. We present a novel technique that takes the concept of N -gram stemming one step ahead and compare our method with an established algorithm in the field, say, Porter's stemmer for English, Spanish, and Portuguese languages. Results indicate that our N -gram stemmer is comparable with the Porter's linguistic stemmer.