Linguistic features evaluation for hadith authenticity through automatic machine learning.

There has not been any research that provides an evaluation of the linguistic features extracted from the matn (text) of a Hadith. Moreover, none of the fairly large corpora are publicly available as a benchmark corpus for Hadith authenticity, and there is a need to build a 'gold standard' corpus fo...

Full description

Bibliographic Details
Published in:Digital Scholarship in the Humanities Vol. 37; no. 3; pp. 830 - 844
Main Authors: Mohamed, Emad, Sarwar, Raheem
Format: Article
Published: Oxford University Press / USA Sep2022
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=158667350&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 158667350
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Sep2022
      vid: 37
      iid: 3
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        158667350
        10.1093/llc/fqab092
      ppf: 830
      ppct: 14
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.5MB
      tig:
        atl: Linguistic features evaluation for hadith authenticity through automatic machine learning.
      aug:
        au:
          Mohamed, Emad
          Sarwar, Raheem
        affil: Research Group in Computational Linguistics, University of Wolverhampton , UK
      su:
        Naive Bayes classification
        Machine learning
        Hadith
        Python programming language
        Digital humanities
      sug:
        subj:
          Naive Bayes classification
          Machine learning
          Hadith
          Python programming language
          Digital humanities
      ab: There has not been any research that provides an evaluation of the linguistic features extracted from the matn (text) of a Hadith. Moreover, none of the fairly large corpora are publicly available as a benchmark corpus for Hadith authenticity, and there is a need to build a 'gold standard' corpus for good practices in Hadith authentication. We write a scraper in Python programming language and collect a corpus of 3,651 authentic prophetic traditions and 3,593 fake ones. We process the corpora with morphological segmentation and perform extensive experimental studies using a variety of machine learning algorithms, mainly through automatic machine learning, to distinguish between these two categories. With a feature set including words, morphological segments, characters, top N words, top N segments, function words, and several vocabulary richness features, we analyze the results in terms of both prediction and interpretability to explain which features are more characteristic of each class. Many experiments have produced good results and the highest accuracy (i.e. 78.28%) is achieved using word n-grams as features using the Multinomial Naive Bayes classifier. Our extensive experimental studies conclude that, at least for Digital Humanities, feature engineering may still be desirable due to the high interpretability of the features. The corpus and software (scripts) will be made publicly available to other researchers in an effort to promote progress and replicability.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2022
    holdings:
      @attributes:
        islocal: N