Fake news article detection datasets for Hindi language.

With the increasing concerns of disinformation shared over digital platforms, detecting fake news articles in resource-poor languages is becoming an important research problem. While several datasets are under circulation in the public domain for the English language for fake news detection research...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 3; pp. 3153 - 3189
Autores principales: Kumar, Sujit, Shankhdhar, Anant, Singal, Divyam, Aggarwal, Bhuvan, Malhotra, Ahaan Sameer, Ranbir Singh, Sanasam
Formato: Conference Paper/Materials
Publicado: Springer Nature Sep2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909048&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186909048
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2025
      vid: 59
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        186909048
        10.1007/s10579-024-09787-z
      ppf: 3153
      ppct: 36
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 3.8MB
      tig:
        atl: Fake news article detection datasets for Hindi language.
      aug:
        au:
          Kumar, Sujit
          Shankhdhar, Anant
          Singal, Divyam
          Aggarwal, Bhuvan
          Malhotra, Ahaan Sameer
          Ranbir Singh, Sanasam
        affil: https://ror.org/0022nd079 Department of Computer Science and Engineering, Indian Institute of Technology Guwahati, North Guwahati, Assam, India
      su:
        Fake news
        Hindi language
        Detection algorithms
        Disinformation
        Text mining
        Research questions
        Data libraries
        Sampling methods
      sug:
        subj:
          Fake news
          Hindi language
          Detection algorithms
          Disinformation
          Text mining
          Research questions
          Data libraries
          Sampling methods
      keyword:
        Communication and Culture Linguistics
        Fake news detection
        Fake news Hindi dataset
        Language
        Misinformation detection in Hindi
      ab: With the increasing concerns of disinformation shared over digital platforms, detecting fake news articles in resource-poor languages is becoming an important research problem. While several datasets are under circulation in the public domain for the English language for fake news detection research, such datasets are not readily available for resource-poor languages languages. In this paper, we curate and propose four types of large-scale hybrid (real samples and fake synthetic samples) Hindi datasets, suitable for fake news detection research in news articles from different content and linguistic aspects for public access. Though few small-scale Hindi datasets for fake news detection are reported in the literature, they are neither readily available nor linguistically annotated. Appropriate annotation is important for developing a linguistically complex model and explainability study. The quality and reliability of the proposed datasets are further evaluated using different state-of-the-art methods over real fake news samples.
      pubtype: Academic Journal
      doctype: Conference Paper/Materials
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N