Fake news article detection datasets for Hindi language.
With the increasing concerns of disinformation shared over digital platforms, detecting fake news articles in resource-poor languages is becoming an important research problem. While several datasets are under circulation in the public domain for the English language for fake news detection research...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 3; pp. 3153 - 3189 |
|---|---|
| Autores principales: | , , , , , |
| Formato: | Conference Paper/Materials |
| Publicado: |
Springer Nature
Sep2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909048&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 186909048 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2025 vid: 59 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 186909048 10.1007/s10579-024-09787-z ppf: 3153 ppct: 36 formats: fmt: – @attributes: type: T – @attributes: type: P size: 3.8MB tig: atl: Fake news article detection datasets for Hindi language. aug: au: Kumar, Sujit Shankhdhar, Anant Singal, Divyam Aggarwal, Bhuvan Malhotra, Ahaan Sameer Ranbir Singh, Sanasam affil: https://ror.org/0022nd079 Department of Computer Science and Engineering, Indian Institute of Technology Guwahati, North Guwahati, Assam, India su: Fake news Hindi language Detection algorithms Disinformation Text mining Research questions Data libraries Sampling methods sug: subj: Fake news Hindi language Detection algorithms Disinformation Text mining Research questions Data libraries Sampling methods keyword: Communication and Culture Linguistics Fake news detection Fake news Hindi dataset Language Misinformation detection in Hindi ab: With the increasing concerns of disinformation shared over digital platforms, detecting fake news articles in resource-poor languages is becoming an important research problem. While several datasets are under circulation in the public domain for the English language for fake news detection research, such datasets are not readily available for resource-poor languages languages. In this paper, we curate and propose four types of large-scale hybrid (real samples and fake synthetic samples) Hindi datasets, suitable for fake news detection research in news articles from different content and linguistic aspects for public access. Though few small-scale Hindi datasets for fake news detection are reported in the literature, they are neither readily available nor linguistically annotated. Appropriate annotation is important for developing a linguistically complex model and explainability study. The quality and reliability of the proposed datasets are further evaluated using different state-of-the-art methods over real fake news samples. pubtype: Academic Journal doctype: Conference Paper/Materials src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|