| Sumario: | With the increasing concerns of disinformation shared over digital platforms, detecting fake news articles in resource-poor languages is becoming an important research problem. While several datasets are under circulation in the public domain for the English language for fake news detection research, such datasets are not readily available for resource-poor languages languages. In this paper, we curate and propose four types of large-scale hybrid (real samples and fake synthetic samples) Hindi datasets, suitable for fake news detection research in news articles from different content and linguistic aspects for public access. Though few small-scale Hindi datasets for fake news detection are reported in the literature, they are neither readily available nor linguistically annotated. Appropriate annotation is important for developing a linguistically complex model and explainability study. The quality and reliability of the proposed datasets are further evaluated using different state-of-the-art methods over real fake news samples.
|