Evaluation of Keyword Extraction using YAKE and KeyBERT in Text Preprocessing for Hoax News Detection Based on Bi-LSTM.
The spread of hoaxes through social media presents a significant challenge to the accuracy of public information. Automated detection based on natural language processing (NLP) offers a potential solution to this issue. This study investigates the impact of keyword extraction methods on the performa...
| Publicado en: | Riwayat: Educational Journal of History & Humanities Vol. 8; no. 3; pp. 4490 - 4501 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Riwayat: Educational Journal of History & Humanities
2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=188746904&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 188746904 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 26143917 NA6W jtl: Riwayat: Educational Journal of History & Humanities issn: 26143917 maglogo: N pubinfo: dt: 2025 vid: 8 iid: 3 pid: 76246 pub: Riwayat: Educational Journal of History & Humanities artinfo: ui: 188746904 10.24815/jr.v8i3.48626 ppf: 4490 ppct: 11 formats: tig: atl: Evaluation of Keyword Extraction using YAKE and KeyBERT in Text Preprocessing for Hoax News Detection Based on Bi-LSTM. aug: au: Kamila, Ahya Radiatul Derhass, Gerry Hudera Surianto Budiyanto, Very affil: Fakultas Teknologi dan Desain, Universitas Bunda Mulia Fakultas, Sekolah Sains Data, Matematika, dan Informatika, Institut Pertanian Bogor Fakultas Ilmu Sosial dan Humaniora, Universitas Bunda Mulia su: Long short-term memory Feature extraction Deception Feature selection Computational linguistics Statistical accuracy sug: subj: Long short-term memory Feature extraction Deception Feature selection Computational linguistics Statistical accuracy keyword: Bi-LSTM Hoax KeyBERT YAKE ab: The spread of hoaxes through social media presents a significant challenge to the accuracy of public information. Automated detection based on natural language processing (NLP) offers a potential solution to this issue. This study investigates the impact of keyword extraction methods on the performance of hoax classification using the Bidirectional Long Short-Term Memory (Bi-LSTM) architecture. Two methods are evaluated: YAKE, which relies on statistical features, and KeyBERT, which utilizes semantic representations from the BERT transformer model. The IDNHoaxCorpus, an Indonesian-language dataset, serves as the experimental basis, undergoing preprocessing, keyword extraction, and model training stages. Evaluation metrics include accuracy, precision, recall, F1-score, and processing time. Results show that KeyBERT achieves higher accuracy and F1-score (82.56% and 73.30%, respectively) compared to YAKE (80.07% and 71.11%), but at the cost of significantly longer processing time (360 seconds vs. 13 seconds). These findings highlight a notable trade-off between accuracy and computational efficiency, which should be considered based on application requirements such as real-time systems or batch processing. This study underscores the importance of selecting appropriate feature extraction strategies in text-based hoax detection systems. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|