EVALUATING PHISHING DETECTION SYSTEMS AGAINST OBFUSCATED URLS: A BENCHMARK AND ANALYSIS.
Cyber security is an arms race between the bad guys and phishing protection systems, as the former continue to seek ways to obfuscate URLs. Machine learning techniques are highly effective in detecting phishing URLs, but it remains to be seen how they will hold up under attack. In this paper we repo...
| Published in: | Scientific Culture Vol. 12; no. 5 Part 1; pp. 350 - 360 |
|---|---|
| Main Author: | |
| Format: | Article |
| Published: |
University of the Aegean
2026
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=193975199&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 193975199 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 24080071 I6HU jtl: Scientific Culture issn: 24080071 maglogo: N pubinfo: dt: 2026 vid: 12 iid: 5 Part 1 pid: 47715 pub: University of the Aegean artinfo: ui: 193975199 10.5281/zenodo.12511030 ppf: 350 ppct: 10 formats: tig: atl: EVALUATING PHISHING DETECTION SYSTEMS AGAINST OBFUSCATED URLS: A BENCHMARK AND ANALYSIS. aug: au: Punjabi, Sunil Kumar affil: Research Scholar, Shridhar University, Pilani Assistant Professor, Dept of Computer Engineering, SIES Graduate School of Technology, Nerul, Navi Mumbai su: Phishing Machine learning Random forest algorithms Internet security Logistic regression analysis Boosting algorithms Multilayer perceptrons sug: subj: Phishing Machine learning Random forest algorithms Internet security Logistic regression analysis Boosting algorithms Multilayer perceptrons keyword: adversarial robustness character-level features cybersecurity machine learning phishing detection URL obfuscation URL semantics ab: Cyber security is an arms race between the bad guys and phishing protection systems, as the former continue to seek ways to obfuscate URLs. Machine learning techniques are highly effective in detecting phishing URLs, but it remains to be seen how they will hold up under attack. In this paper we report baseline testing of phishing protection systems for URL obfuscation. We developed a quantitative set of experiments using a list of URL features from the PhiUSIIL data set. We tested four machine learning models (Logistic Regression, Random Forest, Extreme Gradient Boosting (XGBoost) and Multi-Layer Perceptron (MLP) using baseline test data (a mix of good and bad URLs) and synthetic test data (a mix of good and bad URLs with medium obfuscation, including URLs with subdomain injection, character replacement and encoding). We found all models performed extremely well on baseline (better than macro F1 score of 0.99), but less well on obfuscated. We found Random Forest and XGBoost models were more robust, attaining more than 40% of original performance, while Logistic Regression and MLP were less robust (less than 20%). This is a trade-off between performance and robustness. This work shows the importance of robust models for phishing detection systems and provides a method for evaluating model robustness to inform the development of robust models for detecting phishing under attacks. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2026 holdings: @attributes: islocal: N |
|---|