Design and Implementation of a Machine Learning-Based Malicious URL Detection System

Authors

DOI:

https://doi.org/10.54536/ajiri.v5i3.8109

Keywords:

Cyber Security, Machine Learning, Malicious Url Detection, Phishing Detection, Random Forest, Tf-Idf, Url Classification, Xgboost

Abstract

The internet has rapidly evolved in its communication, commerce and information sharing making it a huge platform for cyber threats, particularly malicious URLs. They pose a serious threat to individuals and to organisations. Phishing attacks, malware distribution and other types of cybercrime frequently are carried out through malicious URLs. In this research, we have created and tested the machine learning models to detect malicious URLs. The labeled URLs used were obtained from a public dataset with more than 651,000 labeled URLs. The dataset was prepared for classification by applying data pre-processing techniques like stratified sampling, label encoding and Term Frequency–Inverse Document Frequency (TF-IDF) vectorization. To train and test the algorithms, five machine learning were used: Support Vector Machine (SVM), K-Nearest Neighbors (KNN), Naïve Bayes (NB), Random Forest (RF) and Extreme Gradient Boosting (XGBoost), which were trained and evaluated by the metrics of accuracy, precision, recall and F1 score. The results indicated that the Random Forest model had the highest classification accuracy (95%) as compared to the other models. Moreover, a web based malicious URL detection system was developed to demonstrate the actual application of the developed models in real time cyber security scenarios. The study provides a conclusion that the machine learning techniques, particularly ensemble learning techniques can be considered an effective and reliable technique to detect malicious URL.

Downloads

Download data is not yet available.

References

Alsad, O., & Abu Al-Haija, Q. (2024). DNS cache poisoning attack detection: A systematic review. IET Conference Proceedings, 2023(44), 426–432. https://doi.org/10.1049/icp.2024.0962

Alzahrani, A., Alenazi, M., Alshammari, S., & Alharthi, R. (2020). Email spoofing attack detection through an end-to-end authorship attribution system. In Proceedings of the 6th International Conference on Information Systems Security and Privacy (ICISSP 2020) (pp. 64–74). SCITEPRESS. https://doi.org/10.5220/0008954600640074

Aryee, B. A., Umoren, J., & Agymang, K. A. (2025). Artificial intelligence for strengthening cybersecurity in U.S. healthcare systems. American Journal of Medical Science and Innovation, 4(2), 150–158. https://doi.org/10.54536/ajmsi.v4i2.6178

Aslan, Ö., Aktuğ, S. S., Ozkan-Okay, M., Yilmaz, A. A., & Akin, E. (2023). A comprehensive review of cyber security vulnerabilities, threats, attacks, and solutions. Electronics, 12(6), 1333. https://doi.org/10.3390/electronics12061333

Catak, F. O., Şahinbaş, K., & Dörtkardeş, V. (2021). Malicious URL detection using machine learning. In A. K. Luhach & A. Elçi (Eds.), Artificial intelligence paradigms for smart cyber-physical systems (pp. 160–180). IGI Global. https://doi.org/10.4018/978-1-7998-5101-1.ch008

Cisco. (2022). Cisco event response: Corporate network security incident. Cisco Security Center. https://sec.cloudapps.cisco.com/security/center/resources/corp_network_security_incident

Dawood, M., & Ibrahim, O. (2020). Exploration of hidden fraudulent website dataset with different perspective. Journal of Computational and Theoretical Nanoscience, 17(2–3), 980–984. https://doi.org/10.1166/jctn.2020.8738

Geldenhuys, K. (2024). Spoofing unmasked: Cheating criminals shatter trust. Servamus Community-based Safety and Security Magazine, 117(10). https://hdl.handle.net/10520/ejc-servamus_v117_n10_a6

Hasan, M. K. (2024). New heuristics method for malicious URLs detection using machine learning. Wasit Journal of Computer and Mathematics Science, 3(3), 60–67. https://doi.org/10.31185/wjcms.267

Ijiga, O. M., Idoko, I. P., Ebiega, G. I., Olajide, F. I., Olatunde, T. I., & Ukaegbu, C. (2024). Harnessing adversarial machine learning for advanced threat detection: AI-driven strategies in cybersecurity risk assessment and fraud prevention. Open Access Research Journal of Science and Technology, 11(1), 1–24. https://doi.org/10.53022/oarjst.2024.11.1.0060

Imani, M., Beikmohammadi, A., & Arabnia, H. R. (2025). Comprehensive analysis of Random Forest and XGBoost performance with SMOTE, ADASYN, and GNUS under varying imbalance levels. Technologies, 13(3), 88. https://doi.org/10.3390/technologies13030088

Kailas, S., & Roopalakshmi, R. (2025). ‘Think before you click’ - Malicious URL detection in cybersecurity: A systematic review and research roadmap. IEEE Access, 13, 154305–154325. https://doi.org/10.1109/ACCESS.2025.3601387

Li, M., Jiang, Y., Zhang, Y., & Zhu, H. (2023). Medical image analysis using deep learning algorithms. Frontiers in Public Health, 11, 1273253. https://doi.org/10.3389/fpubh.2023.1273253

Li, W., Manickam, S., Chong, Y.-W., Leng, W., & Nanda, P. (2024). A state-of-the-art review on phishing website detection techniques. IEEE Access, 12, 187976–188012. https://doi.org/10.1109/ACCESS.2024.3514972

Mehndiratta, M., Jain, N., Malhotra, A., Gupta, I., & Narula, R. (2023). Malicious URL: Analysis and detection using machine learning.

Mosa, D. T., Shams, M. Y., Abohany, A. A., El-Kenawy, E. M., & Thabet, M. (2023). Machine learning techniques for detecting phishing URL attacks. Computers, Materials & Continua, 75(1), 1271–1290. https://doi.org/10.32604/cmc.2023.036422

Ogunbiyi, S. A., Adeleke, O., & Osunade, O. (2026). An ensemble machine learning framework for automated cybersecurity incident classification using structured metadata. SADI International Journal of Science, Engineering and Technology, 13(2), 1–11. https://doi.org/10.5281/zenodo.20025311https://doi.org/10.5281/zenodo.20025311

Olatunji, B. T., Adeleke, O., Osunade, O., & Asoro, B. (2026). Intelligent classification of healthcare incident reports using supervised learning. SADI International Journal of Science, Engineering and Technology (SIJSET), 13(2), 25–36. https://doi.org/10.5281/zenodo.20542825

Osmanoğlu, M., Gupta, D., & Sharma, V. (2026). A comprehensive review of malicious URLs: Detection techniques, features and datasets. Computers & Electrical Engineering, 136, 111186. https://doi.org/10.1016/j.compeleceng.2026.111186

Zhou, J., Ye, Z., Zhang, S., Geng, Z., Han, N., & Yang, T. (2024). Investigating response behavior through TF-IDF and Word2vec text analysis: A case study of PISA 2012 problem-solving process data. Heliyon, 10(16), e35945. https://doi.org/10.1016/j.heliyon.2024.e35945

Downloads

Published

2026-09-01

How to Cite

Adeleke, O. ., Kuranga, J. A. ., Ogunbiyi, S. A. ., & Aneke, E. .O. . (2026). Design and Implementation of a Machine Learning-Based Malicious URL Detection System. American Journal of Interdisciplinary Research and Innovation, 5(3), 91-100. https://doi.org/10.54536/ajiri.v5i3.8109

Similar Articles

11-20 of 49

You may also start an advanced similarity search for this article.