Design and Implementation of a Machine Learning-Based Malicious URL Detection System
DOI:
https://doi.org/10.54536/ajiri.v5i3.8109Keywords:
Cyber Security, Machine Learning, Malicious Url Detection, Phishing Detection, Random Forest, Tf-Idf, Url Classification, XgboostAbstract
The internet has rapidly evolved in its communication, commerce and information sharing making it a huge platform for cyber threats, particularly malicious URLs. They pose a serious threat to individuals and to organisations. Phishing attacks, malware distribution and other types of cybercrime frequently are carried out through malicious URLs. In this research, we have created and tested the machine learning models to detect malicious URLs. The labeled URLs used were obtained from a public dataset with more than 651,000 labeled URLs. The dataset was prepared for classification by applying data pre-processing techniques like stratified sampling, label encoding and Term Frequency–Inverse Document Frequency (TF-IDF) vectorization. To train and test the algorithms, five machine learning were used: Support Vector Machine (SVM), K-Nearest Neighbors (KNN), Naïve Bayes (NB), Random Forest (RF) and Extreme Gradient Boosting (XGBoost), which were trained and evaluated by the metrics of accuracy, precision, recall and F1 score. The results indicated that the Random Forest model had the highest classification accuracy (95%) as compared to the other models. Moreover, a web based malicious URL detection system was developed to demonstrate the actual application of the developed models in real time cyber security scenarios. The study provides a conclusion that the machine learning techniques, particularly ensemble learning techniques can be considered an effective and reliable technique to detect malicious URL.
Downloads
References
Alsad, O., & Abu Al-Haija, Q. (2024). DNS cache poisoning attack detection: A systematic review. IET Conference Proceedings, 2023(44), 426–432. https://doi.org/10.1049/icp.2024.0962
Alzahrani, A., Alenazi, M., Alshammari, S., & Alharthi, R. (2020). Email spoofing attack detection through an end-to-end authorship attribution system. In Proceedings of the 6th International Conference on Information Systems Security and Privacy (ICISSP 2020) (pp. 64–74). SCITEPRESS. https://doi.org/10.5220/0008954600640074
Aryee, B. A., Umoren, J., & Agymang, K. A. (2025). Artificial intelligence for strengthening cybersecurity in U.S. healthcare systems. American Journal of Medical Science and Innovation, 4(2), 150–158. https://doi.org/10.54536/ajmsi.v4i2.6178
Aslan, Ö., Aktuğ, S. S., Ozkan-Okay, M., Yilmaz, A. A., & Akin, E. (2023). A comprehensive review of cyber security vulnerabilities, threats, attacks, and solutions. Electronics, 12(6), 1333. https://doi.org/10.3390/electronics12061333
Catak, F. O., Şahinbaş, K., & Dörtkardeş, V. (2021). Malicious URL detection using machine learning. In A. K. Luhach & A. Elçi (Eds.), Artificial intelligence paradigms for smart cyber-physical systems (pp. 160–180). IGI Global. https://doi.org/10.4018/978-1-7998-5101-1.ch008
Cisco. (2022). Cisco event response: Corporate network security incident. Cisco Security Center. https://sec.cloudapps.cisco.com/security/center/resources/corp_network_security_incident
Dawood, M., & Ibrahim, O. (2020). Exploration of hidden fraudulent website dataset with different perspective. Journal of Computational and Theoretical Nanoscience, 17(2–3), 980–984. https://doi.org/10.1166/jctn.2020.8738
Geldenhuys, K. (2024). Spoofing unmasked: Cheating criminals shatter trust. Servamus Community-based Safety and Security Magazine, 117(10). https://hdl.handle.net/10520/ejc-servamus_v117_n10_a6
Hasan, M. K. (2024). New heuristics method for malicious URLs detection using machine learning. Wasit Journal of Computer and Mathematics Science, 3(3), 60–67. https://doi.org/10.31185/wjcms.267
Ijiga, O. M., Idoko, I. P., Ebiega, G. I., Olajide, F. I., Olatunde, T. I., & Ukaegbu, C. (2024). Harnessing adversarial machine learning for advanced threat detection: AI-driven strategies in cybersecurity risk assessment and fraud prevention. Open Access Research Journal of Science and Technology, 11(1), 1–24. https://doi.org/10.53022/oarjst.2024.11.1.0060
Imani, M., Beikmohammadi, A., & Arabnia, H. R. (2025). Comprehensive analysis of Random Forest and XGBoost performance with SMOTE, ADASYN, and GNUS under varying imbalance levels. Technologies, 13(3), 88. https://doi.org/10.3390/technologies13030088
Kailas, S., & Roopalakshmi, R. (2025). ‘Think before you click’ - Malicious URL detection in cybersecurity: A systematic review and research roadmap. IEEE Access, 13, 154305–154325. https://doi.org/10.1109/ACCESS.2025.3601387
Li, M., Jiang, Y., Zhang, Y., & Zhu, H. (2023). Medical image analysis using deep learning algorithms. Frontiers in Public Health, 11, 1273253. https://doi.org/10.3389/fpubh.2023.1273253
Li, W., Manickam, S., Chong, Y.-W., Leng, W., & Nanda, P. (2024). A state-of-the-art review on phishing website detection techniques. IEEE Access, 12, 187976–188012. https://doi.org/10.1109/ACCESS.2024.3514972
Mehndiratta, M., Jain, N., Malhotra, A., Gupta, I., & Narula, R. (2023). Malicious URL: Analysis and detection using machine learning.
Mosa, D. T., Shams, M. Y., Abohany, A. A., El-Kenawy, E. M., & Thabet, M. (2023). Machine learning techniques for detecting phishing URL attacks. Computers, Materials & Continua, 75(1), 1271–1290. https://doi.org/10.32604/cmc.2023.036422
Ogunbiyi, S. A., Adeleke, O., & Osunade, O. (2026). An ensemble machine learning framework for automated cybersecurity incident classification using structured metadata. SADI International Journal of Science, Engineering and Technology, 13(2), 1–11. https://doi.org/10.5281/zenodo.20025311https://doi.org/10.5281/zenodo.20025311
Olatunji, B. T., Adeleke, O., Osunade, O., & Asoro, B. (2026). Intelligent classification of healthcare incident reports using supervised learning. SADI International Journal of Science, Engineering and Technology (SIJSET), 13(2), 25–36. https://doi.org/10.5281/zenodo.20542825
Osmanoğlu, M., Gupta, D., & Sharma, V. (2026). A comprehensive review of malicious URLs: Detection techniques, features and datasets. Computers & Electrical Engineering, 136, 111186. https://doi.org/10.1016/j.compeleceng.2026.111186
Zhou, J., Ye, Z., Zhang, S., Geng, Z., Han, N., & Yang, T. (2024). Investigating response behavior through TF-IDF and Word2vec text analysis: A case study of PISA 2012 problem-solving process data. Heliyon, 10(16), e35945. https://doi.org/10.1016/j.heliyon.2024.e35945
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Oludele Adeleke, Jimoh Abdulhakeem Kuranga, Samuel Adeolu Ogunbiyi, Endurance .O. Aneke

This work is licensed under a Creative Commons Attribution 4.0 International License.



