A hybrid structural-contextual framework for malicious URL detection based on graph attention networks and knowledge distillation
38 viewsDOI:
https://doi.org/10.54939/1859-1043.j.mst.113.2026.150-157Keywords:
Malicious URLs; Graph attention networks; Graph convolutional network; Knowledge distillation.Abstract
Malicious URLs are evolving with increasingly sophisticated obfuscation techniques, posing a persistent threat to global cybersecurity. Traditional deep learning methods still struggle to identify the intricate topological characteristics and hidden structural dependencies inherent in URL sequences. This paper proposes a structural-semantic integration framework leveraging Graph Attention Networks and Knowledge Distillation for robust multi-layer malicious URL detection. We conceptualize raw URLs as character-level directed graphs to extract intricate structural patterns, fused with high-level lexical context distilled from a pretrained DistilBERT model. A hybrid loss integrating Kullback-Leibler divergence and Negative Log-Likelihood balances computational efficiency and predictive performance. Experimental results on large-scale, imbalanced datasets show our framework achieves 98.86% accuracy and 98.16% F1-score.
References
[1]. J. Devlin, M.-W. Chang, K. Lee and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding”, arXiv, (2019).
[2]. V. Sanh, L. Debut, J. Chaumond and T. Wolf, “DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter”, arXiv, (2020).
[3]. P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Liò and Y. Bengio, “Graph attention networks”, arXiv, (2018).
[4]. A. Yilmaz and R. Das, “A novel hybrid approach combining GCN and GAT for effective anomaly detection from firewall logs in campus networks”, Computer Networks, Vol. 259, p. 111082, (2025). DOI: https://doi.org/10.1016/j.comnet.2025.111082
[5]. G. Hinton, O. Vinyals and J. Dean, “Distilling the knowledge in a neural network”, arXiv preprint arXiv:1503.02531, (2015).
[6]. Z. Y. Lim, Y. HanPang, E. C. KahJun and S. Y. Ooi, “Contra-KD: A lightweight transformer model for malicious URL detection with contrastive representation and model distillation”, Future Internet, Vol. 18, No. 157, (2026). DOI: https://doi.org/10.3390/fi18030157
[7]. K. Clark, M.-T. Luong, Q. V. Le and C. D. Manning, “ELECTRA: Pre-training text encoders as discriminators rather than generators”, Eighth International Conference on Learning Representations, Addis Ababa, Ethiopia, (2020).
[8]. A. Aljofey, S. A. Bello, J. Lu and C. Xu, “BERT-PhishFinder: A robust model for accurate phishing URL detection with optimized DistilBERT”, IEEE Transactions on Dependable and Secure Computing, Vol. 22, No. 4, pp. 4315–4329, (2025). DOI: https://doi.org/10.1109/TDSC.2025.3545771
[9]. O. Niyaoui and O. Reda, “Malicious URL detection using transformers’ NLP models and machine learning”, International Conference on Advanced Intelligent Systems for Sustainable Development (AI2SD'2023), Switzerland, (2024). DOI: https://doi.org/10.1007/978-3-031-54318-0_35
[10]. M.-Y. Su and K.-L. Su, “BERT-based approaches to identifying malicious URLs”, Sensors, Vol. 23, No. 20, (2023). DOI: https://doi.org/10.3390/s23208499
[11]. P. Maneriker, J. W. Stokes, E. G. Lazo, D. Carutasu, F. Tajaddodianfar and A. Gururajan, “URLTran: Improving phishing URL detection using transformers”, (2021). DOI: https://doi.org/10.1109/MILCOM52596.2021.9653028
[12]. Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer and V. Stoyanov, “RoBERTa: A robustly optimized BERT pretraining approach”, (2019).
[13]. C. Wang and Y. Chen, “TCURL: Exploring hybrid transformer and convolutional neural network on phishing URL detection”, Knowledge-Based Systems, Vol. 258, (2022). DOI: https://doi.org/10.1016/j.knosys.2022.109955
[14]. C. Wang, Z. Liu and Y. Zeng, “Detecting phishing URLs with GNN-based network inference”, 2024 International Conference on Networking and Network Applications (NaNA), Yinchuan, China, (2024). DOI: https://doi.org/10.1109/NaNA63151.2024.00089
[15]. A. Chattopadhyay, D. Reti and H. D. Schotten, “GNN-based anomaly detection for encoded network traffic”, arXiv, (2024).
[16]. M. I. Hossain, K. A. A. Arafat, B. Shepard, K. Craig and I. Parvez, “A graph-attentive LSTM model for malicious URL detection”, arXiv, (2025). DOI: https://doi.org/10.1109/IETC69527.2026.11568650
[17]. M. Siddhartha, “Malicious URLs Dataset”, Kaggle, (2025). Available online: Malicious URLs Dataset.
[18]. G. Tongo, F. Farid, A. Al-Areqi and F. Ahamed, “Fine-tuned BERT for malicious URL detection”, Western Sydney University. Available online: Fine-Tuned BERT for Malicious URL Detection.
[19]. P. H. Hussan and S. M. Mangj, “BERTPHIURL: A teacher-student learning approach using DistilRoBERTa and RoBERTa for detecting phishing cyber URLs”, Journal of Future Artificial Intelligence and Technologies, Vol. 1, No. 4, pp. 417–428, (2025). DOI: 10.62411/faith.3048-3719-71. DOI: https://doi.org/10.62411/faith.3048-3719-71
[20]. K. S. Jishnu and B. Arthi, “Real-time phishing URL detection framework using knowledge distilled ELECTRA”, Journal for Control, Measurement, Electronics, Computing and Communications, Vol. 65, No. 4, (2024). DOI: 10.1080/00051144.2024.2415797. DOI: https://doi.org/10.1080/00051144.2024.2415797
[21]. R. Zaimi, K. S. Eljil, H. Mohamed, M. Lamia and F. Nait-Abdesselam, “An enhanced mechanism for malicious URL detection using deep learning and DistilBERT-based feature extraction”, The Journal of Supercomputing, Vol. 81, No. 2, (2025). DOI: 10.1007/s11227-024-06908-x. DOI: https://doi.org/10.1007/s11227-024-06908-x
[22]. G. Goldenits, P. König, S. Raubitzek and A. Ekelhart, “Small language models for phishing website detection: Cost, performance, and privacy trade-offs”, Journal of Cybersecurity and Privacy, Vol. 6, No. 2, p. 48, (2026). DOI: 10.3390/jcp6020048. DOI: https://doi.org/10.3390/jcp6020048
[23]. S. E. Blake, “Phishsense-1B: A technical perspective on an AI-powered phishing detection model”, arXiv:2503.10944, (2025).
[24]. Phan Thi Hai Hong, Truong Thi Thu Hang, Dang Van Giap and Ta Huu Vinh, “EB-UNet++: An enhanced crack segmentation network combining EfficientNet-B2 and UNet++ with boundary extraction module”, Journal of Military Science and Technology, No. IITE, pp. 148–159, (2025). DOI: https://doi.org/10.54939/1859-1043.j.mst.IITE.2025.148-159
