Hybrid CNN–Transformer Model for Real-Time Cybercrime Detection
Keywords:
Cybercrime detection, Deep learning, Intrusion detection system, Multi-Head attention, Network security, SHAP explainability, TensorFlow LiteAbstract
The rapid expansion of Internet-connected systems has increased the volume and diversity of malicious network activity, creating a continuing need for intrusion detection methods that can recognize attacks from network traffic patterns rather than relying only on previously known signatures. This article presents a Hybrid CNN–Transformer (HCNT) architecture for real-time, multi-class cybercrime detection in network traffic. The approach combines one-dimensional convolutional neural networks, which are used to learn local feature interactions, with Transformer encoder layers that model longer-range relationships through multi-head self-attention. The proposed architecture uses three dilated residual convolutional blocks, learnable positional encodings, and a two-layer Transformer encoder followed by a multi-layer classification head. Five traffic classes are considered: normal traffic, Denial-of-Service (DoS), probing, Remote-to-Local (R2L), and User-to-Root (U2R). The study uses 5,000 labelled network connection records with 41 features. Class imbalance is addressed using SMOTE, while feature standardisation is fitted only on the training portion to reduce the risk of data leakage. In the reported experiment, the model reaches 97.20% test accuracy, a macro-averaged F1-score of 0.9639, and a macro ROC-AUC of 0.9971. An INT8 TensorFlow Lite version reduces the model size and records a mean inference latency of 1.87 ms in the reported benchmark. SHAP DeepExplainer is also used to examine feature contributions for different attack classes. The results indicate that combining local feature extraction with global dependency modelling can provide an effective framework for the stated experimental setting, while the limited dataset and absence of adversarial and long-term drift evaluation remain important limitations.
References
Cybersecurity Ventures, "Cybercrime Report 2024," 2024.
V. Paxson, “Bro: a system for detecting network intruders in real-time,” Computer Networks, vol. 31, no. 23–24, pp. 2435–2463, Dec. 1999.
D. E. Denning, “An Intrusion-Detection Model,” IEEE Transactions on Software Engineering, vol. SE-13, no. 2, pp. 222–232, Feb. 1987.
G. Stein, B. Chen, A. S. Wu, and K. A. Hua, “Decision tree classifier for network intrusion detection with GA-based feature selection,” Proceedings of the 43rd annual southeast regional conference on - ACM-SE 43, 2005.
S. Mukkamala, G. Janoski, and A. Sung, “Intrusion detection using neural networks and support vector machines,” IEEE Xplore, 2002.
T. Chen and C. Guestrin, “XGBoost: a Scalable Tree Boosting System,” Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining - KDD ’16, vol. 1, no. 1, pp. 785–794, Aug. 2016.
J. Kim, J. Kim, H. L. Thu, and H. Kim, “Long Short Term Memory Recurrent Neural Network Classifier for Intrusion Detection,” 2016 International Conference on Platform Technology and Service (PlatCon), 2016.
C. Yin, Y. Zhu, J. Fei, and X. He, “A Deep Learning Approach for Intrusion Detection Using Recurrent Neural Networks,” IEEE Access, vol. 5, pp. 21954–21961, 2017.
Y. Mirsky, T. Doitshman, Y. Elovici, and A. Shabtai, “Kitsune: An Ensemble of Autoencoders for Online Network Intrusion Detection,” arXiv.org, May 27, 2018.
A. Vaswani et al., “Attention Is All You Need,” arXiv.org, 2017.
K. Jiang, W. Wang, A. Wang, and H. Wu, “Network Intrusion Detection Combined Hybrid Sampling With Deep Hierarchical Network,” IEEE Access, vol. 8, pp. 32464–32476, 2020.
P. Lin, K. Ye, and C.-Z. Xu, “Dynamic Network Anomaly Detection System by Using Deep Learning Techniques,” Cloud Computing – CLOUD 2019, pp. 161–176, 2019.
X. Gao, C. Shan, C. Hu, Z. Niu, and Z. Liu, “An Adaptive Ensemble Machine Learning Model for Intrusion Detection,” IEEE Access, vol. 7, pp. 82512–82521, 2019.
N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “SMOTE: Synthetic Minority Over-sampling Technique,” Journal of Artificial Intelligence Research, vol. 16, no. 16, pp. 321–357, June 2002.
S. Lundberg and S.-I. Lee, “A Unified Approach to Interpreting Model Predictions,” arXiv.org, Nov. 24, 2017.
J. Bjorck, K. Q. Weinberger, and C. Gomes, “Understanding Decoupled and Early Weight Decay,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 8, pp. 6777–6785, May 2021.
M. Tavallaee, E. Bagheri, W. Lu, and A. A. Ghorbani, “A detailed analysis of the KDD CUP 99 data set,” 2009 IEEE Symposium on Computational Intelligence for Security and Defense Applications, pp. 1–6, July 2009.
D. Hendrycks and K. Gimpel, “Gaussian Error Linear Units (GELUs),” arXiv:1606.08415, June 2016.
W. Wang et al., “HAST-IDS: Learning Hierarchical Spatial-Temporal Features Using Deep Neural Networks to Improve Intrusion Detection,” IEEE Access, vol. 6, pp. 1792–1806, 2018.
H. Zhang, J.-L. Li, X.-M. Liu, and C. Dong, “Multi-dimensional feature fusion and stacking ensemble mechanism for network intrusion detection,” Future Generation Computer Systems, vol. 122, pp. 130–143, Sept. 2021.