Systematic Review of Machine Learning and Deep Learning Approaches for Spam Detection in Online Social Networks
Abstract
Malicious activities such as phishing, disinformation, and bot attacks for spamming in online social networks (OSNs) have become an increasing problem, and online spam detection is an ongoing challenge to keep users safe and OSNs healthy. However, classical machine learning (ML) algorithms like Random Forest (RF), Decision Trees (DT), k-Nearest Neighbors (KNNs), and Naïve Bayes (NB) have various limitations, including class imbalance, overfitting, and scalability, in the context of OSN environments. Emerging deep learning (DL) algorithms, including Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), and Long Short-Term Memory (LSTM) networks, are more tolerant of high dimensionality and temporal data. With recent developments, Graph Neural Networks (GNNs), attention mechanisms, and transformers, along with ensemble learning, are being used to model structural and semantic spam patterns. It presents the results of a systematic review of 25 peer-reviewed papers published between 2020 and 2025 and discusses them in terms of approaches, performance assessment, and limitations. The review found that hybrid DL approaches generally achieved better spam detection performance than traditional machine learning (ML) methods, although at the cost of increased computational complexity and resource requirements. However, since 2020, no systematic synthesis has examined the trade-offs between the performance and efficiency of hybrid models. We observed that hybrid models that integrate GNNs, attention mechanisms, and optimization techniques have great promise for future research and can be used to provide real-time and scalable spam detection for multiple social platforms.
Keywords
Funding
This research received no external funding.
References
- Schmitt, M.; Flechais, I. Digital Deception: Generative Artificial Intelligence in Social Engineering and Phishing. Artif. Intell. Rev. 2024, 57, https://doi.org/10.1007/s10462-024-10973-2.
- Mouncey, E.; Ciobotaru, S. Phishing Scams on Social Media: An Evaluation of Cyber Awareness Education on Impact and Effectiveness. Journal of Economic Criminology 2025, 7, 100125, https://doi.org/10.1016/J.JECONC.2025.100125.
- Simon, M.; Ya’u, B.I.; Gital, A.Y.; Magombe, Y.; Alli, A.; Maishanu, M. A Hybrid Convolutional Neural Network-Gated Recurrent Unit Model for Fake News Detection on Social Media. In Proceedings of the 2026 2nd International Symposium on AI-Driven Engineering Systems (ISADES); IEEE, 2026; pp. 1–5. https://doi.org/10.1109/ISADES69945.2026.11608227
- Pierri, F.; Luceri, L.; Jindal, N.; Ferrara, E. Propaganda and Misinformation on Facebook and Twitter during the Russian Invasion of Ukraine. ACM International Conference Proceeding Series 2023, 65–74, https://doi.org/10.1145/3578503.3583597.
- Kawintiranon, K.; Singh, L.; Budak, C. Traditional and Context-Specific Spam Detection in Low Resource Settings. Mach. Learn. 2022, 111, 2515–2536, https://doi.org/10.1007/S10994-022-06176-X.
- Shaaban, M.A.; Hassan, Y.F.; Guirguis, S.K. Deep Convolutional Forest: A Dynamic Deep Ensemble Approach for Spam Detection in Text. Complex and Intelligent Systems 2022, 8, 4897–4909, https://doi.org/10.1007/s40747-022-00741-6.
- Minaee, S.; Kalchbrenner, N.; Cambria, E.; Nikzad, N.; Chenaghlu, M.; Gao, J. Deep Learning-Based Text Classification. ACM Comput. Surv. 2022, 54, https://doi.org/10.1145/3439726
- Rao, S.; Verma, A.K.; Bhatia, T. Hybrid Ensemble Framework with Self-Attention Mechanism for Social Spam Detection on Imbalanced Data. Expert Syst. Appl. 2023, 217, 119594, https://doi.org/10.1016/J.ESWA.2023.119594.
- Zhou, M.; Zhang, D.; Wang, Y.; Geng, Y.A.; Dong, Y.; Tang, J. LGB: Language Model and Graph Neural Network-Driven Social Bot Detection. IEEE Trans. Knowl. Data Eng. 2024, 37, 4728–4742, https://doi.org/10.1109/TKDE.2025.3573748.
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 Statement: An Updated Guideline for Reporting Systematic Reviews. BMJ 2021, 372, 790–799, https://doi.org/10.1136/BMJ.N71.
- Wang, K.; Ding, Y.; Han, S.C. Graph Neural Networks for Text Classification: A Survey. Artif. Intell. Rev. 2024, 57, https://doi.org/10.1007/s10462-024-10808-0.
- Islam, S.; Elmekki, H.; Elsebai, A.; Bentahar, J.; Drawel, N.; Rjoub, G.; Pedrycz, W. A Comprehensive Survey on Applications of Transformers for Deep Learning Tasks. Expert Syst. Appl. 2023, 241, https://doi.org/10.1016/j.eswa.2023.122666.
- Lundberg, S.M.; Erion, G.; Chen, H.; DeGrave, A.; Prutkin, J.M.; Nair, B.; Katz, R.; Himmelfarb, J.; Bansal, N.; Lee, S.I. From Local Explanations to Global Understanding with Explainable AI for Trees. Nature Machine Intelligence 2020 2:1 2020, 2, 56–67, https://doi.org/10.1038/s42256-019-0138-9.
- Ribeiro, M.T.; Singh, S.; Guestrin, C. “Why Should i Trust You?” Explaining the Predictions of Any Classifier. Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 2016, 13-17-August-2016, 1135–1144, https://doi.org/10.1145/2939672.2939778
- Bhattacharya, M.; Roy, S.; Chattopadhyay, S.; Das, A.K.; Shetty, S. A Comprehensive Survey on Online Social Networks Security and Privacy Issues: Threats, Machine Learning-Based Solutions, and Open Challenges. Security and Privacy 2023, 6, e275, https://doi.org/10.1002/SPY2.275.
- Xiao, A.S.; Liang, Q. Spam Detection for Youtube Video Comments Using Machine Learning Approaches. Machine Learning with Applications 2024, 16, 100550, https://doi.org/10.1016/J.MLWA.2024.100550.
- Gangwar, S.S.; Rathore, S.S.; Chouhan, S.S.; Soni, S. Predictive Modeling for Suspicious Content Identification on Twitter. Soc. Netw. Anal. Min. 2022, 12, 149, https://doi.org/10.1007/S13278-022-00977-7.
- Velammal, B.L.; Aarthy, N. Improvised Spam Detection in Twitter Data Using Lightweight Detectors and Classifiers. International Journal of Web-Based Learning and Teaching Technologies 2021, 16, 12–32, https://doi.org/10.4018/IJWLTT.20210701.OA2.
- Alharthi, R.; Alhothali, A.; Moria, K. A Real-Time Deep-Learning Approach for Filtering Arabic Low-Quality Content and Accounts on Twitter. Inf. Syst. 2021, 99, 101740, https://doi.org/10.1016/J.IS.2021.101740.
- Hayami, R.; Al Amien, J.; Utami, N.T. Tweet Spam Detection Using the Convolutional Neural Network (CNN) Model. AIP Conf. Proc. 2023, 2601, https://doi.org/10.1063/5.0130468/2894141.
- Loucif, H. A Hybrid Deep Learning Approach for Spam Detection in Twitter. Ingenierie des Systemes d’Information 2024, 29, 117–123, https://doi.org/10.18280/ISI.290113.
- Xu, G.; Zhou, D.; Liu, J. Social Network Spam Detection Based on ALBERT and Combination of Bi-LSTM with Self-Attention. Security and Communication Networks 2021, 2021, 5567991, https://doi.org/10.1155/2021/5567991.
- Kumar, C.; Bharti, T.S.; Prakash, S. A Hybrid Data-Driven Framework for Spam Detection in Online Social Network. Procedia Comput. Sci. 2023, 218, 124–132, https://doi.org/10.1016/J.PROCS.2022.12.408.
- Pal, M.B.; Agrawal, S. Graph Neural Network-Based Attention Mechanism to Classify Spam Review over Heterogeneous Social Networks. J. Supercomput. 2024, 80, 27176–27203, https://doi.org/10.1007/S11227-024-06459-1.
- Dou, Y.; Liu, Z.; Sun, L.; Deng, Y.; Peng, H.; Yu, P.S. Enhancing Graph Neural Network-Based Fraud Detectors against Camouflaged Fraudsters. International Conference on Information and Knowledge Management, Proceedings 2020, 315–324, https://doi.org/10.1145/3340531.3411903.
- Zhao, C.; Xin, Y.; Li, X.; Zhu, H.; Yang, Y.; Chen, Y. An Attention-Based Graph Neural Network for Spam Bot Detection in Social Networks. Applied Sciences 2020, Vol. 10, Page 8160 2020, 10, 8160, https://doi.org/10.3390/APP10228160.
- Shen, H.; Liu, X.; Zhang, X. Boosting Social Spam Detection via Attention Mechanisms on Twitter. Electronics 2022, Vol. 11, Page 1129 2022, 11, 1129, https://doi.org/10.3390/ELECTRONICS11071129.
- Sallah, A.; Arbi Abdellaoui Alaoui, E.; Agoujil, S.; Ahmad Wani, M.; Hammad, M.; Maleh, Y.; Abd El-Latif, A.A. Fine-Tuned Understanding: Enhancing Social Bot Detection With Transformer-Based Classification. IEEE Access 2024, 12, 118250–118269, https://doi.org/10.1109/ACCESS.2024.3440657.
- Nasser, M.; Saeed, F.; Da’u, A.; Alblwi, A.; Al-Sarem, M. Topic-Aware Neural Attention Network for Malicious Social Media Spam Detection. Alexandria Engineering Journal 2025, 111, 540–554, https://doi.org/10.1016/J.AEJ.2024.10.073.
- Iqbal, A.; Younas, M.; Iftikhar, S.; Fatima, F.; Saleem, R. Spam Detection Using Hybrid Model on Fusion of Spammer Behavior and Linguistics Features. Egyptian Informatics Journal 2025, 29, 100605, https://doi.org/10.1016/J.EIJ.2024.100605.
- Citlak, O.; Dorterler, M.; Dogru, İ. A Hybrid Spam Detection Framework for Social Networks. Politeknik Dergisi 2023, 26, 823–837, https://doi.org/10.2339/POLITEKNIK.933785.
- Bilgen, Y.; Kaya, M. EGMA: Ensemble Learning-Based Hybrid Model Approach for Spam Detection. Applied Sciences 2024, Vol. 14, Page 9669 2024, 14, 9669, https://doi.org/10.3390/APP14219669.
- Elakkiya, E.; Selvakumar, S. Stratified Hyperparameters Optimization of Feed-Forward Neural Network for Social Network Spam Detection (SON2S). Soft Computing 2022 26:21 2022, 26, 11915–11934, https://doi.org/10.1007/S00500-022-07020-Z.
- Abualhaj, M.M.; Daoud, M.S.; Al-Allawee, A.; Al-Mimi, H.; Al-Khatib, S.; Hyassat, A.S. Detecting Spam Using K-Nearest Neighbors and Metaheuristic Feature Selection Algorithms. International Conference on Social Networks Analysis, Management and Security, SNAMS 2025, 249–252, https://doi.org/10.1109/SNAMS67467.2025.11390963.
- Shukla, P.K.; Veerasamy, B.D.; Alduaiji, N.; Addula, S.R.; Pandey, A.; Shukla, P.K. Fraudulent Account Detection in Social Media Using Hybrid Deep Transformer Model and Hyperparameter Optimization. Scientific Reports 2025 15:1 2025, 15, 38447-, https://doi.org/10.1038/s41598-025-24326-8.
- Ghatasheh, N.; Altaharwa, I.; Aldebei, K. Modified Genetic Algorithm for Feature Selection and Hyper Parameter Optimization: Case of XGBoost in Spam Prediction. IEEE Access 2022, 10, 84365–84383, https://doi.org/10.1109/ACCESS.2022.3196905.
- Rashidi, A.; Salehi, M.; Najari, S. CGANS: A Code-Based GAN for Spam Detection in Social Media. Social Network Analysis and Mining 2024 14:1 2024, 14, 218-, https://doi.org/10.1007/S13278-024-01379-7.
- Chandana; Anuradha; Naik, A.J.; Kumar, A.; Banu, S. A Framework for Twitter Spam Detection and Reporting. AIP Conf. Proc. 2024, 2742, https://doi.org/10.1063/5.0184161/3263654.