A comprehensive analysis of adversarial attacks against spam filters
Computers and Security, cilt.171, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 171
- Basım Tarihi: 2026
- Doi Numarası: 10.1016/j.cose.2026.105066
- Dergi Adı: Computers and Security
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, ABI/INFORM, Aerospace Database, Applied Science & Technology Source, Compendex, Criminal Justice Abstracts, INSPEC, Criminal Justice Periodical Index, Social Science Premium Collection (ProQuest), Business Source Ultimate (EBSCO), Criminology Collection (ProQuest), Technology Collection (ProQuest)
- Anahtar Kelimeler: Adversarial learning, Deep learning, Email security, Natural language processing, Spam detection
- Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
- Hacettepe Üniversitesi Adresli: Evet
Özet
Deep learning has revolutionized email filtering, which is critical to protect users from cyber threats such as spam, malware, and phishing. However, the increasing sophistication of adversarial attacks poses a significant challenge to the effectiveness of these filters. This study investigates the impact of adversarial attacks on deep learning-based spam detection systems using real-world datasets. Six prominent deep learning models are evaluated on these datasets, analyzing attacks at the word, character sentence, and AI-generated paragraph-levels. Novel scoring functions, including spam weights and attention weights, are introduced to improve attack effectiveness. A key contribution of this study is the analysis of spam-weight- and attention-weight-based scoring functions, highlighting their role in improving the effectiveness and efficiency of adversarial attacks. This comprehensive analysis sheds light on the vulnerabilities of spam filters and contributes to efforts to improve their security against evolving adversarial threats. Experimental results show that word- and sentence-level attacks markedly increase false negatives, while character-level perturbations disrupt token representations with minimal semantic change. AI-generated paragraph-level attacks remain challenging even for transformer-based models. In addition, spam-weight-based scoring consistently enables more effective adversarial attacks than alternative scoring strategies with lower computational cost.