Hybrid Transformer Architecture for Context-Aware Spam Email Classification
Abstract
1. Introduction
- Construct a large, multi-source, multilingual email corpus and audit it for source–label association and campaign-level duplication prior to partitioning.
- Fine-tune a Transformer-based model (DistilBERT) and integrate its contextual embeddings with engineered structural features to build a multimodal classification pipeline.
- Evaluate all models under identical training configurations and two evaluation protocols, quantifying the generalisation gap between in-distribution and cross-source conditions.
- Establish which structural features contribute to detection through retraining-based ablation rather than attribution proxies, and verify stability through cross-validation, multi-seed replication, and bootstrap confidence intervals.
- Apply SHAP and LIME to the multimodal classifier to provide transparency consistent with GDPR accountability requirements.
- Characterise the class composition of each language above a minimum-count threshold to establish which support classification metrics.
- Multimodal fusion as a robustness mechanism, not only an accuracy improvement: Structural features fused with a Transformer encoder produce a small but statistically significant benefit under in-distribution evaluation (F1 = 0.0014, 95% CI [0.0005, 0.0023]) that grows roughly five-and-a-half-fold under cross-source evaluation (F1 = 0.0078, CI [0.0036, 0.0120]). This shift-dependent scaling is invisible to the protocol used throughout the prior literature, which reports only in-distribution accuracy, and is established here through a three-arm design in which head architecture is held constant and only the structural block varies.
- A leakage-audited multilingual corpus: A 99,707-email corpus is assembled and audited for metadata leakage, with sender domain excluded as a provenance identifier rather than a phishing signal.
- A dual-protocol evaluation that separates memorisation from generalisation: Nine models spanning classical, convolutional, Transformer, and multimodal families are evaluated under in-distribution and cross-source protocols.
- Retraining-based ablation of structural features: Rather than inferring feature importance from attribution scores, each structural-feature configuration is retrained and re-measured. In distribution, the top-3 SHAP-ranked structural features reduce false negatives by up to 17.7% (124 → 102) with F1 essentially unchanged (0.9862–0.9878 across configurations); under cross-source evaluation, the same features yield a significant advantage over both text-only and other variants.
- Interval-based robustness reporting: Results are validated through five-fold stratified cross-validation, five-seed replication, bootstrap confidence intervals, and McNemar tests, with intervals reported rather than single-run point estimates. Seed variance under cross-source evaluation proves an order of magnitude larger than in distribution, and model ordering is shown to be unstable there—a result that qualifies single-run rankings throughout this literature.
- Cross-lingual evaluability analysis: The class composition of every language above a minimum-count threshold is characterised, establishing which languages can support classification metrics and which can support detection rate alone—a prerequisite routinely omitted when multilingual accuracy is reported in aggregate.
- Joint SHAP and LIME attribution: Both explainers are applied to the same instances on the multimodal classifier, yielding cross-validated attributions that support model debugging and regulatory transparency under GDPR, alongside the demonstration that attribution magnitude does not predict a feature’s conditional contribution under shift.
2. Literature Review
2.1. From Rule-Based Filtering to Classical Machine Learning
2.2. Deep Learning and Transformer Architectures
2.3. Multimodal Fusion and Explainable AI
2.4. Evaluation Practice and Research Gaps
3. Methodology
3.1. Corpus Construction
3.2. Leakage Auditing
3.3. Split Design
3.4. Models
3.5. Training Configuration
3.6. Evaluation Protocol
3.7. Explainability Methods
3.8. Ethical Considerations and Reproducibility
4. Results and Discussion
4.1. Training Dynamics
4.2. In-Distribution Performance
4.3. Cross-Source Generalisation
4.4. Ablation
4.5. Statistical Robustness
4.6. Operating Point Selection
4.7. Cross-Lingual Evaluability
4.8. Explainability Results
4.9. Comparison with Prior Work
5. Conclusions
6. Future Work
- Rotating cross-source evaluation refers to withholding each of the three TREC collections in turn, each of which contains both classes, so that provenance shift can be isolated from the change in class composition that accompanies withholding the honeypot source. Adding temporal splits would further separate general provenance sensitivity from an artefact of the particular collection held out here.
- Balanced multilingual evaluation refers to assembling non-English legitimate mail so full classification metrics can be reported beyond English, with multilingual encoders such as XLM-R as comparison.
- Generality of the conditional-contribution effect refers to testing whether other feature families with small in-distribution contributions become significantly more important under shift and whether the effect observed here holds for shift types other than provenance. If it generalises, in-distribution ablation is unsafe as a feature selection procedure wherever deployment data may differ from training data.
- Adversarial robustness refers to stress-testing both feature families against character substitution, homoglyph attacks, and perturbed text. The cross-source results suggest a testable hypothesis: structural features, which transfer better across sources, may prove more brittle under deliberate evasion precisely because they are simple to manipulate.
- AI-generated phishing refers to evaluating detection of generative-model output, which lacks the templated repetition that campaign-level deduplication exploits.
- Human-centred explainability assessment refers to measuring whether SHAP and LIME attributions improve analyst triage accuracy and calibrated trust, rather than assuming that available explanations are useful ones.
- Deployment cost characterisation refers to profiling the full pipeline including embedding generation, where the encoder rather than the fusion stage is the binding constraint, and assessing quantisation or distillation.
Author Contributions
Funding
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Cormack, G.V. Email spam filtering: A systematic review. Found. Trends Inf. Retr. 2008, 1, 335–455. [Google Scholar] [CrossRef] [Scilit]
- Hong, J. The state of phishing attacks. Commun. ACM 2012, 55, 74–81. [Google Scholar] [CrossRef] [Scilit]
- Jáñez-Martino, F.; Alaiz-Rodríguez, R.; González-Castro, V.; Fidalgo, E.; Alegre, E. A review of spam email detection: Analysis of spammer strategies and the dataset shift problem. Artif. Intell. Rev. 2023, 56, 1145–1173. [Google Scholar] [CrossRef] [Scilit]
- Zimba, A.; Wang, Z.; Chen, H. Multi-stage crypto ransomware attacks: A new emerging cyber threat to critical infrastructure and industrial control systems. ICT Express 2018, 4, 14–18. [Google Scholar] [CrossRef] [Scilit]
- Liu, X. Deciphering Spam Through AI: From Traditional Methods to Deep Learning Advancements in Email Security. In Proceedings of the 1st International Conference on Engineering Management, Information Technology and Intelligence (EMITI 2024); SCITEPRESS: Setúbal, Portugal, 2024; pp. 553–558. [Google Scholar]
- Pathak, A.; Hu, Y.C.; Mao, Z.M. Peeking into spammer behavior from a unique vantage point. In Proceedings of the 1st USENIX Workshop on Large-Scale Exploits and Emergent Threats (LEET), San Francisco, CA, USA, 15 April 2008. [Google Scholar]
- Kallepalli, K.; Chaudhry, U.B. Intelligent Security: Applying Artificial Intelligence to Detect Advanced Cyber Attacks. In Challenges in the IoT and Smart Environments: A Practitioners’ Guide to Security, Ethics and Criminal Threats; Springer: Cham, Switzerland, 2021. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Pachare, R.; Banarase, P.; Dhanke, P.; Dhakade, A. Advancements in Email Spam Detection: A Systematic Review of Machine Learning and Deep Learning Techniques. Int. J. Res. Appl. Sci. Eng. Technol. 2025, 13, 1908–1914. [Google Scholar] [CrossRef] [Scilit]
- Ma, J.; Saul, L.K.; Savage, S.; Voelker, G.M. Beyond blacklists: Learning to detect malicious web sites from suspicious URLs. In Proceedings of the 15th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Paris, France, 28 June–1 July 2009; pp. 1245–1254. [Google Scholar] [CrossRef] [Scilit]
- Asliyuksek, H.; Tonkal, O.; Kocaoglu, R. A Comparative Evaluation of a Multimodal Approach for Spam Email Classification Using DistilBERT and Structural Features. Electronics 2025, 14, 3855. [Google Scholar] [CrossRef] [Scilit]
- Metsis, V.; Androutsopoulos, I.; Paliouras, G. Spam filtering with Naive Bayes—Which Naive Bayes? In Proceedings of the 3rd Conference on Email and Anti-Spam (CEAS), Mountain View, CA, USA, 27–28 July 2006. [Google Scholar]
- Klimt, B.; Yang, Y. The Enron corpus: A new dataset for email classification research. In Proceedings of the 15th European Conference on Machine Learning (ECML); Springer: Berlin/Heidelberg, Germany, 2004; pp. 217–226. [Google Scholar] [CrossRef] [Scilit]
- Jamal, S.; Wimmer, H. An Improved Transformer-based Model for Detecting Phishing, Spam, and Ham: A Large Language Model Approach. arXiv 2023, arXiv:2311.04913. [Google Scholar]
- Shirvani, G.; Ghasemshirazi, S. Advancing Email Spam Detection: Leveraging Zero-Shot Learning and Large Language Models. arXiv 2025, arXiv:2505.02362. [Google Scholar]
- Sanh, V.; Debut, L.; Chaumond, J.; Wolf, T. DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter. arXiv 2020, arXiv:1910.01108. [Google Scholar]
- Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. BERT: Pre-training of deep bidirectional Transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), Minneapolis, MN, USA, 2–7 June 2019; pp. 4171–4186. [Google Scholar]
- Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I. Improving Language Understanding by Generative Pre-Training; Technical Report; OpenAI: San Francisco, CA, USA, 2018. [Google Scholar]
- Kaufman, S.; Rosset, S.; Perlich, C.; Stitelman, O. Leakage in data mining: Formulation, detection, and avoidance. ACM Trans. Knowl. Discov. Data 2012, 6, 15. [Google Scholar] [CrossRef] [Scilit]
- Kapoor, S.; Narayanan, A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns 2023, 4, 100804. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Conneau, A.; Khandelwal, K.; Goyal, N.; Chaudhary, V.; Wenzek, G.; Guzmán, F.; Grave, E.; Ott, M.; Zettlemoyer, L.; Stoyanov, V. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL); Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 8440–8451. [Google Scholar] [CrossRef] [Scilit]
- Hotoğlu, E.; Sen, S.; Can, B. A Comprehensive Analysis of Adversarial Attacks against Spam Filters. arXiv 2025, arXiv:2505.03831. [Google Scholar]
- Reddy, M.A.S.K. Advanced Techniques in Spam Detection: A Comparative Analysis of Traditional and AI-Based Approaches. Int. J. Res. Trends Innov. 2025, 10, a598–a609. [Google Scholar]
- Kontsewaya, Y.; Antonov, E.; Artamonov, A. Evaluating the effectiveness of machine learning methods for spam detection. Procedia Comput. Sci. 2021, 190, 479–486. [Google Scholar] [CrossRef] [Scilit]
- Kim, Y. Convolutional neural networks for sentence classification. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar, 25–29 October 2014; pp. 1746–1751. [Google Scholar] [CrossRef] [Scilit]
- Lai, S.; Xu, L.; Liu, K.; Zhao, J. Recurrent convolutional neural networks for text classification. In Proceedings of the 29th AAAI Conference on Artificial Intelligence, Austin, TX, USA, 25–30 January 2015; pp. 2267–2273. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Jin, R.; Zhou, Z.H. Understanding bag-of-words model: A statistical framework. Int. J. Mach. Learn. Cybern. 2010, 1, 43–52. [Google Scholar] [CrossRef] [Scilit]
- Mikolov, T.; Sutskever, I.; Chen, K.; Corrado, G.; Dean, J. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems; Curran Associates Inc.: Red Hook, NY, USA, 2013; Volume 26, pp. 3111–3119. [Google Scholar]
- Pennington, J.; Socher, R.; Manning, C.D. GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar, 25–29 October 2014; pp. 1532–1543. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention Is All You Need. In Advances in Neural Information Processing Systems; Curran Associates Inc.: Red Hook, NY, USA, 2017; Volume 30, pp. 5998–6008. [Google Scholar]
- Liu, Y.; Ott, M.; Goyal, N.; Stoyanov, V. RoBERTa: A robustly optimized BERT pretraining approach. arXiv 2019, arXiv:1907.11692. [Google Scholar]
- Young, T.; Hazarika, D.; Poria, S.; Cambria, E. Recent trends in deep learning based natural language processing. IEEE Comput. Intell. Mag. 2018, 13, 55–75. [Google Scholar] [CrossRef] [Scilit]
- Ubale, K.S.; Shirsath, K.A. SpamNet: A hybrid deep learning framework for robust spam email detection using multi-modal features. Int. J. Appl. Math. 2025, 38, 670–693. [Google Scholar] [CrossRef] [Scilit]
- Hijji, M.; Alam, G. A Multivocal Literature Review on Growing Social Engineering Based Cyber-Attacks/Threats During the COVID-19 Pandemic: Challenges and Prospective Solutions. IEEE Access 2021, 9, 7152–7169. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lundberg, S.M.; Lee, S.I. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems; Curran Associates Inc.: Red Hook, NY, USA, 2017; Volume 30, pp. 4765–4774. [Google Scholar]
- Ribeiro, M.T.; Singh, S.; Guestrin, C. “Why should I trust you?”: Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 1135–1144. [Google Scholar] [CrossRef] [Scilit]
- Arrieta, A.B.; Díaz-Rodríguez, N.; Del Ser, J.; Bennetot, A.; Herrera, F. Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Inf. Fusion 2020, 58, 82–115. [Google Scholar] [CrossRef] [Scilit]
- Jain, S.; Wallace, B.C. Attention is not explanation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), Minneapolis, MN, USA, 2–7 June 2019; pp. 3543–3556. [Google Scholar] [CrossRef] [Scilit]
- EPSRC. Framework for Responsible Innovation; Engineering and Physical Sciences Research Council: Swindon, UK, 2020.
- Gorman, K.; Bedrick, S. We need to talk about standard splits. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), Florence, Italy, 28 July–2 August 2019; pp. 2786–2791. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Torralba, A.; Efros, A.A. Unbiased look at dataset bias. In Proceedings of the 2011 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Colorado Springs, CO, USA, 20–25 June 2011; pp. 1521–1528. [Google Scholar] [CrossRef] [Scilit]
- Geirhos, R.; Jacobsen, J.H.; Michaelis, C.; Zemel, R.; Brendel, W.; Bethge, M.; Wichmann, F.A. Shortcut learning in deep neural networks. Nat. Mach. Intell. 2020, 2, 665–673. [Google Scholar] [CrossRef] [Scilit]
- Recht, B.; Roelofs, R.; Schmidt, L.; Shankar, V. Do ImageNet classifiers generalize to ImageNet? In Proceedings of the 36th International Conference on Machine Learning (ICML), Long Beach, CA, USA, 9–15 June 2019; pp. 5389–5400. [Google Scholar]
- Cormack, G.V.; Lynam, T.R. TREC 2005 Spam Track Overview. In Proceedings of the Fourteenth Text REtrieval Conference (TREC 2005); Voorhees, E.M., Buckland, L.P., Eds.; NIST Special Publication 500-266; NIST: Gaithersburg, MD, USA, 2005. Available online: https://trec.nist.gov/pubs/trec14/papers/SPAM.OVERVIEW.pdf (accessed on 18 August 2026).
- Cormack, G.V. TREC 2006 Spam Track Overview. In Proceedings of the Fifteenth Text REtrieval Conference (TREC 2006); Voorhees, E.M., Buckland, L.P., Eds.; NIST Special Publication 500-272; NIST: Gaithersburg, MD, USA, 2006. Available online: https://trec.nist.gov/pubs/trec15/papers/SPAM06.OVERVIEW.pdf (accessed on 18 August 2026).
- Cormack, G.V. TREC 2007 Spam Track Overview. In Proceedings of the Sixteenth Text REtrieval Conference (TREC 2007); Voorhees, E.M., Buckland, L.P., Eds.; NIST Special Publication 500-274; NIST: Gaithersburg, MD, USA, 2007. Available online: https://trec.nist.gov/pubs/trec16/papers/SPAM.OVERVIEW16.pdf (accessed on 18 August 2026).
- Peixoto, R.F. phishing_pot: A Collection of Phishing Samples. GitHub Repository. 2024. Available online: https://github.com/rf-peixoto/phishing_pot (accessed on 18 August 2026).
- Baltrušaitis, T.; Ahuja, C.; Morency, L.P. Multimodal machine learning: A survey and taxonomy. IEEE Trans. Pattern Anal. Mach. Intell. 2019, 41, 423–443. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ramachandram, D.; Taylor, G.W. Deep multimodal learning: A survey on recent advances and trends. IEEE Signal Process. Mag. 2017, 34, 96–108. [Google Scholar] [CrossRef] [Scilit]
- Loshchilov, I.; Hutter, F. Decoupled weight decay regularization. In Proceedings of the 7th International Conference on Learning Representations (ICLR), New Orleans, LA, USA, 6–9 May 2019. [Google Scholar]
- Lipton, Z.C. The mythos of model interpretability: In supervised learning, the highest-performing models are often the least transparent. Commun. ACM 2018, 61, 36–43. [Google Scholar] [CrossRef] [Scilit]
- Goodman, B.; Flaxman, S. EU regulations on algorithmic decision-making and a “right to explanation”. AI Mag. 2017, 38, 50–57. [Google Scholar] [CrossRef] [Scilit]
- Mireshghallah, F.; Taram, M.; Vepakomma, P.; Singh, A.; Raskar, R.; Esmaeilzadeh, H. Privacy in deep learning: A survey. arXiv 2020, arXiv:2004.12254. [Google Scholar]

















| Study | Method | Langs. | Explainability | Cross-Source Eval. |
|---|---|---|---|---|
| Kontsewaya et al. [23] | SVM, NB, LR | 1 | None | No |
| Kim [24] | CNN | 1 | None | No |
| Lai et al. [25] | Recurrent CNN | 1 | None | No |
| Jamal & Wimmer [13] | Fine-tuned BERT | 1 | Attention | No |
| Asliyuksek et al. [10] | DistilBERT + structural | 1 | None | No |
| Ubale & Shirsath [32] | SpamNet (DL) | 1 | None | No |
| Shirvani et al. [14] | Zero-shot FLAN-T5 | 1 | Partial | No |
| This paper | DistilBERT + MLP | 91 (4 eval.) | SHAP + LIME | Yes |
| Setting | Value |
|---|---|
| Transformer encoder | |
| Checkpoint | distilbert-base-uncased (67.0 M parameters per encoder) |
| Max sequence length | 256 tokens |
| Batch size | 32 |
| Learning rate | |
| Epoch budget policy | Start 10, extend by 5 to ceiling 25; stop on early stopping (patience 3 on validation loss) or plateau (range < 0.002 over last 4 epochs); checkpoint = epoch of lowest validation loss |
| Epochs run (A, B) | 11, 7 (best epoch: 8, 4)—both terminated by early stopping |
| Regularisation (weight decay/ label smoothing/dropout) | 0.05/0.1/0.2 (hidden, attn), 0.3 (classifier) |
| Frozen parameters | Embeddings + lowest 2 of 6 blocks (28.9 M trainable) |
| Embedding pooling | Mean over non-padding tokens of the final layer |
| Shared optimisation | |
| Optimiser | AdamW, decoupled weight decay (0.05 DistilBERT, 0.01 elsewhere) |
| LR schedule | Linear warmup over 10% of steps, then linear decay to zero |
| Gradient clipping | Max-norm 1.0 |
| Mixed precision | fp16 |
| Loss | Cross-entropy with class weights from each design’s training partition |
| Imbalance handling | Class weighting only; no oversampling, undersampling, or synthetic data |
| Classification heads | |
| MLP heads | (256, 64) ReLU, dropout 0.2, AdamW lr , batch 256, early stopping patience 10 (max 100 epochs) |
| CNN | Embedding dim 128 trained from scratch, filters 3/4/5 × 100, max-pool, dropout 0.5, lr , unchanged val-F1-based adaptive budget policy |
| Classical baselines | |
| TF-IDF | 1–2 g, max 50,000 features, min_df 3, max_df 0.95, sublinear tf, first 5000 characters |
| Logistic regression/SVM | , class_weight balanced |
| Naive Bayes | Multinomial, |
| Random forest | 300 trees, min_samples_leaf 2, class_weight balanced, 13 features; refit per seed |
| Environment | |
| Seeds | 42, 43, 44, 45, 46 |
| Hardware | NVIDIA Tesla T4 |
| Software | Python 3.13.15, PyTorch 2.11.0+cu128, Transformers 5.16.1, scikit-learn 1.6.1, NumPy 2.1.3, pandas 2.2.3, SciPy 1.16.3 |
| Comparison | Split | F1 | 95% CI | Recall | McNemar p |
|---|---|---|---|---|---|
| Fusion − ablated head | A | +0.0014 | [0.0005, 0.0023] | +0.0030 | 0.0037 |
| Fusion − text-only | A | +0.0001 | [−0.0012, 0.0012] | +0.0045 | 1.000 |
| Fusion − ablated head | B | +0.0078 | [0.0036, 0.0120] | +0.0107 | <0.001 |
| Fusion − text-only | B | +0.0044 | [−0.0001, 0.0092] | +0.0067 | 0.065 |
| Model | F1 (A) | F1 (B) | Gap | Recall (B) | AUC (B) |
|---|---|---|---|---|---|
| Structural only (RF) | 0.9256 | 0.9195 | 0.0061 | 0.9067 | 0.9743 |
| Structural + lexical (LR) | 0.9803 | 0.9510 | 0.0293 | 0.9292 | 0.9925 |
| Lexical only (TF-IDF + LR) | 0.9797 | 0.9439 | 0.0358 | 0.9176 | 0.9924 |
| DistilBERT + Structural (Fusion) | 0.9878 | 0.9488 | 0.0390 | 0.9149 | 0.9844 |
| DistilBERT (text-only) | 0.9877 | 0.9444 | 0.0433 | 0.9082 | 0.9904 |
| DistilBERT embeddings + MLP | 0.9864 | 0.9410 | 0.0454 | 0.9042 | 0.9934 |
| CNN (word-embedding, from scratch) | 0.9766 | 0.9189 | 0.0577 | 0.8702 | 0.9907 |
| SVM (TF-IDF, linear) | 0.9857 | 0.9188 | 0.0669 | 0.8655 | 0.9910 |
| Naive Bayes (TF-IDF) | 0.9613 | 0.8641 | 0.0972 | 0.7739 | 0.9862 |
| Model | Split | 42 | 43 | 44 | 45 | 46 | Mean | SD |
|---|---|---|---|---|---|---|---|---|
| Structural only (RF) | A | 0.9256 | 0.9264 | 0.9261 | 0.9256 | 0.9268 | 0.9261 | 0.0005 |
| Structural only (RF) | B | 0.9195 | 0.9191 | 0.9190 | 0.9172 | 0.9194 | 0.9188 | 0.0009 |
| Naive Bayes † | A | 0.9613 | 0.9613 | 0.9613 | 0.9613 | 0.9613 | 0.9613 | 0.000 |
| Naive Bayes † | B | 0.8641 | 0.8641 | 0.8641 | 0.8641 | 0.8641 | 0.8641 | 0.000 |
| SVM † | A | 0.9857 | 0.9857 | 0.9857 | 0.9857 | 0.9857 | 0.9857 | 0.000 |
| SVM † | B | 0.9188 | 0.9188 | 0.9188 | 0.9188 | 0.9188 | 0.9188 | 0.000 |
| Lexical only † | A | 0.9797 | 0.9797 | 0.9797 | 0.9797 | 0.9797 | 0.9797 | 0.000 |
| Lexical only † | B | 0.9439 | 0.9439 | 0.9439 | 0.9439 | 0.9439 | 0.9439 | 0.000 |
| Structural + lexical † | A | 0.9803 | 0.9803 | 0.9803 | 0.9803 | 0.9803 | 0.9803 | 0.000 |
| Structural + lexical † | B | 0.9510 | 0.9510 | 0.9510 | 0.9510 | 0.9510 | 0.9510 | 0.000 |
| CNN † | A | 0.9766 | 0.9766 | 0.9766 | 0.9766 | 0.9766 | 0.9766 | 0.000 |
| CNN † | B | 0.9189 | 0.9189 | 0.9189 | 0.9189 | 0.9189 | 0.9189 | 0.000 |
| DistilBERT (text-only) † | A | 0.9877 | 0.9877 | 0.9877 | 0.9877 | 0.9877 | 0.9877 | 0.000 |
| DistilBERT (text-only) † | B | 0.9444 | 0.9444 | 0.9444 | 0.9444 | 0.9444 | 0.9444 | 0.000 |
| DistilBERT + MLP | A | 0.9864 | 0.9867 | 0.9864 | 0.9863 | 0.9866 | 0.9865 | 0.0002 |
| DistilBERT + MLP | B | 0.9410 | 0.9433 | 0.9407 | 0.9507 | 0.9471 | 0.9445 | 0.0043 |
| Fusion | A | 0.9878 | 0.9875 | 0.9881 | 0.9885 | 0.9876 | 0.9879 | 0.0004 |
| Fusion | B | 0.9488 | 0.9424 | 0.9513 | 0.9323 | 0.9482 | 0.9446 | 0.0076 |
| Language | n | Legit. | Phish. | Det. Rate | Det. Rate | Majority |
|---|---|---|---|---|---|---|
| (Text-Only) | (Fusion) | Baseline | ||||
| English | 18,757 | 11,180 | 7577 | 0.9817 | 0.9864 | 0.5960 |
| German | 285 | 12 | 273 | 0.9853 | 0.9927 | 0.9579 |
| Portuguese | 139 | 2 | 137 | 1.0000 | 1.0000 | 0.9856 |
| Russian | 138 | 1 | 137 | 1.0000 | 1.0000 | 0.9928 |
| French | 95 | 8 | 87 | 0.9885 | 1.0000 | 0.9158 |
| Japanese | 83 | 2 | 81 | 1.0000 | 1.0000 | 0.9759 |
| Spanish | 79 | 3 | 76 | 1.0000 | 1.0000 | 0.9620 |
| Chinese | 74 | 1 | 73 | 1.0000 | 1.0000 | 0.9865 |
| Dutch | 46 | 0 | 46 | 1.0000 | 1.0000 | 1.0000 |
| Polish | 27 | 0 | 27 | 1.0000 | 1.0000 | 1.0000 |
| Ukrainian | 23 | 2 | 21 | 1.0000 | 1.0000 | 0.9130 |
| Italian | 21 | 1 | 20 | 1.0000 | 1.0000 | 0.9524 |
| Other (pooled) | 175 | 7 | 168 | 0.9880 | 0.9880 | 0.9600 |
| Study | Method | Corpus | Acc. | Explain. | Cross-Source |
|---|---|---|---|---|---|
| Jamal & Wimmer [13] | Fine-tuned BERT | Custom | 0.9892 | Attention | No |
| Asliyuksek et al. [10] | DistilBERT + structural | Combined (81,586) | 0.9962 | None | No |
| Kontsewaya et al. [23] | SVM, NB, LR | Multiple | ∼0.990 | None | No |
| Ubale & Shirsath [32] | SpamNet (multimodal DL) | Balanced (6000) | 0.9881 | None | No |
| Shirvani et al. [14] | Zero-shot FLAN-T5 | Custom | n/r | Partial | No |
| This paper | DistilBERT + MLP | MEPC (99,707) | 0.9894 | SHAP + LIME | Yes |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Xhafaj, G.; Chaudhry, U.B.; Jahankhani, H. Hybrid Transformer Architecture for Context-Aware Spam Email Classification. Future Internet 2026, 18, 542. https://doi.org/10.3390/fi18100542
Xhafaj G, Chaudhry UB, Jahankhani H. Hybrid Transformer Architecture for Context-Aware Spam Email Classification. Future Internet. 2026; 18(10):542. https://doi.org/10.3390/fi18100542
Chicago/Turabian StyleXhafaj, Geldi, Umair B. Chaudhry, and Hamid Jahankhani. 2026. "Hybrid Transformer Architecture for Context-Aware Spam Email Classification" Future Internet 18, no. 10: 542. https://doi.org/10.3390/fi18100542
APA StyleXhafaj, G., Chaudhry, U. B., & Jahankhani, H. (2026). Hybrid Transformer Architecture for Context-Aware Spam Email Classification. Future Internet, 18(10), 542. https://doi.org/10.3390/fi18100542

