TERA: A Trade-Off Evaluation and Resource-Aware Framework for Spam and Phishing Email Detection
Abstract
1. Introduction
- We formulate phishing detection model selection as a constraint-aware decision problem under deployment conditions, explicitly separating feasibility requirements from preference-based trade-offs.
- We introduce a multi-dimensional evaluation framework that jointly represents detection effectiveness, latency, and resource usage without reducing them to a single scalar objective.
- We demonstrate, through empirical evaluation, that models with comparable accuracy can exhibit substantial differences in operational cost, highlighting the importance of deployment-aware model selection.
- RQ1: How do classical and transformer-based models compare in detection effectiveness under identical conditions?
- RQ2: How do latency, memory usage, and model footprint influence model selection among similarly accurate candidates?
- RQ3: How does a constraint-aware evaluation framework improve model selection compared to accuracy-based ranking?
2. Background and Related Works
2.1. Email Spam and Phishing as an Evolving Threat
2.2. Machine Learning Approaches for Phishing Detection
2.3. Accuracy-Centric Evaluation Practices
2.4. Transformer-Based and Resource-Efficient Detection Models
2.5. Multi-Objective Evaluation and Research Gap
3. TERA Framework
3.1. Layer 1: Constraint Formulation
3.2. Layer 2: Model Evaluation Space
3.3. Layer 3: Decision and Deployment Mapping
- 1.
- Feasibility enforcement: Restrict consideration to models in .
- 2.
- Preference-based selection: Rank and select among feasible models according to deployment-specific priorities.
3.4. Practical Interpretation and Framework Characteristics
4. Research Methodology
4.1. Operationalization of the TERA Framework
4.2. System Architecture and Execution Model
4.2.1. Unified Inference Pipeline
4.2.2. Model Integration and Processing Flow
4.2.3. Feature Representation Strategies
4.2.4. TERA-Based Model Selection Algorithm
| Algorithm 1 TERA-Based Constraint-Aware Model Selection |
| Require: Candidate model set ; latency constraint ; resource constraint ; optional preference weights Ensure: Selected model set
|
4.3. Datasets and Preprocessing
4.3.1. Dataset Construction and Consolidation
4.3.2. Preprocessing Pipeline
4.3.3. Dataset Characteristics
4.3.4. Dataset Representativeness and Limitations
4.4. Experimental Setup
4.4.1. Execution Environment
4.4.2. Deployment Constraints
4.5. Evaluated Models
- Classical machine learning models are included as efficiency-oriented baselines. These models operate on sparse lexical representations and are characterized by low inference latency and minimal resource usage. Representative algorithms include Multi-Layer Perceptron (MLP), as well as ensemble-based methods such as XGBoost and LightGBM. These models reflect commonly deployed solutions in high-throughput and resource- constrained environments.
- Transformer-based models are incorporated to capture advanced semantic and contextual representations. Models from the BERT family, including both standard and lightweight variants, are used to evaluate the impact of increased representational capacity on detection effectiveness. These models typically incur higher computational cost but offer improved robustness against context-aware and linguistically sophisticated phishing attacks.
4.6. Evaluation Metrics
4.6.1. Predictive Performance Metrics
- Accuracy, defined aswhere , , , and denote true positives, true negatives, false positives, and false negatives, respectively.
- Precision, defined as
- Recall, defined as
- F1-score, defined as
4.6.2. Operational Metrics
- Inference Latency: the end-to-end processing time required to classify a single email, including preprocessing, feature encoding, and model inference. Given N samples, average latency is computed aswhere is the processing time of the i-th instance.
- Peak Memory Usage: the maximum memory consumption observed during inference, representing worst-case resource demand.
- Model Footprint: the serialized size of the model, reflecting storage requirements relevant to deployment and update scenarios.
4.6.3. Role in the TERA Framework
4.7. Deployment-Aware Evaluation Protocol
5. Results and Analysis
5.1. Baseline Detection Performance
5.2. Deployment Cost and Trade-Off Analysis
5.3. TERA-Based Model Selection and Comparison
5.4. Robustness and Sensitivity Analysis
5.5. Discussion and Deployment Implications
6. Discussion
6.1. From Predictive Equivalence to Deployment Differentiation
6.2. Implications for Constraint-Aware Model Selection
6.3. Revisiting Model Complexity and Representation Trade-Offs
6.4. Generalization and Broader Implications
6.5. Limitations and Future Work
6.6. Implications for Deployment-Aware Evaluation
6.7. Ablation Perspective on Evaluation Dimensions
7. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Correction Statement
References
- Bhowmick, A.; Hazarika, S.M. E-Mail Spam Filtering: A Review of Techniques and Trends. In Advances in Electronics, Communication and Computing; Kalam, A., Das, S., Sharma, K., Eds.; Lecture Notes in Electrical Engineering; Springer: Singapore, 2018; Volume 443, pp. 583–590. [Google Scholar] [CrossRef]
- Kumar Birthriya, S.; Jain, A.K. A comprehensive survey of phishing email detection and protection techniques. Inf. Secur. J. A Glob. Perspect. 2022, 31, 411–440. [Google Scholar] [CrossRef]
- Patra, C.; Giri, D.; Nandi, S.; Das, A.K.; Alenazi, M.J. Phishing email detection using vector similarity search leveraging transformer-based word embedding. Comput. Electr. Eng. 2025, 124, 110403. [Google Scholar] [CrossRef]
- Verizon. 2023 Data Breach Investigations Report; Technical Report; Verizon: Basking Ridge, NJ, USA, 2023. [Google Scholar]
- Hasanov, I.; Virtanen, S.; Hakkala, A.; Isoaho, J. Application of Large Language Models in Cybersecurity: A Systematic Literature Review. IEEE Access 2024, 12, 176751–176778. [Google Scholar] [CrossRef]
- Tooher, P.; Lallie, H.S. A Two-Stage Deep Learning Framework for AI-Driven Phishing Email Detection Based on Persuasion Principles. Computers 2025, 14, 523. [Google Scholar] [CrossRef]
- Salloum, S.A.; Gaber, T.; Vadera, S.; Shaalan, K. A Systematic Literature Review on Phishing Email Detection Using Natural Language Processing Techniques. IEEE Access 2022, 10, 65703–65727. [Google Scholar] [CrossRef]
- Hosseinzadeh, M.; Ali, U.; Ali, S.; Abbaszadi, R.; Gharehchopogh, F.S. Improving phishing email detection performance through deep learning with adaptive optimization. Sci. Rep. 2025, 15, 20668. [Google Scholar] [CrossRef] [PubMed]
- Chakraborty, P.; Arumugam, K.K.; Alfadel, M.; Nagappan, M.; McIntosh, S. Revisiting the Performance of Deep Learning-Based Vulnerability Detection on Realistic Datasets. IEEE Trans. Softw. Eng. 2024, 50, 2163–2177. [Google Scholar] [CrossRef]
- Sharma, S.; Kumar, V.; Dutta, K. Multi-objective optimization algorithms for intrusion detection in IoT networks: A systematic review. Internet Things Cyber-Phys. Syst. 2024, 4, 258–267. [Google Scholar] [CrossRef]
- Abedin, N.F.; Bawm, R.; Sarwar, T.; Saifuddin, M.; Rahman, M.A.; Hossain, S. Phishing Attack Detection using Machine Learning Classification Techniques. In Proceedings of the 2020 3rd International Conference on Intelligent Sustainable Systems (ICISS); IEEE: New York City, NY, USA, 2020; pp. 1125–1130. [Google Scholar] [CrossRef]
- Brioua, H.; Siyambaş, H.; Şahin, D.Ö. Phishing E-mail Detection with Machine Learning and Deep Learning: Improving Classification Performance with Proposed New Features. Bulg. J. Electr. Comput. Eng. 2025, 13, 183–193. [Google Scholar] [CrossRef]
- Altwaijry, N.; Al-Turaiki, I.; Alotaibi, R.; Alakeel, F. Advancing Phishing Email Detection: A Comparative Study of Deep Learning Models. Sensors 2024, 24, 2077. [Google Scholar] [CrossRef] [PubMed]
- Arshad, J.; Azad, M.A.; Abdeltaif, M.M.; Salah, K. An intrusion detection framework for energy constrained IoT devices. Mech. Syst. Signal Process. 2020, 136, 106436. [Google Scholar] [CrossRef]
- Maseer, Z.K.; Robiah, Y.; Bahaman, N.; Mostafa, S.A. Benchmarking of Machine Learning for Anomaly Based Intrusion Detection Systems in the CICIDS2017 Dataset. IEEE Access 2021, 9, 3056614. [Google Scholar] [CrossRef]
- Sommer, R.; Paxson, V. Outside the Closed World: On Using Machine Learning for Network Intrusion Detection. In Proceedings of the 2010 IEEE Symposium on Security and Privacy; IEEE: New York City, NY, USA, 2010; pp. 305–316. [Google Scholar] [CrossRef]
- Sfaxi, H.; Lahyani, I.; Yangui, S.; Torjmen, M. Latency-Aware and Proactive Service Placement for Edge Computing. IEEE Trans. Netw. Serv. Manag. 2024, 21, 4243–4254. [Google Scholar] [CrossRef]
- Noor, K.; Imoize, A.L.; Li, C.T.; Weng, C.Y. A Review of Machine Learning and Transfer Learning Strategies for Intrusion Detection Systems in 5G and Beyond. Mathematics 2025, 13, 1088. [Google Scholar] [CrossRef]
- Otieno, D.O.; Siami Namin, A.; Jones, K.S. The Application of the BERT Transformer Model for Phishing Email Classification. In Proceedings of the 2023 IEEE 47th Annual Computers, Software, and Applications Conference (COMPSAC); IEEE: New York City, NY, USA, 2023; pp. 1303–1310. [Google Scholar] [CrossRef]
- Ayodele, T.O. Impact of AI-Generated Phishing Attacks: A New Cybersecurity Threat. In Intelligent Computing. CompCom 2025; Arai, K., Ed.; Lecture Notes in Networks and Systems; Springer: Cham, Switzerland, 2025; Volume 1424. [Google Scholar] [CrossRef]
- Strubell, E.; Ganesh, A.; McCallum, A. Energy and Policy Considerations for Modern Deep Learning Research. Proc. AAAI Conf. Artif. Intell. 2020, 34, 13693–13696. [Google Scholar] [CrossRef]
- Sanh, V.; Debut, L.; Chaumond, J.; Wolf, T. DistilBERT, a Distilled Version of BERT: Smaller, Faster, Cheaper and Lighter. In Proceedings of the 5th Workshop on Energy Efficient Machine Learning and Cognitive Computing, Vancouver, BC, Canada, 13 December 2019. [Google Scholar] [CrossRef]
- Jiao, X.; Yin, Y.; Shang, L.; Jiang, X.; Chen, X.; Li, L.; Wang, F.; Liu, Q. TinyBERT: Distilling BERT for Natural Language Understanding. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP; IEEE: New York City, NY, USA, 2020. [Google Scholar] [CrossRef]
- Lan, Z.; Chen, M.; Goodman, S.; Gimpel, K.; Sharma, P.; Soricut, R. ALBERT: A Lite BERT for Self-Supervised Learning of Language Representations. In Proceedings of the International Conference on Learning Representations (ICLR); IEEE: New York City, NY, USA, 2020. [Google Scholar] [CrossRef]
- Milliken, M.; Bi, Y.; Galway, L.; Hawe, G. Multi-objective optimization of base classifiers in StackingC by NSGA-II for intrusion detection. In Proceedings of the 2016 IEEE Symposium Series on Computational Intelligence (SSCI); IEEE: New York City, NY, USA, 2016; pp. 1–8. [Google Scholar] [CrossRef]
- Telikani, A.; Rudbardeh, N.E.; Soleymanpour, S.; Shahbahrami, A.; Shen, J.; Gaydadjiev, G.; Hassanpour, R. A Cost-Sensitive Machine Learning Model with Multitask Learning for Intrusion Detection in IoT. IEEE Trans. Ind. Inform. 2024, 20, 3880–3890. [Google Scholar] [CrossRef]
- Abu Al-Haija, Q.; Al-Fayoumi, M. An intelligent identification and classification system for malicious uniform resource locators (URLs). Neural Comput. Appl. 2023, 35, 16995–17011. [Google Scholar] [CrossRef] [PubMed]
- Abu Al-Haija, Q.; Al Badawi, A. URL-based Phishing Websites Detection via Machine Learning. In Proceedings of the 2021 International Conference on Data Analytics for Business and Industry (ICDABI); IEEE: New York City, NY, USA, 2021; pp. 644–649. [Google Scholar] [CrossRef]
- Al-Fayoumi, M.; Alhijawi, B.; Abu Al-Haija, Q.; Armoush, R. XAI-PhD: Fortifying Trust of Phishing URL Detection Empowered by Shapley Additive Explanations. Int. J. Online Biomed. Eng. (iJOE) 2024, 20, 80–101. [Google Scholar] [CrossRef]
- Elqasass, A.; Aljundi, I.; Al-Fayoumi, M.; Abu Al-Haija, Q. Facilitating Secure Web Browsing by Utilizing Supervised Filtration of Malicious URLs. In IoT Based Control Networks and Intelligent Systems; Joby, P.P., Alencar, M.S., Falkowski-Gilski, P., Eds.; Lecture Notes in Networks and Systems; Springer: Singapore, 2024; Volume 789. [Google Scholar] [CrossRef]
- Abu-Nimeh, S.; Nappa, D.; Wang, X.; Nair, S. A Comparison of Machine Learning Techniques for Phishing Detection. In Proceedings of the Anti-Phishing Working Groups 2nd Annual eCrime Researchers Summit, Pittsburgh, PA, USA, 4–5 October 2007; pp. 60–69. [Google Scholar] [CrossRef]
- Al-Subaiey, A.; Al-Thani, M.; Alam, N.A.; Antora, K.F.; Khandakar, A.; Zaman, S.A.U. Novel Interpretable and Robust Web-based AI Platform for Phishing Email Detection. arXiv 2024, arXiv:2405.11619. [Google Scholar] [CrossRef]
- Raschka, S.; Mirjalili, V. Machine Learning and Deep Learning with Python; Packt Publishing: Birmingham, UK, 2018. [Google Scholar]
- Probst, P.; Wright, M.N.; Boulesteix, A.L. Tunability: Importance of Hyperparameters of Machine Learning Algorithms. J. Mach. Learn. Res. 2019, 20, 1–32. [Google Scholar]
- Olson, R.S.; Bartley, N.; Urbanowicz, R.J.; Moore, J.H. Evaluation of a tree-based pipeline optimization tool for automating data science. In Proceedings of the Genetic and Evolutionary Computation Conference, New York, NY, USA, 20–24 July 2016; pp. 485–492. [Google Scholar]






| Study | Year | Domain | Multi-Dimensional Focus | Key Limitation | RQ1 | RQ2 | RQ3 |
|---|---|---|---|---|---|---|---|
| [25] | 2020 | Intrusion Detection | Accuracy–cost trade-off via multi-objective optimization | Treats factors as jointly tradeable; lacks explicit feasibility constraints | ✓ | ✓ | × |
| [19] | 2021 | Phishing Detection | Transformer-based semantic robustness | Focuses on representation; ignores operational feasibility | ✓ | × | × |
| [26] | 2021 | Security ML | Cost-sensitive learning objective | Combines objectives; no separation of feasibility and optimization | ✓ | ✓ | × |
| [9] | 2021 | Security ML Evaluation | Dataset-driven evaluation | Highlights benchmarking issues; lacks decision support | ✓ | × | × |
| [17] | 2022 | Edge Computing Security | Latency-aware evaluation | Reports latency but not integrated into decision | ✓ | ✓ | × |
| [28] | 2022 | Intelligent Systems | Performance-oriented evaluation | Focus on performance; lacks decision formulation | ✓ | × | × |
| [14] | 2023 | Edge-based IDS | Resource-aware deployment constraints | Limited to specific models; lacks cross-model framework | × | ✓ | × |
| [27] | 2023 | ML Optimization | Multi-objective optimization (accuracy, cost, efficiency) | Joint optimization; no feasibility separation | ✓ | ✓ | × |
| [31] | 2023 | Intelligent Systems | Accuracy-driven evaluation | Evaluation-centric; no constraint-aware selection | ✓ | × | × |
| [13] | 2024 | Phishing Detection | Comparative deep learning evaluation | Remains accuracy-centric; ignores constraints | ✓ | × | × |
| [30] | 2024 | AI Applications | Model-centric optimization | No constrained decision formulation | ✓ | × | × |
| [29] | 2024 | Intelligent Systems | Multi-metric system evaluation | Descriptive metrics; no structured decision process | ✓ | ✓ | × |
| TERA | 2026 | Deployment-aware ML | Constraint-aware multi-dimensional evaluation | Separates feasibility from preference; structured decision formulation | ✓ | ✓ | ✓ |
| Weight | Interpretation | Deployment Implication |
|---|---|---|
| (Accuracy) | Emphasizes detection effectiveness | Favors models with higher predictive performance; suitable for high-security scenarios |
| (Latency) | Penalizes inference delay | Prioritizes low-latency models; suitable for real-time systems |
| (Resource) | Penalizes memory and model size | Prefers lightweight models; suitable for resource-constrained environments |
| Class | Records | Mean | Median | 90th Percentile |
|---|---|---|---|---|
| Ham | 11,322 (48.18%) | 245.7 | 168 | 555 |
| Phishing | 7328 (31.19%) | 206.6 | 142 | 538 |
| Spam | 4848 (20.63%) | 123.0 | 58 | 371 |
| Feature | Model | Accuracy | Precision | Recall | F1-Score |
|---|---|---|---|---|---|
| TF-IDF | LightGBM | 0.9681 | 0.9685 | 0.9681 | 0.9682 |
| TF-IDF | XGBoost | 0.9598 | 0.9601 | 0.9598 | 0.9599 |
| MPNet | MLP | 0.9560 | 0.9565 | 0.9560 | 0.9561 |
| TF-IDF | MLP | 0.9438 | 0.9450 | 0.9438 | 0.9442 |
| MPNet | LightGBM | 0.9238 | 0.9256 | 0.9238 | 0.9241 |
| MiniLM-L12 | MLP | 0.9194 | 0.9205 | 0.9194 | 0.9198 |
| MPNet | XGBoost | 0.9181 | 0.9194 | 0.9181 | 0.9180 |
| MiniLM-L6 | MLP | 0.9006 | 0.9024 | 0.9006 | 0.9012 |
| MiniLM-L12 | LightGBM | 0.8885 | 0.8898 | 0.8885 | 0.8883 |
| MiniLM-L12 | XGBoost | 0.8783 | 0.8791 | 0.8783 | 0.8772 |
| Model Pair | Metric | 95% CI | |
|---|---|---|---|
| TF-IDF + LightGBM vs. MPNet + MLP | F1 | 0.0121 | [−0.0065, 0.0182] |
| TF-IDF + LightGBM vs. TF-IDF + XGBoost | F1 | 0.0083 | [−0.0051, 0.0157] |
| MPNet + MLP vs. TF-IDF + XGBoost | F1 | −0.0038 | [−0.0102, 0.0069] |
| Feature | Model | Acc ↑ | F1 ↑ | P95 Latency (s) ↓ | Peak Memory (MB) ↓ | Model Size (MB) ↓ | Score ↑ | Deployment Insight |
|---|---|---|---|---|---|---|---|---|
| TF-IDF | XGB | 0.9615 | 0.9599 | 0.0365 | 0.0103 | 1.74 | 0.9130 | Best overall trade-off under deployment constraints |
| MPNet | MLP | 0.9580 | 0.9561 | 0.0513 | 0.1575 | 1.51 | 0.8710 | Balanced alternative with moderate resource requirements |
| TF-IDF | MLP | 0.9465 | 0.9442 | 0.0189 | 0.3138 | 1.95 | 0.8520 | Efficient in latency but with reduced detection performance |
| MPNet | XGB | 0.9203 | 0.9180 | 0.0108 | 0.0102 | 2.79 | 0.8420 | Highly resource-efficient but lowest detection effectiveness |
| TF-IDF | LGBM | 0.9701 | 0.9682 | 1.7196 | 22.01 | 4.91 | 0.5000 | Highest predictive performance but dominated by latency and memory cost |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Jandaeng, C.; Koad, P.; Zolkipli, M.F.; Phuttharak, J. TERA: A Trade-Off Evaluation and Resource-Aware Framework for Spam and Phishing Email Detection. Informatics 2026, 13, 72. https://doi.org/10.3390/informatics13050072
Jandaeng C, Koad P, Zolkipli MF, Phuttharak J. TERA: A Trade-Off Evaluation and Resource-Aware Framework for Spam and Phishing Email Detection. Informatics. 2026; 13(5):72. https://doi.org/10.3390/informatics13050072
Chicago/Turabian StyleJandaeng, Chanankorn, Peeravit Koad, Mohamad Fadli Zolkipli, and Jurairat Phuttharak. 2026. "TERA: A Trade-Off Evaluation and Resource-Aware Framework for Spam and Phishing Email Detection" Informatics 13, no. 5: 72. https://doi.org/10.3390/informatics13050072
APA StyleJandaeng, C., Koad, P., Zolkipli, M. F., & Phuttharak, J. (2026). TERA: A Trade-Off Evaluation and Resource-Aware Framework for Spam and Phishing Email Detection. Informatics, 13(5), 72. https://doi.org/10.3390/informatics13050072

