Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (50)

Search Parameters:
Keywords = spam email

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
30 pages, 3682 KB  
Article
Toward Equitable Arabic Cybersecurity Literacy: A Rubric-Constrained LLM Framework for Phishing Detection and Bilingual Translation Fidelity
by Taher M. Ghazal, Fareeha Anwar, Sumaia Mohammed Al-Ghuribi, Amjed A. Ahmed, Ali Hamzah Najim, Omar Almomani, Prabu Pachiyannan and Hesham A. Sakr
Math. Comput. Appl. 2026, 31(5), 168; https://doi.org/10.3390/mca31050168 - 23 Aug 2026
Viewed by 121
Abstract
Arabic-speaking populations face disproportionate cybersecurity risks due to the predominantly English-centric design of existing awareness materials, which fail to accommodate Arabic dialectal diversity, script complexity, and culturally embedded communication patterns. These deficiencies impair users’ ability to interpret phishing messages, authentication requests, and security [...] Read more.
Arabic-speaking populations face disproportionate cybersecurity risks due to the predominantly English-centric design of existing awareness materials, which fail to accommodate Arabic dialectal diversity, script complexity, and culturally embedded communication patterns. These deficiencies impair users’ ability to interpret phishing messages, authentication requests, and security alerts, increasing susceptibility to social engineering, identity theft, and data breaches. This paper presents SECURE-A2RC, a rubric-constrained, Arabic-aware large language model framework designed to deliver scalable, interpretable, and culturally relevant cybersecurity education. The framework comprises two coupled components. The first, the Arabic-Aware Secure Communication Encoder (A-SCE), employs an instruction-tuned LLM to produce multidimensional encodings that capture three learner competencies: security intent comprehension; linguistic deception cue recognition encompassing urgency, authority impersonation, and incentive framing; and action-critical translation fidelity across Arabic dialectal registers and Arabic–English bilingual contexts. The second, the Rubric-Constrained Adaptive Feedback Generator (RCAFG), translates A-SCE encodings into personalized, expert-aligned instructional feedback and proficiency-calibrated adaptive tasks, ensuring pedagogical consistency, security correctness, and dialect awareness throughout the learning cycle. The framework is evaluated on three domain-relevant corpora: the English–Arabic Parallel Phishing Email Corpus, the Open MalSec dataset, and the Arabic Spam and Ham Tweets dataset. SECURE-A2RC achieves a 31% improvement in phishing identification accuracy and a 26% reduction in action-critical translation errors compared to conventional awareness materials. A comparative evaluation against SERENA, a Multi-Agent LLM, and the Arabic Multitask Learning Model confirms consistent superiority across detection accuracy, F1-score, dialectal robustness, and educational effectiveness metrics, affirming rubric-constrained LLM integration as a viable approach to equitable multilingual cybersecurity education. Full article
Show Figures

Figure 1

26 pages, 1349 KB  
Article
ML-Based SMS Messaging Spam Detection: Impacts of Text Feature Extraction Techniques
by Ahmad Ababneh and Maram Bani Younes
J. Cybersecur. Priv. 2026, 6(4), 125; https://doi.org/10.3390/jcp6040125 - 18 Jul 2026
Viewed by 397
Abstract
Spam detection on SMS messaging has not received as much attention from researchers recently as the spam detection studies on emails or social media platforms. However, spam SMS messaging can be more intrusive, annoying, and harmful. Thus, detecting and filtering spam SMS messages [...] Read more.
Spam detection on SMS messaging has not received as much attention from researchers recently as the spam detection studies on emails or social media platforms. However, spam SMS messaging can be more intrusive, annoying, and harmful. Thus, detecting and filtering spam SMS messages is becoming a priority that saves human productivity. This work aims to introduce a dynamic, accurate, and efficient machine learning-based spam detection technique for SMS messaging. It aims at protecting users and businesses from spam SMS attacks. It aims to detect and identify suspicious messages that contain promotional, misleading, irrelevant, or harmful content. It primarily aims to test and evaluate the impact of feature extraction methods on the performance of machine-learning-based spam detection. Several text feature extraction techniques have been used and tested, including classical, statistical, contextual, and advanced embedding techniques. An extensive set of experiments has been presented on benchmark datasets in this field. From the comparative study, we can infer that all investigated feature extraction techniques have achieved high accuracy (90%+) on the in-domain dataset. However, their performance decreased when they were tested on the out-of-domain dataset (70%+). The advanced embedding techniques achieved the best performance across both datasets compared to the other tested feature extraction models. Full article
(This article belongs to the Section Security Engineering & Applications)
Show Figures

Figure 1

54 pages, 9796 KB  
Article
Multimodal Zone-Aware Graph-Based Transformer with Continual Learning and Bio-Inspired Optimization for Email Spam Detection
by Neomi Nelin Nicholas and V. Nirmalrani
Appl. Sci. 2026, 16(14), 7107; https://doi.org/10.3390/app16147107 - 15 Jul 2026
Viewed by 261
Abstract
Cyberattacks via email remain a major menace to people, companies, and critical infrastructures, and effective spam and phishing detection is a social concern. Nevertheless, the current methods, such as NetSpam and SMART, tend to have issues with non-homogenous data streams, lack of contextual [...] Read more.
Cyberattacks via email remain a major menace to people, companies, and critical infrastructures, and effective spam and phishing detection is a social concern. Nevertheless, the current methods, such as NetSpam and SMART, tend to have issues with non-homogenous data streams, lack of contextual knowledge, poor generalization, and inability to adapt to changing attack patterns. The existing techniques are not strong in terms of multimodal fusion and cannot effectively transfer trust or update risk scores in dynamic conditions. To overcome these shortcomings, this paper presents a Zone-Aware Multimodal Graph-Based Transformer that combines text, image, video, and metadata streams in a smooth manner to detect threats in emails. The three key novelties of the proposed framework include AAGFusion to contrastively align multimodal features and hierarchically fuse them using transformers; MAGNN-SASO to classify zones, compute similarity across zones, and optimize bio-inspired optimization; and Q-BayesTrustNet-X to propagate trust, risk score, Bayesian calibration, and continual learning, and provide interpretable feedback by using LRP-based explainability. The experimental findings prove that the proposed system has a high level of performance, with the accuracy, precision, and specificity reaching 98.57, 97.51, and 99.48, respectively, which proves its efficiency in high-fidelity and real-world spam and phishing detection. Full article
(This article belongs to the Section Computing and Artificial Intelligence)
Show Figures

Figure 1

46 pages, 2150 KB  
Article
The Fragility of Phishing Detection Models: Evidence from Cross-Corpus Transfer, Prevalence Shift, Artifact Learning, and Evasion Risk
by Istiaque Bhuiyan and Tanvir Bhuiyan
Big Data Cogn. Comput. 2026, 10(7), 211; https://doi.org/10.3390/bdcc10070211 - 29 Jun 2026
Viewed by 608
Abstract
Phishing detection models often report strong benchmark performance, yet their reliability under realistic deployment conditions remains uncertain. This study examines this problem by investigating three failure modes of cross-dataset phishing email detection: corpus generalization failure, asymmetric prevalence-shift failure, and artifact-driven spurious learning. Using [...] Read more.
Phishing detection models often report strong benchmark performance, yet their reliability under realistic deployment conditions remains uncertain. This study examines this problem by investigating three failure modes of cross-dataset phishing email detection: corpus generalization failure, asymmetric prevalence-shift failure, and artifact-driven spurious learning. Using six public email corpora, CEAS_08, Enron, Ling, Nazario, Nigerian Fraud, and SpamAssassin, the study evaluates Term Frequency (TF) and Inverse Document Frequency (IDF)-based Logistic Regression and Linear Support Vector Classifier (SVC) models across pooled baseline testing, single-corpus cross-dataset transfer, leave-one-corpus-out pooled training, prevalence-shift simulation, training prevalence manipulation, dataset-identification analysis, top-feature inspection, artifact-removal ablation, and targeted feature-sensitivity masking. The findings show that single-corpus models are unstable under cross-dataset transfer, with F1-scores varying substantially across source–target combinations. In contrast, leave-one-corpus-out pooled training improves robustness, with Logistic Regression achieving sustained F1-scores between 0.8201 and 0.8994, and Linear SVC achieving F1-scores between 0.7607 and 0.8910 across unseen corpora. Prevalence-shift experiments reveal that failure is asymmetric and threshold-dependent. High-prevalence-trained models maintain high recall under fixed thresholds but suffer sharp recall degradation when operational alert-budget constraints are imposed. Conversely, low-prevalence-trained models become overly conservative in high-threat environments, producing high precision but substantially lower recall and poorer calibration. Artifact analyses further show that source corpus identity is highly learnable, with dataset-identification accuracy reaching 0.9722 for Logistic Regression and 0.9806 for Linear SVC. Top-feature and masking analyses indicate that models rely partly on corpus markers, date tokens, URL/domain terms, headers, and other artifact-like features rather than only general phishing indicators. The study contributes a deployment-aware and adversary-aware evaluation framework for phishing detection. It shows that benchmark accuracy alone is insufficient for assessing real-world robustness and that reliable phishing detection requires cross-corpus validation, prevalence-aware thresholding, and systematic testing for artifact-driven spurious learning. Full article
(This article belongs to the Special Issue Big Data and Cognitive Computing in 2026)
Show Figures

Figure 1

1 pages, 129 KB  
Correction
Correction: Jandaeng et al. TERA: A Trade-Off Evaluation and Resource-Aware Framework for Spam and Phishing Email Detection. Informatics 2026, 13, 72
by Chanankorn Jandaeng, Peeravit Koad, Mohamad Fadli Zolkipli and Jurairat Phuttharak
Informatics 2026, 13(7), 101; https://doi.org/10.3390/informatics13070101 - 25 Jun 2026
Viewed by 456
Abstract
In the published publication [...] Full article
27 pages, 437 KB  
Article
Adaptive Semi-Personalized Email Classification Model (ASPEC) with Incremental Learning
by Worawit Kitikusoun and Nawaporn Wisitpongphan
Informatics 2026, 13(6), 79; https://doi.org/10.3390/informatics13060079 - 29 May 2026
Viewed by 803
Abstract
The volume of daily email traffic continues to grow rapidly, creating challenges in efficiently distinguishing important from irrelevant messages. Beyond spam detection, modern email systems classify messages into categories such as promotions, social, updates, and forums, many of which are ignored or deleted [...] Read more.
The volume of daily email traffic continues to grow rapidly, creating challenges in efficiently distinguishing important from irrelevant messages. Beyond spam detection, modern email systems classify messages into categories such as promotions, social, updates, and forums, many of which are ignored or deleted without review. To address this issue, researchers have explored intelligent classification systems to predict the importance of emails, enhance user productivity, and improve organizational communication efficiency. This study proposes an email classification model that adapts to different users’ work functions and communication patterns within an organizational context. Using three-month historical real corporate anonymized email data from 9788 individuals across 12 work functions, the proposed Adaptive Semi-Personalized Email Classification Model (ASPEC) automatically retrieves each employee’s occupational profile—including job category and years of work experience—from the organization’s Human Resources (HR) system, enabling seamless personalization without manual configuration. ASPEC significantly improves email classification accuracy over the best-performing baseline of 73.50%, with incremental learning further enabling continuous adaptation to evolving data streams and achieving accuracy up to 92.57% in stable user segments. Unlike most existing email classification frameworks, which rely on static batch-learning models and lack memory-based or incremental update mechanisms, ASPEC addresses this gap by continuously adapting to evolving communication patterns without requiring full model retraining. The adoption of this incremental learning framework offers tangible benefits for organizations, including reduced manual email filtering workload, improved communication efficiency, and decreased operational burden on IT departments in managing email-related tasks and issues. Full article
(This article belongs to the Section Machine Learning)
Show Figures

Figure 1

30 pages, 4692 KB  
Article
TERA: A Trade-Off Evaluation and Resource-Aware Framework for Spam and Phishing Email Detection
by Chanankorn Jandaeng, Peeravit Koad, Mohamad Fadli Zolkipli and Jurairat Phuttharak
Informatics 2026, 13(5), 72; https://doi.org/10.3390/informatics13050072 - 12 May 2026
Cited by 1 | Viewed by 1356 | Correction
Abstract
Email spam and phishing detection is typically evaluated using accuracy-centric metrics under implicitly unconstrained computational settings. However, in practical deployment scenarios—particularly in real-time and resource-constrained environments—models with comparable predictive performance may differ substantially in inference latency and resource usage, directly affecting their operational [...] Read more.
Email spam and phishing detection is typically evaluated using accuracy-centric metrics under implicitly unconstrained computational settings. However, in practical deployment scenarios—particularly in real-time and resource-constrained environments—models with comparable predictive performance may differ substantially in inference latency and resource usage, directly affecting their operational feasibility. This paper introduces TERA, a deployment-aware evaluation framework that formulates model assessment as a constraint-aware decision problem. Instead of aggregating performance and efficiency into a single objective, TERA treats predictive performance as a feasibility requirement that defines an admissible set of models. Within this feasible region, operational factors such as latency and resource usage are used to differentiate among candidates through structured, multi-dimensional analysis. Experiments on benchmark email datasets show that multiple models achieve comparable detection performance, forming a region of predictive equivalence. Within this region, significant variations in latency and resource consumption are observed, indicating that predictive equivalence does not imply deployment equivalence. These findings demonstrate that accuracy-based evaluation alone may provide limited guidance for deployment-oriented model selection. By explicitly separating feasibility constraints from preference-based trade-offs, TERA enables transparent and deployment-aligned model evaluation. The framework supports consistent comparison and selection among accuracy-comparable models without altering the role of detection effectiveness as a primary requirement, thereby complementing existing evaluation practices with a structured decision-oriented perspective. Full article
Show Figures

Figure 1

21 pages, 1162 KB  
Review
Machine Learning Based Spam Detection in Digital Communication Systems: A Comparative Analysis
by Maram Bani Younes and Ahmad Ababneh
Systems 2026, 14(3), 229; https://doi.org/10.3390/systems14030229 - 24 Feb 2026
Cited by 3 | Viewed by 3509
Abstract
Spam messages are unwanted, irrelevant, or potentially harmful messages sent in bulk to large numbers of recipients via email, SMS, or social media. These messages pose a threat of spam to individual users and commercial companies. They threaten digital communication platforms by enabling [...] Read more.
Spam messages are unwanted, irrelevant, or potentially harmful messages sent in bulk to large numbers of recipients via email, SMS, or social media. These messages pose a threat of spam to individual users and commercial companies. They threaten digital communication platforms by enabling phishing, malware distribution, service disruption, and unsolicited advertisements. Several mechanisms have been used in the literature to detect spam over digital communication systems. This includes rule-based filtering, Bayesian filtering, heuristic analysis, and machine learning (ML) techniques. Traditional rule-based and heuristic analyses were insufficient to cope with evolving attack patterns. Meanwhile, ML models can present modern, dynamic, appropriate, and efficient solutions in this manner. This study aims to evaluate and compare several basic ML models for spam detection, considering popular benchmark datasets on several communication platforms as a comprehensive comparative study. The experimental results demonstrate that the tested models achieve good accuracy, precision, recall, and F1-score on each investigated benchmark dataset. However, the performance of all models has decreased drastically when the trained models are tested on an unseen dataset. Recommendations for future required enhancements to handle this reduction in the performance of ML techniques for unseen datasets are provided. Finally, extra experimental tests have shown the positive impact of applying some of these recommendations. Full article
Show Figures

Figure 1

21 pages, 1311 KB  
Article
A Novel Dual-Layer Deep Learning Architecture for Phishing and Spam Email Detection
by Sarmad Rashed and Caner Ozcan
Electronics 2026, 15(3), 630; https://doi.org/10.3390/electronics15030630 - 2 Feb 2026
Cited by 2 | Viewed by 1316
Abstract
Phishing and spam emails continue to pose a serious cybersecurity threat, leading to financial loss, information leakage, and reputational damage. Traditional email filtering approaches struggle to keep pace with increasingly sophisticated attack strategies, particularly those involving malicious content and deceptive attachments. This study [...] Read more.
Phishing and spam emails continue to pose a serious cybersecurity threat, leading to financial loss, information leakage, and reputational damage. Traditional email filtering approaches struggle to keep pace with increasingly sophisticated attack strategies, particularly those involving malicious content and deceptive attachments. This study proposes a dual-layer deep learning architecture designed to enhance email security by improving the detection of phishing and spam messages. The first layer employs deep learning models, including LSTM- and transformer-based classifiers, to analyze email content and structural features across legitimate, phishing, and spam emails. The second layer focuses on spam emails containing attachments and applies advanced transformer models, such as GPT-2 and XLM-RoBERTa, to assess contextual and semantic patterns associated with malicious attachments. By integrating textual analysis with attachment-level inspection, the proposed architecture overcomes limitations of single-layer approaches that rely solely on email body content. Experimental evaluation using accuracy and F1-score demonstrates that the dual-layer framework achieves a minimum F1-score of 98.75 percent in spam–ham classification and attains an attachment detection accuracy of up to 99.46 percent. These results indicate that the proposed approach offers a reliable and scalable solution for enhancing real-world email security systems. Full article
Show Figures

Figure 1

36 pages, 8254 KB  
Article
A Comparative Evaluation of a Multimodal Approach for Spam Email Classification Using DistilBERT and Structural Features
by Halim Asliyuksek, Ozgur Tonkal and Ramazan Kocaoglu
Electronics 2025, 14(19), 3855; https://doi.org/10.3390/electronics14193855 - 29 Sep 2025
Cited by 5 | Viewed by 5386
Abstract
This study aims to improve the automatic detection of unwanted emails using advanced machine learning and deep learning methods. By reviewing current research over the past five years, a comprehensive combined dataset structure was created containing a total of 81,586 email samples from [...] Read more.
This study aims to improve the automatic detection of unwanted emails using advanced machine learning and deep learning methods. By reviewing current research over the past five years, a comprehensive combined dataset structure was created containing a total of 81,586 email samples from seven different spam datasets. Class imbalance was addressed through the application of random oversampling and class-weighted loss, and the decision threshold was subsequently tuned for deployment. Among classical machine learning solutions, Random Forest (RF) emerged as the most successful method, while deep learning approaches, such as Transformer-based models like Distilled Bidirectional Encoder Representations from Transformers (DistilBERT) and Robustly Optimized BERT Pretraining Approach (RoBERTa), demonstrated superior performance. The highest test score (99.62%) on a combined static dataset was achieved with a multimodal architecture that combines deep meaningful text representations from DistilBERT with structural text features. Beyond this static performance benchmark, the study investigates the critical challenge of concept drift by performing a temporal analysis on datasets from different eras. The results reveal a significant performance degradation in all models when tested on modern spam, highlighting a critical vulnerability of statically trained systems. Notably, the Transformer-based model demonstrated greater robustness against this temporal decay compared to traditional methods. This study offers not only an effective classification solution but also provides crucial empirical evidence on the necessity of adaptive, continually learning systems for robust spam detection. Full article
(This article belongs to the Special Issue Role of Artificial Intelligence in Natural Language Processing)
Show Figures

Figure 1

26 pages, 3423 KB  
Article
Federated Learning Spam Detection Based on FedProx and Multi-Level Multi-Feature Fusion
by Yunpeng Xiong, Junkuo Cao and Guolian Chen
Informatics 2025, 12(3), 93; https://doi.org/10.3390/informatics12030093 - 12 Sep 2025
Cited by 6 | Viewed by 3430
Abstract
Traditional spam detection methodologies often neglect user privacy preservation, potentially incurring data leakage risks. Furthermore, current federated learning models for spam detection face several critical challenges: (1) data heterogeneity and instability during server-side parameter aggregation, (2) training instability in single neural network architectures [...] Read more.
Traditional spam detection methodologies often neglect user privacy preservation, potentially incurring data leakage risks. Furthermore, current federated learning models for spam detection face several critical challenges: (1) data heterogeneity and instability during server-side parameter aggregation, (2) training instability in single neural network architectures leading to mode collapse, and (3) constrained expressive capability in multi-module frameworks due to excessive complexity. These issues represent fundamental research pain points in federated learning-based spam detection systems. To address this technical challenge, this study innovatively integrates federated learning frameworks with multi-feature fusion techniques to propose a novel spam detection model, FPW-BC. The FPW-BC model addresses data distribution imbalance through the FedProx aggregation algorithm and enhances stability during server-side parameter aggregation via a horse-racing selection strategy. The model effectively mitigates limitations inherent in both single and multi-module architectures through hierarchical multi-feature fusion. To validate FPW-BC’s performance, comprehensive experiments were conducted on six benchmark datasets with distinct distribution characteristics: CEAS, Enron, Ling, Phishing_email, Spam_email, and Fake_phishing, with comparative analysis against multiple baseline methods. Experimental results demonstrate that FPW-BC achieves exceptional generalization capability for various spam patterns while maintaining user privacy preservation. The model attained 99.40% accuracy on CEAS and 99.78% on Fake_phishing, representing significant dual improvements in both privacy protection and detection efficiency. Full article
Show Figures

Figure 1

18 pages, 1417 KB  
Article
A Fusion-Based Approach with Bayes and DeBERTa for Efficient and Robust Spam Detection
by Ao Zhang, Kelei Li and Haihua Wang
Algorithms 2025, 18(8), 515; https://doi.org/10.3390/a18080515 - 15 Aug 2025
Cited by 2 | Viewed by 2061
Abstract
Spam emails pose ongoing risks to digital security, including data breaches, privacy violations, and financial losses. Addressing the limitations of traditional detection systems in terms of accuracy, adaptability, and resilience remains a significant challenge. In this paper, we propose a hybrid spam detection [...] Read more.
Spam emails pose ongoing risks to digital security, including data breaches, privacy violations, and financial losses. Addressing the limitations of traditional detection systems in terms of accuracy, adaptability, and resilience remains a significant challenge. In this paper, we propose a hybrid spam detection framework that integrates a classical multinomial naive Bayes classifier with a pre-trained large language model, DeBERTa. The framework employs a weighted probability fusion strategy to combine the strengths of both models—lexical pattern recognition and deep semantic understanding—into a unified decision process. We evaluate the proposed method on a widely used spam dataset. Experimental results demonstrate that the hybrid model achieves superior performance in terms of accuracy and robustness when compared with other classifiers. The findings support the effectiveness of hybrid modeling in advancing spam detection techniques. Full article
(This article belongs to the Section Evolutionary Algorithms and Machine Learning)
Show Figures

Figure 1

21 pages, 804 KB  
Article
Spam Email Detection Using Long Short-Term Memory and Gated Recurrent Unit
by Samiullah Saleem, Zaheer Ul Islam, Syed Shabih Ul Hasan, Habib Akbar, Muhammad Faizan Khan and Syed Adil Ibrar
Appl. Sci. 2025, 15(13), 7407; https://doi.org/10.3390/app15137407 - 1 Jul 2025
Cited by 9 | Viewed by 4045
Abstract
In today’s business environment, emails are essential across all sectors, including finance and academia. There are two main types of emails: ham (legitimate) and spam (unsolicited). Spam wastes consumers’ time and resources and poses risks to sensitive data, with volumes doubling daily. Current [...] Read more.
In today’s business environment, emails are essential across all sectors, including finance and academia. There are two main types of emails: ham (legitimate) and spam (unsolicited). Spam wastes consumers’ time and resources and poses risks to sensitive data, with volumes doubling daily. Current spam identification methods, such as Blocklist approaches and content-based techniques, have limitations, highlighting the need for more effective solutions. These constraints call for detailed and more accurate approaches, such as machine learning (ML) and deep learning (DL), for realistic detection of new scams. Emphasis has since been placed on the possibility that ML and DL technologies are present in detecting email spam. In this work, we have succeeded in developing a hybrid deep learning model, where Long Short-Term Memory (LSTM) and the Gated Recurrent Unit (GRU) are applied distinctly to identify spam email. Despite the fact that the other models have been applied independently (CNNs, LSTM, GRU, or ensemble machine learning classifier) in previous studies, the given research has provided a contribution to the existing body of literature since it has managed to combine the advantage of LSTM in capturing the long-term dependency and the effectiveness of GRU in terms of computational efficiency. In this hybridization, we have addressed key issues such as the vanishing gradient problem and outrageous resource consumption that are usually encountered in applying standalone deep learning. Moreover, our proposed model is superior regarding the detection accuracy (90%) and AUC (98.99%). Though Transformer-based models are significantly lighter and can be used in real-time applications, they require extensive computation resources. The proposed work presents a substantive and scalable foundation to spam detection that is technically and practically dissimilar to the familiar approaches due to the powerful preprocessing steps, including particular stop-word removal, TF-IDF vectorization, and model testing on large, real-world size dataset (Enron-Spam). Additionally, delays in the feature comparison technique within the model minimize false positives and false negatives. Full article
Show Figures

Figure 1

16 pages, 2334 KB  
Article
PhiShield: An AI-Based Personalized Anti-Spam Solution with Third-Party Integration
by Hyunsol Mun, Jeeeun Park, Yeonhee Kim, Boeun Kim and Jongkil Kim
Electronics 2025, 14(8), 1581; https://doi.org/10.3390/electronics14081581 - 13 Apr 2025
Cited by 2 | Viewed by 3067
Abstract
In this paper, we present PhiShield, which is a spam filter system designed to offer real-time email collection and analysis at the end node. Before our work, most existing spam detection systems focused more on detection accuracy rather than usability and privacy. PhiShield [...] Read more.
In this paper, we present PhiShield, which is a spam filter system designed to offer real-time email collection and analysis at the end node. Before our work, most existing spam detection systems focused more on detection accuracy rather than usability and privacy. PhiShield is introduced to enhance both of these features by precisely choosing the deployment location where it achieves personalization and proactive defense. The PhiShield system is designed to allow enhanced compatibility and proactive phishing prevention for users. Phishield is implemented as a browser extension and is compatible with third-party email services such as Gmail. As it is implemented as a browser extension, it assesses emails before a user clicks on them. It offers proactive prevention for users by showing a personalized report, not the content of the phishing email, when a phishing email is detected. Therefore, it provides users with transparency surrounding phishing mechanisms and helps them mitigate phishing risks in practice. We test various locally trained Artificial Intelligence (AI)-based detection models and show that a Long Short-Term Memory (LSTM) model is suitable for practical phishing email detection (>98% accuracy rate) with a reasonable training cost. This means that an organization or user can develop their own private detection rules and supplementarily use the private rules in addition to the third-party email service. In this paper, we implement PhiShield to show the scalability and practicality of our solution and provide a performance evaluation of approximately 300,000 emails from various sources. Full article
(This article belongs to the Special Issue New Technologies for Network Security and Anomaly Detection)
Show Figures

Figure 1

30 pages, 3133 KB  
Article
In-Depth Analysis of Phishing Email Detection: Evaluating the Performance of Machine Learning and Deep Learning Models Across Multiple Datasets
by Abeer Alhuzali, Ahad Alloqmani, Manar Aljabri and Fatemah Alharbi
Appl. Sci. 2025, 15(6), 3396; https://doi.org/10.3390/app15063396 - 20 Mar 2025
Cited by 40 | Viewed by 24266
Abstract
Phishing emails remain a primary vector for cyberattacks, necessitating advanced detection mechanisms. Existing studies often focus on limited datasets or a small number of models, lacking a comprehensive evaluation approach. This study develops a novel framework for implementing and testing phishing email detection [...] Read more.
Phishing emails remain a primary vector for cyberattacks, necessitating advanced detection mechanisms. Existing studies often focus on limited datasets or a small number of models, lacking a comprehensive evaluation approach. This study develops a novel framework for implementing and testing phishing email detection models to address this gap. A total of fourteen machine learning (ML) and deep learning (DL) models are evaluated across ten datasets, including nine publicly available datasets and a merged dataset created for this study. The evaluation is conducted using multiple performance metrics to ensure a comprehensive comparison. Experimental results demonstrate that DL models consistently outperform their ML counterparts in both accuracy and robustness. Notably, transformer-based models BERT and RoBERTa achieve the highest detection accuracies of 98.99% and 99.08%, respectively, on the balanced merged dataset, outperforming traditional ML approaches by an average margin of 4.7%. These findings highlight the superiority of DL in phishing detection and emphasize the potential of AI-driven solutions in strengthening email security systems. This study provides a benchmark for future research and sets the stage for advancements in cybersecurity innovation. Full article
Show Figures

Figure 1

Back to TopTop