Next Article in Journal
Integrity Diagnostics for Cross-Silo Federated Deepfake Speech Detection Under Heterogeneity and Poisoning
Previous Article in Journal
Validating Resource-Competitive DDoS Defense Through Experiment
Previous Article in Special Issue
An Explainable AI Framework for Identity Document Authentication in AML/KYC Verification
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Hybrid Transformer Architecture for Context-Aware Spam Email Classification

1
Faculty of Science and Environment, School of Computer Science, Northumbria University (NUL), London E1 7HT, UK
2
School of Electronic Engineering and Computer Science, Queen Mary University of London (QMUL), London E1 4NS, UK
*
Author to whom correspondence should be addressed.
Future Internet 2026, 18(10), 542; https://doi.org/10.3390/fi18100542
Submission received: 26 August 2026 / Revised: 30 September 2026 / Accepted: 1 October 2026 / Published: 9 October 2026
(This article belongs to the Special Issue Securing Artificial Intelligence Against Attacks)

Abstract

Spam and phishing email remain a persistent cybersecurity threat and a primary delivery vector for credential theft, malware, and ransomware. Detection accuracies above 99% are routinely reported. Such figures are typically obtained on small, single-source, monolingual corpora in which near-duplicate campaign templates span the training and test partitions. This study re-examines what they measure. A corpus of 99,707 emails is assembled from four independent public collections spanning multiple languages, campaign-aware deduplication is applied before splitting, and nine models—classical, convolutional, Transformer, and multimodal—are evaluated under two protocols: an in-distribution split and a cross-source split in which the evaluation source is withheld entirely from training. The multimodal model fine-tunes DistilBERT and fuses its contextual embeddings with thirteen structural features into a 781-dimensional representation classified by a multilayer perceptron. In distribution, six of the nine models fall within one F1 point of each other ( Δ F1 = 0.0001, 95% CI [−0.0012, 0.0012], fusion vs. the fine-tuned text-only baseline). Against an identical classification head without the structural block—the controlled comparison isolating the structural contribution—fusion produces a small but statistically significant gain in distribution ( Δ F1 = 0.0014, 95% CI [0.0005, 0.0023]; Δ recall = 0.0030, CI [0.0016, 0.0044]) that grows roughly five-and-a-half-fold under cross-source evaluation ( Δ F1 = 0.0078, CI [0.0036, 0.0120]; Δ recall = 0.0107, CI [0.0031, 0.0180]). The contribution of multimodal fusion therefore scales with distribution shift rather than appearing only under it and is largest exactly where deployment conditions differ most from training. Against the strongest classical baseline (a linear SVM), the fusion model attains a false-positive rate less than one-fifth as large under cross-source evaluation (0.91% vs. 5.40%) alongside significantly higher F1 (0.9488 vs. 0.9188, non-overlapping bootstrap CIs); in distribution, it attains marginally higher AUC (0.9990 vs. 0.9985) at a marginally higher false-positive rate (0.94% vs. 0.37%). SHAP places 97.6% of signal in the text embeddings, yet that residual share is what produces the cross-source gain—attribution magnitude measured in distribution does not predict conditional contribution under shift.

1. Introduction

Email remains one of the most widely used means of communication across the globe [1], and that ubiquity has made it among the most attractive targets for attackers [2]. Spam—unsolicited and misleading bulk email—is estimated to constitute between 45% and 85% of global email traffic [3,4]. Phishing is the subset that solicits credentials, payment, or the execution of malicious attachments, and it carries the greater share of direct harm through credential theft, ransomware, and identity theft [5,6,7]. Billions are lost annually to email-based fraud globally [8].
The prediction task in this study considers a binary: legitimate mail against unsolicited mail. The positive class is heterogeneous by construction. The TREC collections supply bulk commercial spam, which accounts for approximately 90% of positive training instances; the honeypot collection supplies credential phishing exclusively. This composition is a design property rather than an artefact, and Section 4.3 exploits it. Because the cross-source protocol withholds the honeypot entirely, a model evaluated under Split B is trained almost wholly on commercial spam and tested wholly on credential phishing, and so faces a shift in category as well as in provenance. Where this paper reports phishing-class metrics, it refers to this combined positive class; where the distinction between commercial spam and credential phishing is material, it is stated explicitly.
Going beyond the need to achieve high detection accuracy, modern spam-filtering systems are being required to be more transparent and to comply with data protection regulations like General Data Protection Regulation (GDPR) [9]. Multimodal classifiers combining textual and structural features are widely proposed on the rationale that structural cues persist when lexical content is obfuscated [3,10]. Whether that rationale holds is an empirical question, and this study shows the answer depends on the evaluation protocol applied.
Spam filtering has evolved from rule-based blacklisting in the 1990s through classical machine learning models such as Naive Bayes and SVMs, which improved adaptability but struggled with complex semantics [11]. Modern spam increasingly mimics legitimate communication, embedding deceptive links and social engineering cues that evade static detection [3]. Reliance on static datasets such as Enron [12] further limits generalization to current spamming activity [8]. These trends, combined with the lack of transparency in AI-based filters, motivate detection systems that combine textual and structural cues with explainable outputs [13,14].
Significant improvements have been made to address spam filtration, especially in deep learning techniques. The use of Transformer models like Bidirectional Encoder Representations from Transformers (BERT) and DistilBERT has shown great success in metrics like accuracy, capturing semantic relations in spam emails [10,13,15]. Traditional models that analyse data one element after another use only left-to-right analysis; Transformers reverse this analysis method as well, making them more adept at understanding spam emails [16]. In practical applications, Transformer models fine-tuned with spam data illustrate substantial enhancements in precision, accuracy, and recall rates [14]. Structural features combined with sender-derived data have further improved spam resilience [10,17], and Explainable AI (XAI) techniques like SHapley Additive exPlanations (SHAP) and Local Interpretable Model-agnostic Explanations (LIME) are increasingly adopted to visualize classifier decisions, addressing transparency concerns and GDPR compliance [3,13].
These reported gains, however, rely on an evaluation practice that is rarely examined. Spam detection results above 99% accuracy are now routine, yet they are typically obtained on small, single-source, monolingual corpora in which the same conditions that produce the headline figure also inflate it. Spam arrives in campaigns: thousands of messages are generated from a shared template, differing only in recipient name or a tracking parameter. When such a corpus is partitioned by random stratified sampling, members of the same campaign appear in both the training and test partitions, and the reported test score partly measures template memorisation rather than generalisation. This is a documented failure mode across machine-learning-based science [18,19], and it is particularly acute in email corpora, where provenance itself is often predictive: if legitimate mail is drawn from one collection and spam from another, source-specific artefacts such as header conventions, encoding, or collection period become shortcuts that no longer exist at deployment time.
A second limitation follows from corpus construction. Benchmark spam datasets are predominantly English, and non-English content is commonly discarded during preprocessing as a cleaning step. This produces classifiers whose behaviour outside English is simply unmeasured, even though phishing campaigns are multilingual and cross-lingual transfer in pre-trained encoders is an active research area [20]. Where multilingual evaluation is attempted, an accompanying difficulty is that non-English portions of public corpora are frequently dominated by a single class, so aggregate accuracy on such subsets can approach that of a trivial majority-class predictor.
Alongside these evaluation concerns, practical challenges remain. Class imbalance between normal and spam mail biases reported results [8]; weaknesses in feature extraction expose models to adversarial manipulation [21]; and Transformer models consume substantial computational resources, raising concerns for real-time filtering.
This study therefore reframes the question from how accurately a multimodal spam classifier can score to how much of that score survives when the conditions that inflate it are removed. A corpus of 99,707 emails is assembled spanning multiple languages, near-duplicate campaign groups are identified and collapsed before partitioning, and every model is evaluated under two protocols: an in-distribution split and a cross-source split in which the evaluation source is withheld entirely from training. The proposed multimodal architecture—DistilBERT embeddings fused with thirteen structural features and classified by a multilayer perceptron (MLP)—is assessed within this protocol alongside eight comparison models, with SHAP and LIME applied to establish which signals drive the decision.
The objectives are to
  • Construct a large, multi-source, multilingual email corpus and audit it for source–label association and campaign-level duplication prior to partitioning.
  • Fine-tune a Transformer-based model (DistilBERT) and integrate its contextual embeddings with engineered structural features to build a multimodal classification pipeline.
  • Evaluate all models under identical training configurations and two evaluation protocols, quantifying the generalisation gap between in-distribution and cross-source conditions.
  • Establish which structural features contribute to detection through retraining-based ablation rather than attribution proxies, and verify stability through cross-validation, multi-seed replication, and bootstrap confidence intervals.
  • Apply SHAP and LIME to the multimodal classifier to provide transparency consistent with GDPR accountability requirements.
  • Characterise the class composition of each language above a minimum-count threshold to establish which support classification metrics.
While the objectives above define the specific research tasks undertaken, the contributions below summarize the resulting novel outcomes and their significance to the field of spam detection research.
  • Multimodal fusion as a robustness mechanism, not only an accuracy improvement: Structural features fused with a Transformer encoder produce a small but statistically significant benefit under in-distribution evaluation ( Δ F1 = 0.0014, 95% CI [0.0005, 0.0023]) that grows roughly five-and-a-half-fold under cross-source evaluation ( Δ F1 = 0.0078, CI [0.0036, 0.0120]). This shift-dependent scaling is invisible to the protocol used throughout the prior literature, which reports only in-distribution accuracy, and is established here through a three-arm design in which head architecture is held constant and only the structural block varies.
  • A leakage-audited multilingual corpus: A 99,707-email corpus is assembled and audited for metadata leakage, with sender domain excluded as a provenance identifier rather than a phishing signal.
  • A dual-protocol evaluation that separates memorisation from generalisation: Nine models spanning classical, convolutional, Transformer, and multimodal families are evaluated under in-distribution and cross-source protocols.
  • Retraining-based ablation of structural features: Rather than inferring feature importance from attribution scores, each structural-feature configuration is retrained and re-measured. In distribution, the top-3 SHAP-ranked structural features reduce false negatives by up to 17.7% (124 → 102) with F1 essentially unchanged (0.9862–0.9878 across configurations); under cross-source evaluation, the same features yield a significant advantage over both text-only and other variants.
  • Interval-based robustness reporting: Results are validated through five-fold stratified cross-validation, five-seed replication, bootstrap confidence intervals, and McNemar tests, with intervals reported rather than single-run point estimates. Seed variance under cross-source evaluation proves an order of magnitude larger than in distribution, and model ordering is shown to be unstable there—a result that qualifies single-run rankings throughout this literature.
  • Cross-lingual evaluability analysis: The class composition of every language above a minimum-count threshold is characterised, establishing which languages can support classification metrics and which can support detection rate alone—a prerequisite routinely omitted when multilingual accuracy is reported in aggregate.
  • Joint SHAP and LIME attribution: Both explainers are applied to the same instances on the multimodal classifier, yielding cross-validated attributions that support model debugging and regulatory transparency under GDPR, alongside the demonstration that attribution magnitude does not predict a feature’s conditional contribution under shift.

2. Literature Review

2.1. From Rule-Based Filtering to Classical Machine Learning

Early spam filters applied handcrafted keyword rules and sender blacklists [22]. Both are cheap and auditable, but static: attackers evade keyword rules through character insertion and defeat blacklists through domain spoofing and rapid IP turnover [3]. Bayesian classifiers improved on this by modelling word distributions probabilistically [11,23], reducing false alarms while remaining vulnerable to token-injection attacks that manipulate the underlying estimates [3]. Discriminative classifiers—SVMs, decision trees, and k-NN—learned patterns from labelled data instead [23], with SVMs suiting high-dimensional sparse text. All three families nonetheless depend on manually engineered features that treat content words and header cues independently [5,10], capturing little semantics beyond surface lexical statistics [5]. This limitation motivated the shift to models that learn representations directly from data [13]. This progression is conventionally presented as a succession, with each family superseding the last. Section 4.3 qualifies that reading: under cross-source evaluation, the two most stable models of the nine evaluated here are both classical, and the manually engineered features on which they depend transfer across sources better than learned representations do.

2.2. Deep Learning and Transformer Architectures

CNNs capture local n-gram patterns [24], and RNNs model sequential dependencies [5,25], both outperforming classical methods given sufficient data [22], though RNNs degrade over long sequences. Representation quality proved equally decisive: TF-IDF discards word order [26], while Word2Vec and GloVe encode semantic similarity but assign each word one fixed vector regardless of context [27,28]. Self-attention resolved this by conditioning token representations on the full sequence [29], and BERT extended it through bidirectional pre-training [16]. BERT is now a standard baseline for phishing classification, where contextual encoding exposes social-engineering cues embedded in otherwise plausible text [13,30], routinely exceeding 97% accuracy against CNN and RNN baselines [13,31]. DistilBERT retains most of this accuracy at roughly 40% lower inference cost [15], making it the practical choice where latency matters [10]. Multilingual variants—multilingual DistilBERT and XLM-R [20]—extend the same architecture across languages at greater parameter cost. The present study uses the English uncased checkpoint; Section 3.4 reports the tokenisation consequences of that choice for the non-English portion of the corpus.

2.3. Multimodal Fusion and Explainable AI

Multimodal detection combines Transformer text representations with structural attributes—message length, hyperlink count, HTML markup, sender characteristics—on the rationale that structural cues persist when lexical content is obfuscated [10,32]; the present study excludes sender-derived attributes for reasons given in Section 3.2. SpamNet fuses linguistic, metadata, and temporal branches to reach 98.8% accuracy, outperforming its single-modality ablations [32]. Fusion may occur early, concatenating feature vectors before classification, or late, combining independent model outputs; early fusion is simpler but sensitive to feature scaling and cross-modal class imbalance [10,33]. Reported gains are typically modest relative to the added complexity, a trade-off few studies quantify. These gains are also, without exception among the studies surveyed here, measured under in-distribution evaluation, with training and test partitions drawn from the same collection. Whether multimodal fusion confers an advantage when the evaluation source differs from the training source is consequently untested—a gap Section 4.3 addresses directly and one where the answer proves to differ from the in-distribution result.
SHAP derives additive attributions from cooperative game theory [34], and LIME fits a local interpretable surrogate around individual predictions [35]. Both support triage and debugging [13], and both are approximations rather than faithful accounts of model computation [36]; attention weights are used similarly, though their reliability as explanations is contested [37]. A further limitation is less often noted. Attribution magnitude is an aggregate computed over a particular evaluation distribution, and carries no guarantee of predicting how much a feature contributes under a different one. Section 4.8 demonstrates that discrepancy directly, and it motivates the retraining-based ablation adopted in Section 3.6 in preference to attribution-based feature selection. Interpretability is also a regulatory requirement under GDPR accountability provisions and Responsible AI frameworks [3,38], with on-demand explanation proposed to contain cost [13].

2.4. Evaluation Practice and Research Gaps

How these systems are evaluated has received less scrutiny. Four issues recur.
Spam is generated in campaigns: many messages instantiated from one template with minor per-recipient variation. Random stratified partitioning distributes campaign members across training and test sets, so reported scores partly measure template memorisation. Leakage of this kind is a documented failure mode across machine-learning-based research [18,19], and standard-split evaluation demonstrably overstates performance in NLP [39].
Benchmark corpora are also commonly assembled by pairing a legitimate mail collection with a separately sourced spam collection. Provenance then correlates with the label, and artefacts such as header conventions, encoding, or collection period become shortcuts unavailable at deployment [40,41]. Evaluation on genuinely held-out sources reduces accuracy substantially even where in-distribution scores approach saturation [42].
Results are commonly reported as single-run point estimates. Where several architectures are compared on one partition, differences of a fraction of a percentage point are presented as an ordering without confidence intervals, significance tests, or replication across initialisations. It therefore cannot be determined whether an apparent ranking reflects a property of the models or of a particular random seed; this is a concern established for system comparison in NLP more generally [39].
Finally, public corpora are overwhelmingly English, and non-English content is often discarded in preprocessing, leaving behaviour outside English unmeasured despite multilingual phishing being common and cross-lingual transfer in pre-trained encoders being well studied [20]. Where non-English subsets are evaluated, they are frequently single-class, so aggregate accuracy need not exceed a majority-class baseline.
Table 1 positions representative studies accordingly. Reported accuracies are deferred to Section 4.9: each derives from a different corpus and partitioning scheme, and that incomparability is itself part of what motivates the protocol adopted here. Four gaps follow: single in-distribution scores are reported without testing whether performance survives a change of source; structural feature contributions are inferred from attribution rather than established by retraining; results are reported as single-run point estimates rather than as intervals, so apparent differences between architectures cannot be separated from initialisation variance; and linguistic coverage is rarely reported. This study addresses all four.

3. Methodology

This section describes the corpus, its preprocessing and leakage auditing, the model architectures, and the evaluation protocol. The design goal is not only to build a multimodal classifier but to evaluate it under conditions that distinguish generalisation from memorisation and to establish whether the structural component of that classifier contributes anything the in-distribution protocol can detect. All code, notebooks, seeds, and logs are available in the project repository (https://github.com/geldixhafaj/multimodal-email-spam-detection (accessed on 29 September 2026)).
Figure 1 presents an overview of the framework, from corpus assembly and leakage auditing through the two split designs to model training, dual-protocol evaluation, and attribution.

3.1. Corpus Construction

Experiments use MEPC, a multi-source email phishing corpus assembled from four independent public collections: the TREC 2005, 2006, and 2007 spam tracks [43,44,45]. They supply the legitimate-mail backbone together with approximately 90% of the unsolicited messages, and phishing_pot [46], a public honeypot collection of contemporary phishing messages, used here as a fixed snapshot taken on the access date given in the reference. Combining independently collected sources is deliberate; it enables the cross-source evaluation described in Section 3.3, which no single-source corpus can support.
From 101,229 raw records, cleaning proceeded in logged stages. Records carrying no usable class label were removed first (245 records). Records empty in both subject and body were removed next, of which there were none remaining at that stage. Two-stage deduplication followed: exact deduplication on an MD5 hash of the case-folded, whitespace-collapsed subject and body, then near-duplicate deduplication on a shingle signature—the first 400 characters of the normalised body with digits masked. This collapsed campaign variants that differ only in a recipient name, tracking number, or amount, which exact hashing cannot detect. The two deduplication stages together removed 1277 records. Each retained email carried a dup_group identifier used to enforce split integrity. The final corpus contained 99,707 emails (1.5% removed overall), comprising 56,095 legitimate and 43,612 phishing messages, a class ratio of 1.29:1 (Figure 2). Bodies were parsed to plain text with scripts and styles stripped and encoded as UTF-8. Personal identifiers were replaced with typed placeholders in the source collections.
Non-English content was retained rather than discarded. The corpus spans 91 distinct primary languages, 16 with at least 50 instances, with English accounting for 94.2% (Figure 3). Only four languages—English, German, French and Spanish—contain sufficient instances of both classes to support full classification metrics; the remainder of the non-English subset is predominantly phishing, which constrains how it may be reported. Section 3.6 specifies the resulting protocol.
The thirteen structural features are not individually strong discriminators, and two assumptions commonly built into structural spam features do not hold in this corpus. Figure 4 reports the direction-adjusted univariate AUC of each feature over the full corpus. Only html_flag (0.824) and the URL count and length family (0.614–0.715) separate the classes to any useful degree; the remaining features fall between 0.506 and 0.560. Exaggerated capitalisation, a reliable indicator in SMS-derived spam corpora, is not one here: uppercase_ratio reaches 0.511, and its weak tendency runs towards the legitimate class rather than the phishing class. attachment_count (0.509) and digit_count (0.506) are indistinguishable from chance. The structural block is therefore not expected to carry the classification alone, and Section 4.4 measures what it contributes in combination.

3.2. Leakage Auditing

Because provenance can correlate with label in multi-source corpora, metadata fields were probed for predictive power before any modelling (Figure 5). Classifiers trained on sender domain, receiver domain, and content-type headers alone were evaluated, and every field that separated the classes on provenance rather than content was excluded from the modelling feature set. The exposure is substantial. Against a majority-class baseline of 56.3%, sender domain alone predicts the label with 80.2% accuracy (AUC 0.846), source identity alone with 61.6% (AUC 0.649), collection year alone with 62.0% (AUC 0.667), and all metadata fields jointly with 85.7% (AUC 0.918), without inspecting a single word of message content. Sender domain in particular is treated as a corpus artefact, not a phishing signal: a pipeline retaining it would report a strong result reflecting collection origin rather than detection capability.
Campaign structure was audited separately (Figure 6). Phishing arrives as templated batches, so any partitioning permitting members of one campaign to fall on both sides of the train/test boundary yields a score that partly measures template recall. Near-duplicate deduplication removes this exposure at source rather than managing it at partition time: each campaign is collapsed to a single retained exemplar, so every dup_group in the cleaned corpus is a singleton. The split-integrity constraint described in Section 3.3 therefore verifies campaign separation rather than enforcing it and is retained as a guard against residual duplication that the shingle signature may not have caught.

3.3. Split Design

Two independent partitionings were constructed (Figure 7).
Split A (in-distribution) stratifies by class into approximately 72% training, 8% validation, and 20% test, subject to the constraint that no dup_group may span training and test. This measures performance under the conventional assumption that deployment data resembles training data.
Split B (cross-source) withholds the phishing_pot source entirely, assigning it exclusively to the test partition. No model sees any instance from this source during training or validation. This measures whether performance survives a change of provenance and collection period and constitutes the held-out dataset evaluation that single-corpus protocols cannot provide.
One property of Split B must be stated explicitly, because it bounds what the protocol can demonstrate. The withheld source is composed entirely of phishing messages (4490 of 4490); its test partition draws legitimate mail from the TREC collections instead. Source and class are therefore not fully separable under this split: a model evaluated on Split B faces both a change of provenance and a change in the composition of the positive class, which shifts from predominantly commercial spam in training to credential phishing at test. The design establishes that performance degrades under this combined shift and supports comparison between models evaluated on the identical partition but does not isolate provenance as the sole cause. Section 6 identifies rotating leave-one-source-out evaluation over the three class-mixed TREC collections as the direct remedy.
Five integrity assertions were verified programmatically before modelling: complete assignment under Split A, partition disjointness, no duplicate group spanning train and test under either split, both classes present in every partition, and containment of the held-out source to Split B’s test partition. All five passed; the assertion halts the pipeline on failure.
Class imbalance is addressed through class-weighted loss rather than resampling, avoiding both the duplicated patterns introduced by oversampling and the information loss of undersampling. Weights are derived from the training partition alone (0.8887 legitimate, 1.1431 phishing under Split A). Structural feature scalers are likewise fitted on training data only and applied unchanged to validation and test. Because the cross-source protocol evaluates on a collection whose URL length distribution differs sharply from the training sources (PSI 1.750 for url_length_max; training mean 38.8 against 120.0 on the held-out source), structural features are additionally winsorised at the 99.5th percentile of the training partition before standardisation. Without this step, held-out source values fall several standard deviations outside the fitted range and saturate scale-sensitive linear models.

3.4. Models

Nine models are evaluated. Three classical references establish performance floors: a random forest on structural features alone, TF-IDF with logistic regression, and their concatenation. Three further baselines cover common alternatives, i.e., multinomial Naive Bayes and a linear SVM, both on TF-IDF, and a word-embedding CNN trained from scratch. TF-IDF was selected as the classical lexical representation in preference to Word2Vec because the model set already spans the representational hierarchy described in Section 2.2, i.e., a dense static representation in the CNN’s from-scratch embedding layer and a dense contextual one in DistilBERT. A sparse, interpretable lexical representation complements these rather than duplicating the CNN arm.
Three Transformer arms are evaluated rather than two, so that the contribution of the structural block can be isolated with architecture held constant.
The text-only baseline fine-tunes distilbert-base-uncased—6 encoder layers, 12 attention heads, 768 hidden dimensions, and approximately 66 M parameters, being 40% smaller than BERT-base while retaining 97% of its language understanding [15]. A single linear classification head is applied to the pooled [CLS] representation, and all six encoder layers are updated end-to-end [16]. The uncased variant was selected because case-sensitive tokenisation would fragment the subword vocabulary across capitalisation variants; case information is retained separately as a structural feature, though as Figure 4 shows, it contributes little.
The ablated head takes fixed-size mean-pooled embeddings from that same fine-tuned encoder and classifies them with the feed-forward network described below, receiving no structural input. It exists as a control rather than as a competitive baseline: it shares the encoder, the pooling strategy, and the head architecture of the multimodal model and differs from it only in the width of its input layer.
The multimodal model extracts fixed-size contextual embeddings from the fine-tuned encoder and concatenates them with the scaled structural vector, producing a 781-dimensional representation classified by a feed-forward network with ReLU hidden layers and dropout. Early fusion was selected over attention-based or late fusion for its lower computational overhead [47,48]. Embeddings and structural features are aligned by dataset identifier, with sample counts and missing values verified. A separate encoder is fine-tuned per split, so no Split B model is ever exposed to held-out-source data at any stage. Because the ablated head receives 768 dimensions and the multimodal model 781, any difference between the two is attributable to the thirteen structural features alone and not to encoder capacity, pooling strategy, or head design. Section 4.3 reports that difference under both evaluation protocols, and the comparison is the basis of this study’s principal finding.
One consequence of the uncased English checkpoint should be stated here. Non-English text is tokenised with the same English vocabulary rather than a multilingual one, and the effect is measurable: subword fragmentation rises from 1.79 wordpieces per word in English to 16.40 in Russian, 21.38 in Japanese, and 24.23 in Chinese, while the share of messages truncated at 256 tokens rises from 44.0% in English to 60.8% in Japanese and 75.3% in Ukrainian. The per-language results in Section 4.7 therefore characterise how an English-centric encoder behaves on other languages; they are not evidence of multilingual competence.

3.5. Training Configuration

All neural models train under an identical configuration. The epoch budget is adaptive rather than fixed: training begins with a budget of 10 epochs and extends in increments of 5 to a ceiling of 25, unless one of two stopping conditions is met first—for DistilBERT, early stopping on validation loss with a patience of 3 epochs, or a validation-loss plateau, defined as a range below 0.002 across the last four epochs; the checkpoint retained is the epoch of lowest validation loss, not highest validation F1. The CNN baseline retains the original validation-F1-based criterion (patience 3, plateau tolerance 0.0015), since only the DistilBERT encoder was regularised. Under Split A, DistilBERT ran 11 epochs before patience exhausted, with epoch 8 selected as the checkpoint; under Split B, it ran 7 epochs, with epoch 4 selected. Optimisation uses AdamW with decoupled weight decay [49], linear warmup over the first 10% of steps followed by linear decay to zero, gradient clipping at max-norm 1.0, mixed-precision training, and class-weighted cross-entropy. No model receives a different epoch budget, patience value, or optimiser setting; classical baselines are fitted to convergence, having no epoch budget by construction. Where models differ, the difference is the point at which early stopping triggered, which is reported per model.
Class imbalance is handled by loss weighting alone, with weights n / ( 2 n c ) derived from each design’s training partition. No oversampling, undersampling, or synthetic data generation is applied at any stage of the pipeline.
Results reported as single runs use seed 42; Section 4.5 reports replication across seeds 42–46. Table 2 provides the full configuration for every model family.

3.6. Evaluation Protocol

Every model is evaluated on both splits using accuracy, per-class precision, recall and F1, macro-F1, confusion matrices, and ROC-AUC. The generalisation gap—the difference between Split A and Split B performance—is reported as a primary result rather than a diagnostic.
Five analyses support the findings. Retraining-based ablation refits the fusion model under four configurations (text only; text with the three highest-attribution structural features; text with the six features named in the reference study; text with all thirteen), measuring contribution directly rather than inferring it from attribution scores. All four configurations reuse the frozen embeddings of the same fine-tuned encoder and differ only in the width of the structural block appended to them (768, 771, 774, and 781 dimensions); the text-only configuration is therefore not the independently fine-tuned baseline of Section 3.4 but its frozen-embedding counterpart. Five-fold stratified cross-validation and multi-seed replication establish stability, reported as mean and standard deviation. Statistical testing accompanies every headline comparison: bootstrap confidence intervals are computed over test predictions for F1 and recall, and paired bootstrap differences are reported alongside McNemar tests for the four contrasts that bear on the study’s claims. A difference is described as significant only where its confidence interval excludes zero and the McNemar test survives Bonferroni correction.Threshold sensitivity analysis varies the decision boundary from the default 0.5 to characterise the precision–recall trade-off. Per-language analysis establishes which of the sixteen languages above the minimum-count threshold can support classification metrics; because most non-English subsets contain phishing instances only, full metrics are reportable solely for languages containing both classes, of which there are four.

3.7. Explainability Methods

SHAP [34] and LIME [35] are applied to the multimodal classifier to quantify feature contributions globally and for individual predictions, covering both correct classifications and errors to support model debugging and the identification of systematic failure patterns [50]. Global attribution is computed over a 300-instance subset of the Split A test partition; local attribution is computed for one correctly classified phishing message and one misclassified instance, with both explainers fitted to the same two instances so that their attributions can be compared directly.
Both are approximations rather than faithful accounts of model computation and are interpreted accordingly. A further limitation bears directly on the use made of them here. An attribution score is an aggregate computed over a particular evaluation distribution, and carries no guarantee of describing a feature’s contribution under a different one. Attribution is therefore reported alongside the retraining-based ablation of Section 3.6 rather than in place of it, and Section 4.8 quantifies the discrepancy between what attribution suggests and what retraining measures.

3.8. Ethical Considerations and Reproducibility

All source corpora are publicly available and previously published for research use. The honeypot collection is distributed under the MIT license and was accessed as a fixed snapshot on the date recorded in its reference [46]; it no longer receives public updates, so that archived snapshot rather than the live repository is the reproducible artifact. Sender and receiver fields were consumed during leakage auditing and dropped before modeling; no personal identifiers enter the feature set. Deployment on live traffic would require a GDPR lawful basis and, where automated decisions materially affect users, meaningful explanation of those decisions [51,52]. Explainability supports but does not constitute compliance, which remains an organisational responsibility.
Code, notebooks, configuration files, seeds, and training logs are version-controlled in a public repository (https://github.com/geldixhafaj/multimodal-email-spam-detection (accessed on 29 September 2026)) with replication instructions covering environment setup, preprocessing, training, and interpretability artefact generation. The complete training configuration for every model family, together with the software versions and hardware used, is given in Table 2; all reported runs use seeds 42–46.

4. Results and Discussion

Results are reported under both split designs throughout. Split A measures in-distribution performance; Split B withholds the phishing_pot source entirely from training. The difference between them is treated as a primary finding rather than a diagnostic.
Two distinct comparisons run through this section and should be kept separate. The first is between protocols for a given model—the generalisation gap. The second is between models evaluated on the identical test partition, and in particular between the multimodal model and its own ablation, which differ only in the presence of the structural block. The two are not equally exposed to the composition of the held-out source described in Section 3.3: the first confounds provenance shift with the change in class composition, whereas the second holds the test instances fixed and varies only the model. Headline comparisons are therefore reported with bootstrap confidence intervals and McNemar tests rather than as point estimates, and a difference is called significant only where both agree.
The shift structure of all four candidate held-out sources was characterised before committing to this design (Figure 8). Measured by the Population Stability Index over the thirteen core structural features, the withheld honeypot source exhibits major shift ( PSI > 0.25 ) on six features. The three class-mixed TREC collections are not milder alternatives: holding out trec5 produces major shift on nine features and trec7 on eight, while trec6 produces it on three. Withholding phishing_pot is therefore not a weak choice of held-out source, and the rotating evaluation identified in Section 6 would probe a range of shift severities rather than a single regime. What the TREC folds additionally offer is the absence of the source–class confound, since each retains both classes.

4.1. Training Dynamics

Figure 9 shows the full fine-tuning trajectory under both designs. Under Split A, training ran 11 epochs, with the checkpoint selected at epoch 8 on the validation-loss minimum (0.0818, against a training loss of 0.0529 measured in evaluation mode—a gap of 0.0289); validation F1 at that epoch is 0.9840 and accuracy 0.9861. Under Split B, training ran 7 epochs, with the checkpoint at epoch 4 (validation loss 0.0703 against training loss 0.0557, a gap of 0.0146); validation F1 at that epoch is 0.9868 and accuracy 0.9884.
The adaptive budget’s ceiling of 25 epochs was approached under neither design. Training ended on the patience criterion rather than on budget exhaustion, and the checkpoint carried forward is in each case the epoch of lowest validation loss rather than the last. A longer budget could therefore not have altered the reported results: validation loss had already stopped improving three epochs before training halted under both designs.
Training loss, measured in evaluation mode for direct comparability with validation loss, falls from 0.117 at epoch 1 to 0.053 at the epoch-8 checkpoint under Split A and from 0.108 to 0.056 at the epoch-4 checkpoint under Split B, tracking validation loss down to its minimum of 0.0818 (Split A) and 0.0703 (Split B) at the same checkpoints. After the checkpoint, validation loss rises only slightly over the remaining patience window—by 0.0039 over three epochs under Split A and 0.0040 under Split B. The gap between training and validation loss at the selected checkpoint is 0.0289 under Split A and 0.0146 under Split B, an order of magnitude narrower than the divergence the unregularised configuration produced by its final epoch. Validation-expected calibration error stays flat across training under both designs (0.040–0.047 under Split A and 0.038–0.045 under Split B).

4.2. In-Distribution Performance

Under Split A, the text-only DistilBERT baseline achieves precision of 0.992, recall of 0.983, F1 of 0.9877 and AUC of 0.999; the multimodal model achieves precision of 0.988, recall of 0.988, F1 of 0.9878 and AUC of 0.999.
Fusing structural features with the Transformer produces a small but statistically significant in-distribution benefit, visible only in the controlled comparison. Against the fine-tuned text-only baseline, the difference remains negligible: Δ F1 = 0.0001, 95% CI [−0.0012, 0.0012], McNemar p = 1.0. Against an identical classification head with the structural block removed—the controlled comparison, holding encoder, pooling and head architecture constant—it is Δ F1 = 0.0014, 95% CI [0.0005, 0.0023], Δ recall = 0.0030, CI [0.0016, 0.0044], McNemar p = 0.0037, significant after Bonferroni correction. These two comparisons diverge because the fully fine-tuned text-only model and the frozen-embedding ablated head are not the same classifier: the former updates all encoder weights end-to-end under its own head, and the latter reuses frozen embeddings from that same encoder under a separately trained MLP. Against its true counterfactual (an identical head differing only in the presence of the structural block), the multimodal architecture is measurably, if modestly, better even in distribution. Section 4.3 shows this advantage widening substantially under cross-source evaluation: the contribution of structural fusion is present under both protocols and amplified, not created, by distribution shift.
The spread across the full model set is the second observation (Figure 10, amber bars). Nine architectures spanning multinomial Naive Bayes to a fine-tuned Transformer separate by roughly six F1 points, and six of the nine fall within a single point of one another: TF-IDF logistic regression scores 0.9797, the classical structural-plus-lexical model 0.9803, a linear SVM 0.9857, and the three DistilBERT variants 0.9864 to 0.9878. The Transformer’s margin over the classical lexical baseline is narrow but statistically real ( Δ F1 = 0.0080, 95% CI [0.0057, 0.0103], p < 0.001 ). Under this protocol, in-distribution evaluation has limited power to discriminate between architectures—precisely the condition producing the near-saturated figures common in this area of the literature.

4.3. Cross-Source Generalisation

Under Split B, performance falls for every model (Figure 10, red bars). For the text-only baseline, F1 declines from 0.9877 to 0.9444 and recall from 98.3% to 90.8%; for the multimodal model, F1 declines from 0.9878 to 0.9488 and recall from 98.8% to 91.5%.
The operational reading is sharper when misses are expressed as a share of the phishing instances present, since the two test partitions differ in size. For the text-only baseline, missed phishing rises from 146 of 8723 instances (1.67%) to 412 of 4490 (9.18%)—a 5.5-fold increase in miss rate—while its false-positive rate rises slightly, from 0.60% to 1.01%. For the multimodal model, missed phishing rises from 107 of 8723 (1.23%) to 382 of 4490 (8.51%), a 6.9-fold increase, while its false-positive rate is essentially flat (0.9% to 0.91%) (Figure 11 and Figure 12). Only the fusion model becomes more conservative under shift; the text-only model’s error profile worsens on both sides of the confusion matrix.
Structural fusion’s contribution scales with distribution shift. Section 4.2 established that the controlled ablation—an identical head differing only in the presence of the thirteen structural features—is statistically significant even in distribution, though small ( Δ F1 = 0.0014, 95% CI [0.0005, 0.0023]). Under cross-source evaluation, the same comparison grows substantially ( Δ F1 = 0.0078, 95% CI [0.0036, 0.0120], and Δ recall = 0.0107, 95% CI [0.0031, 0.0180]) roughly 5.6 times the in-distribution F1 gain. Both McNemar tests survive Bonferroni correction (p = 0.0037 in distribution, p < 0.001 under shift). Against the fine-tuned text-only baseline (a different classifier trained end-to-end rather than the controlled counterfactual), the comparison is less informative; neither split’s F1 difference is significant (A: Δ F1 = 0.0001, CI [−0.0012, 0.0012], p = 1.000; B: Δ F1 = 0.0044, CI [−0.0001, 0.0092], p = 0.065), which is expected, since the text-only model differs from the fusion model in more ways than the structural block alone. Table 3 places all four comparisons side by side.
The contribution of multimodal fusion is therefore conditional on distribution shift. It is not an accuracy improvement, and the in-distribution protocol standard in this literature cannot detect it: an ablation run under Split A alone would have justified discarding the structural block. The mechanism is consistent with what the features encode. Lexical and contextual signals are properties of a campaign’s language, which changes with the source; URL structure, HTML declaration and message geometry are properties of message form, which changes less. When semantic evidence becomes unreliable, the structural block supplies a residual signal that remains valid—which is precisely when it matters and precisely when the standard protocol cannot see it.
Comparison with classical baselines. The fusion model ranks second on held-out-source F1 at 0.9488, behind structural-plus-lexical logistic regression at 0.9510, but the difference is not significant ( Δ F1 = −0.0022, CI [−0.0077, 0.0038], McNemar p = 0.635 ), and across five seeds the fusion model records 0.9446 ± 0.0076, a range that straddles the classical value. The text-only arm is likewise tied with lexical-only logistic regression on F1 ( Δ F1 = 0.0005, CI [−0.0049, 0.0065]) and lower, though not significantly, on recall ( Δ = −0.0094, CI [−0.0186, 0.0004]). A fine-tuned Transformer does not, by itself, generalise across sources better than a well-specified classical model; structural fusion is what closes that gap.
The two are not equivalent, however: they occupy different operating points. The fusion model operates at a lower false-positive rate (0.91% against 1.66%) and higher precision (0.9854 against 0.9739), while the classical model has the higher AUC (0.9925 against 0.9844) and sits at a significantly higher-recall point ( Δ recall = −0.0143, CI [−0.0235, −0.0042]). For email filtering, where blocking legitimate correspondence carries the higher operational cost, the fusion model’s operating point is the more useful one.
Generalisation gap. Model ordering also changes (Table 4). The structural-only random forest—weakest in distribution—is by a wide margin the most stable, losing 0.0061, while Naive Bayes degrades furthest at 0.0972. The semantic component is what fails to transfer, and consistently with the ablation, the Transformer arm carrying structural features has the smallest gap of the three.

4.4. Ablation

Rather than inferring structural feature contribution from attribution scores, the fusion model was retrained under four configurations (Figure 13). This is feasible because the fusion head trains on frozen embeddings in under a second, so ablation costs no additional encoder fine-tuning. As described in Section 3.6, all four configurations share the same fine-tuned encoder and differ only in the width of the structural block appended to the frozen embeddings, so the text-only arm here is not the independently fine-tuned baseline of Section 4.2 but its frozen-embedding counterpart.
The pattern is consistent and narrow. Every configuration containing structural features improves recall over text alone, raising phishing recall from 0.9858 to between 0.9860 and 0.9883 and reducing false negatives from 124 to between 102 and 122—a reduction of up to 17.7%. F1 varies more than previously observed: from 0.9862 (text only) to 0.9875–0.9878 across the three structural configurations, a range of 0.0016; macro-F1 varies by 0.0014. The three structural configurations remain above the text-only arm, though the spread across configurations is now comparable in size to the gain over text alone, so individual-configuration differences should not be over-interpreted without repeated fits. Structural features purchase recall at negligible precision cost rather than improving overall classification quality in distribution. The common claim that multimodal integration improves spam detection performance is, on this evidence and under this protocol, true only if performance is defined as recall.
Feature count is also not monotonic with benefit, and more starkly than previously observed: the top three SHAP-ranked features (102 false negatives) outperform both the six named features (107) and all thirteen (122), with false negatives increasing monotonically as more structural features are added beyond the top three, indicating that most of the extended feature set is redundant or actively unhelpful in distribution. This matches the univariate analysis in Figure 4, where several features sit near chance individually, and the correlation structure, where two URL feature pairs exceed r = 0.85 .
This ablation must be read alongside Section 4.3, and the contrast between them is the point. Measured in distribution, the structural block’s point estimate in this ablation is small (recall margin of 0.02 percentage points; F1 margin of 0.0016), directionally consistent with the modest effect size established in Table 3 ( Δ F1 = 0.0014, Δ recall = 0.30 percentage points). The two analyses use independently trained MLP heads on the same frozen embeddings, so some difference in point estimate between them is expected; both agree that the in-distribution effect, where it exists, is small. Measured on a held-out source, the same block yields a substantially larger, statistically significant advantage on both F1 and recall over two independent text-only variants. Ablation conducted under in-distribution evaluation systematically understates the contribution of features whose value is conditional on distribution shift. This is a methodological finding independent of the particular features used here.

4.5. Statistical Robustness

Three stability checks were applied: five-fold stratified cross-validation with campaign groups held within folds, retraining across five random seeds, and bootstrap confidence intervals with paired differences. Cross-validation was applied to the two neural heads—the MLP on embeddings alone and the multimodal fusion head—on both test pools rather than to all nine models, since only these refit cheaply enough on frozen embeddings. The fusion head records F1 0.9871 ± 0.0022 on the Split A pool and 0.9951 ± 0.0007 on the Split B pool, exceeding its ablation on five of five folds under Split B and four of five folds under Split A (losing fold 4, 0.9852 against 0.9869). Grouped cross-validation of the classical baselines over the whole corpus was not performed; their stability is established instead through seed replication.
Seed replication spans seeds 42–46, with per-seed phishing F1 for every model given in Table 5. Under Split A, the fusion model records 0.9879 ± 0.0004 and the ablated head 0.9865 ± 0.0002; under Split B, 0.9446 ± 0.0076 and 0.9445 ± 0.0043. Variance under the cross-source protocol is an order of magnitude larger than in distribution: the stability that in-distribution replication appears to demonstrate is partly a property of the protocol rather than of the models, and single-run cross-source figures should be treated accordingly.
Model ordering is stable under Split A, where all five seeds reproduce the seed-42 ranking exactly. Under Split B, it is not: four of five seeds produce at least one rank swap, most often exchanging the fusion model with the classical structural-plus-lexical baseline, the pair whose confidence intervals overlap in Section 4.3. The fusion model’s advantage over its own ablation is stable across every seed under Split A (won in 5 of 5), but not under Split B, where it is won in only 3 of 5 seeds despite the significant effect reported in Table 3 at seed 42. This reinforces rather than contradicts the caution above: the cross-source comparison’s statistical significance at one seed coexists with real seed-to-seed instability in which model wins. Single-seed cross-source superiority claims, including this study’s own headline comparison, should be read as significant on average evidence rather than a guaranteed ranking.

4.6. Operating Point Selection

Varying the decision boundary across validation data (Figure 14) locates the F1-optimal threshold at 0.49, essentially identical to the 0.50 default, yielding a test F1 of 0.9878—the same as the default threshold’s own F1. The class-weighted objective has already positioned the boundary close to optimal under the in-distribution protocol, so no material post hoc threshold adjustment is required—a result that does not hold for models trained on heavily imbalanced corpora, where substantial displacement from 0.5 is typical. The right-hand panel expresses the trade-off in deployment terms: under a false-positive budget of 1%, phishing recall of approximately 98% is achievable.
The same does not hold under cross-source evaluation, and the difference is practically consequential. On Split B, the F1-optimal threshold falls only to 0.46, close to the 0.50 default, and adopting it raises F1 marginally, from 0.9488 to 0.9497. Selecting instead for a 1% false-positive budget gives a threshold of 0.42 and recall of 92.0% against 91.5% at the default 21 additional phishing messages recovered from the same 4490, while remaining inside the budget.

4.7. Cross-Lingual Evaluability

Reporting per-language results requires first establishing which languages can support them. Figure 15 gives the class composition of the sixteen languages with at least fifty instances. English is close to balanced at 40.5% phishing (38,003 phishing against 55,902 legitimate), but every other language is dominated by phishing: German 97.3% (1297 against 36), French 90.0% (369 against 41), Spanish 93.3% (374 against 27), and the remainder between 89.9% and 100%. Slovenian, Finnish, and Persian contain no legitimate messages at all.
Applying a criterion of at least fifteen instances in both classes, four languages—English, German, French and Spanish—support full classification metrics at corpus level. After partitioning, however, only English retains sufficient instances of both classes within the test partition: German contributes twelve legitimate messages to the Split A test set, French eight, and Spanish three. Full metrics are therefore reported for English alone, and phishing detection rate for every other language (Table 6). For the remaining languages, detection rate is the only metric carrying information; precision, accuracy, and F1 are omitted rather than reported at values the class composition would inflate, since a classifier predicting phishing unconditionally would exceed 0.96 accuracy on all but one of these subsets. Detection rates are therefore reported against an explicit majority-class baseline, and languages contributing fewer than twenty test instances, together with those below the fifty-instance corpus threshold, are pooled so that no metric is quoted from a handful of messages.
These figures should be read with the tokenisation caveat of Section 3.4 in mind. Non-English text is processed by an English-trained vocabulary, and detection rates at or near 1.000 are computed on subsets where a majority-class predictor scores almost as well. Chinese remains the tightest case: both models now detect all 73 phishing messages in this subset (1.0000), only 1.35 points above the 0.9865 a classifier predicting phishing unconditionally would already achieve on the same 74-message subset containing a single legitimate example. Such figures still record the composition of the subset more than the capability of the model.

4.8. Explainability Results

SHAP attribution over a 300-instance test subset shows the text embedding block accounting for 97.6% of total mean |SHAP| (Figure 16), with mean absolute value of 0.48 against structural contributions two orders of magnitude smaller. Within the structural block, HTML declaration, mean URL length, and body length lead, while digit_count, punct_density and attachment_count rank last—digit_count and attachment_count consistent with their near-chance univariate AUCs in Figure 4.
This is consistent with the ablation but must not be conflated with it. A 2.4% attribution share coincides with a 17.7% reduction in false negatives in distribution and with a statistically significant advantage over its own no-structural ablation under cross-source evaluation (Section 4.3). Attribution magnitude is therefore not a substitute for retraining: a feature can carry negligible attributed weight overall while determining the outcome on precisely those instances where the semantic signal is ambiguous. Because attribution here was computed on in-distribution predictions, it measures the structural block under exactly the conditions in which its contribution is smallest. This is the principal methodological reason the ablation refits models rather than reading Figure 16.
Local explanations (Figure 17) show the same asymmetry at instance level. For a correctly classified phishing message, the semantic signal dominates at +0.61, with no structural feature registering materially against it. In the error case—a legitimate message assigned to the phishing class—the semantic signal is again positive at +0.55, and the structural block reinforces rather than opposes it, with html_flag, body_length positive, uppercase_ratio and url_length_avg dissenting at magnitudes an order smaller. Structural features do not act as a corrective on individual in-distribution errors; their contribution, as Section 4.3 shows, emerges in aggregate and under shift.
LIME, fitted to the same two instances, agrees with SHAP on the dominant term but diverges on the structural block: url-subdom-avg carries negligible weight under SHAP in the error case but is LIME’s largest negative contributor. Where attributed magnitudes are this small, neither rank nor sign is stable across methods, which bounds how much weight the structural attributions can bear and supports treating both explainers as post hoc approximations rather than accounts of model computation.

4.9. Comparison with Prior Work

Table 7 positions the proposed framework against representative systems. Values are taken from the respective publications and were obtained on different corpora under different partitioning schemes; direct numerical comparison is therefore not meaningful, and the table characterises evaluation practice rather than ranking performance.
The pattern across the columns is the substantive observation. Reported accuracies cluster between 0.96 and 0.996 across markedly different architectures; none of the compared systems reports cross-source evaluation or linguistic coverage. The in-distribution results in Section 4.2 reproduce that clustering exactly, and the cross-source results in Section 4.3 indicate what it conceals.
Asliyuksek et al. [10] is the most architecturally comparable work—DistilBERT fused with structural features, the same combination evaluated here—and reports marginally higher in-distribution accuracy. The contribution here is not an accuracy improvement over that result. It is the demonstration, under a protocol that study did not apply, that in-distribution accuracy at this level does not discriminate between architectures; that the structural component of such a model contributes a small but statistically significant improvement even when training and evaluation data share a source ( Δ F1 = 0.0014, 95% CI [0.0005, 0.0023] against the same model without it); and that its contribution grows roughly fivefold when they do not ( Δ F1 = 0.0078, 95% CI [0.0036, 0.0120]). That amplification under distribution shift is undetectable under any evaluation protocol the compared systems employ.

5. Conclusions

This study re-examined a result that spam detection research routinely reports, that is, accuracy of above 99% on a held-out test set. Using a corpus of 99,707 emails from four independent sources spanning 91 languages, which were campaign-deduplicated before partitioning, nine models were evaluated under an in-distribution split and a cross-source split withholding an entire source from training.
In-distribution evaluation proved to have limited power to discriminate between architectures: nine models from multinomial Naive Bayes to a fine-tuned Transformer separate by roughly six F1 points, with a TF-IDF logistic regression scoring 0.9797 against the Transformer’s 0.9877. Fusing structural features with the Transformer produced a small but statistically significant gain even under this protocol; against an identical head with the structural block removed, Δ F1 = 0.0014, 95% CI [0.0005, 0.0023].
Under cross-source evaluation, that small effect grows substantially larger, and this is the study’s principal finding. The fusion model significantly outperforms its own no-structural ablation on the held-out source, gaining Δ F1 = 0.0078, 95% CI [0.0036, 0.0120], with McNemar significant under correction; the advantage over the separately trained text-only DistilBERT variant does not reach significance on this split. The contribution of multimodal fusion is not conditional but scale-dependent; small and only sometimes detectable under the protocol standard in this area of the literature; and reliably measurable once provenance changes.
A measurable portion of overall performance does not survive that change of provenance, and the ranking does not survive it either. Cross-source F1 falls to 0.9444 for the text-only baseline and 0.9488 for the multimodal model, with missed phishing rising from under 2% to roughly 9% of phishing instances—concentrated entirely on the class that matters operationally. The structural-only random forest, weakest in-distribution, proves most stable at a gap of 0.0061; a fine-tuned Transformer does not by itself generalise across sources better than a well-specified classical model, and structural fusion is what closes that gap.
Retraining-based ablation qualifies a common claim in this literature. Structural features reduce false negatives by up to 17.7% in distribution, with F1 varying by 0.0016 across all four configurations; measured on a held-out source, the same features yield a statistically significant advantage. SHAP places 97.6% of the total signal in the text embedding block, yet the remaining 2.4% carries that entire cross-source effect. Ablation and attribution conducted in distribution systematically understate the contribution of features whose value is amplified by shift.
Four constraints bound these findings. The corpus remains 94.2% English and the encoder is English-trained, so non-English results characterise an English model’s behaviour on other languages rather than multilingual competence. The cross-source protocol withholds a single source, composed entirely of phishing, so provenance shift and the change in the positive class are not separable in the degradation figures; between-model comparisons, conducted on identical test instances, are unaffected. Adversarial robustness was not evaluated. SHAP and LIME are post hoc approximations offering no formal guarantee of faithfulness.
More broadly, the results suggest that a single in-distribution figure is insufficient to characterise a spam classifier and that components evaluated only in distribution may be discarded precisely because the protocol cannot see what they are for. Where a corpus permits it, evaluation on a genuinely held-out source costs little and reveals differences—including differences favourable to the architecture under proposal—that in-distribution testing cannot.

6. Future Work

Seven directions follow from the limitations identified above.
  • Rotating cross-source evaluation refers to withholding each of the three TREC collections in turn, each of which contains both classes, so that provenance shift can be isolated from the change in class composition that accompanies withholding the honeypot source. Adding temporal splits would further separate general provenance sensitivity from an artefact of the particular collection held out here.
  • Balanced multilingual evaluation refers to assembling non-English legitimate mail so full classification metrics can be reported beyond English, with multilingual encoders such as XLM-R as comparison.
  • Generality of the conditional-contribution effect refers to testing whether other feature families with small in-distribution contributions become significantly more important under shift and whether the effect observed here holds for shift types other than provenance. If it generalises, in-distribution ablation is unsafe as a feature selection procedure wherever deployment data may differ from training data.
  • Adversarial robustness refers to stress-testing both feature families against character substitution, homoglyph attacks, and perturbed text. The cross-source results suggest a testable hypothesis: structural features, which transfer better across sources, may prove more brittle under deliberate evasion precisely because they are simple to manipulate.
  • AI-generated phishing refers to evaluating detection of generative-model output, which lacks the templated repetition that campaign-level deduplication exploits.
  • Human-centred explainability assessment refers to measuring whether SHAP and LIME attributions improve analyst triage accuracy and calibrated trust, rather than assuming that available explanations are useful ones.
  • Deployment cost characterisation refers to profiling the full pipeline including embedding generation, where the encoder rather than the fusion stage is the binding constraint, and assessing quantisation or distillation.

Author Contributions

G.X.: Methodology, Validation, Formal analysis, Investigation, Data curation, Writing—original draft. U.B.C.: Conceptualization, Visualization, Resources, Writing—review and editing, Supervision, Funding acquisition. H.J.: Project administration. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Informed Consent Statement

All authors give consent for the publication of identifiable details, which can include photographs and/or videos and/or case history and/or details within the text to be published in the above Journal and Article.

Data Availability Statement

The data that support the findings of this study are available from the corresponding author, Umair B. Chaudhry, upon reasonable request.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Cormack, G.V. Email spam filtering: A systematic review. Found. Trends Inf. Retr. 2008, 1, 335–455. [Google Scholar] [CrossRef] [Scilit]
  2. Hong, J. The state of phishing attacks. Commun. ACM 2012, 55, 74–81. [Google Scholar] [CrossRef] [Scilit]
  3. Jáñez-Martino, F.; Alaiz-Rodríguez, R.; González-Castro, V.; Fidalgo, E.; Alegre, E. A review of spam email detection: Analysis of spammer strategies and the dataset shift problem. Artif. Intell. Rev. 2023, 56, 1145–1173. [Google Scholar] [CrossRef] [Scilit]
  4. Zimba, A.; Wang, Z.; Chen, H. Multi-stage crypto ransomware attacks: A new emerging cyber threat to critical infrastructure and industrial control systems. ICT Express 2018, 4, 14–18. [Google Scholar] [CrossRef] [Scilit]
  5. Liu, X. Deciphering Spam Through AI: From Traditional Methods to Deep Learning Advancements in Email Security. In Proceedings of the 1st International Conference on Engineering Management, Information Technology and Intelligence (EMITI 2024); SCITEPRESS: Setúbal, Portugal, 2024; pp. 553–558. [Google Scholar]
  6. Pathak, A.; Hu, Y.C.; Mao, Z.M. Peeking into spammer behavior from a unique vantage point. In Proceedings of the 1st USENIX Workshop on Large-Scale Exploits and Emergent Threats (LEET), San Francisco, CA, USA, 15 April 2008. [Google Scholar]
  7. Kallepalli, K.; Chaudhry, U.B. Intelligent Security: Applying Artificial Intelligence to Detect Advanced Cyber Attacks. In Challenges in the IoT and Smart Environments: A Practitioners’ Guide to Security, Ethics and Criminal Threats; Springer: Cham, Switzerland, 2021. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Pachare, R.; Banarase, P.; Dhanke, P.; Dhakade, A. Advancements in Email Spam Detection: A Systematic Review of Machine Learning and Deep Learning Techniques. Int. J. Res. Appl. Sci. Eng. Technol. 2025, 13, 1908–1914. [Google Scholar] [CrossRef] [Scilit]
  9. Ma, J.; Saul, L.K.; Savage, S.; Voelker, G.M. Beyond blacklists: Learning to detect malicious web sites from suspicious URLs. In Proceedings of the 15th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Paris, France, 28 June–1 July 2009; pp. 1245–1254. [Google Scholar] [CrossRef] [Scilit]
  10. Asliyuksek, H.; Tonkal, O.; Kocaoglu, R. A Comparative Evaluation of a Multimodal Approach for Spam Email Classification Using DistilBERT and Structural Features. Electronics 2025, 14, 3855. [Google Scholar] [CrossRef] [Scilit]
  11. Metsis, V.; Androutsopoulos, I.; Paliouras, G. Spam filtering with Naive Bayes—Which Naive Bayes? In Proceedings of the 3rd Conference on Email and Anti-Spam (CEAS), Mountain View, CA, USA, 27–28 July 2006. [Google Scholar]
  12. Klimt, B.; Yang, Y. The Enron corpus: A new dataset for email classification research. In Proceedings of the 15th European Conference on Machine Learning (ECML); Springer: Berlin/Heidelberg, Germany, 2004; pp. 217–226. [Google Scholar] [CrossRef] [Scilit]
  13. Jamal, S.; Wimmer, H. An Improved Transformer-based Model for Detecting Phishing, Spam, and Ham: A Large Language Model Approach. arXiv 2023, arXiv:2311.04913. [Google Scholar]
  14. Shirvani, G.; Ghasemshirazi, S. Advancing Email Spam Detection: Leveraging Zero-Shot Learning and Large Language Models. arXiv 2025, arXiv:2505.02362. [Google Scholar]
  15. Sanh, V.; Debut, L.; Chaumond, J.; Wolf, T. DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter. arXiv 2020, arXiv:1910.01108. [Google Scholar]
  16. Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. BERT: Pre-training of deep bidirectional Transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), Minneapolis, MN, USA, 2–7 June 2019; pp. 4171–4186. [Google Scholar]
  17. Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I. Improving Language Understanding by Generative Pre-Training; Technical Report; OpenAI: San Francisco, CA, USA, 2018. [Google Scholar]
  18. Kaufman, S.; Rosset, S.; Perlich, C.; Stitelman, O. Leakage in data mining: Formulation, detection, and avoidance. ACM Trans. Knowl. Discov. Data 2012, 6, 15. [Google Scholar] [CrossRef] [Scilit]
  19. Kapoor, S.; Narayanan, A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns 2023, 4, 100804. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  20. Conneau, A.; Khandelwal, K.; Goyal, N.; Chaudhary, V.; Wenzek, G.; Guzmán, F.; Grave, E.; Ott, M.; Zettlemoyer, L.; Stoyanov, V. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL); Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 8440–8451. [Google Scholar] [CrossRef] [Scilit]
  21. Hotoğlu, E.; Sen, S.; Can, B. A Comprehensive Analysis of Adversarial Attacks against Spam Filters. arXiv 2025, arXiv:2505.03831. [Google Scholar]
  22. Reddy, M.A.S.K. Advanced Techniques in Spam Detection: A Comparative Analysis of Traditional and AI-Based Approaches. Int. J. Res. Trends Innov. 2025, 10, a598–a609. [Google Scholar]
  23. Kontsewaya, Y.; Antonov, E.; Artamonov, A. Evaluating the effectiveness of machine learning methods for spam detection. Procedia Comput. Sci. 2021, 190, 479–486. [Google Scholar] [CrossRef] [Scilit]
  24. Kim, Y. Convolutional neural networks for sentence classification. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar, 25–29 October 2014; pp. 1746–1751. [Google Scholar] [CrossRef] [Scilit]
  25. Lai, S.; Xu, L.; Liu, K.; Zhao, J. Recurrent convolutional neural networks for text classification. In Proceedings of the 29th AAAI Conference on Artificial Intelligence, Austin, TX, USA, 25–30 January 2015; pp. 2267–2273. [Google Scholar] [CrossRef] [Scilit]
  26. Zhang, Y.; Jin, R.; Zhou, Z.H. Understanding bag-of-words model: A statistical framework. Int. J. Mach. Learn. Cybern. 2010, 1, 43–52. [Google Scholar] [CrossRef] [Scilit]
  27. Mikolov, T.; Sutskever, I.; Chen, K.; Corrado, G.; Dean, J. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems; Curran Associates Inc.: Red Hook, NY, USA, 2013; Volume 26, pp. 3111–3119. [Google Scholar]
  28. Pennington, J.; Socher, R.; Manning, C.D. GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar, 25–29 October 2014; pp. 1532–1543. [Google Scholar] [CrossRef] [Scilit]
  29. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention Is All You Need. In Advances in Neural Information Processing Systems; Curran Associates Inc.: Red Hook, NY, USA, 2017; Volume 30, pp. 5998–6008. [Google Scholar]
  30. Liu, Y.; Ott, M.; Goyal, N.; Stoyanov, V. RoBERTa: A robustly optimized BERT pretraining approach. arXiv 2019, arXiv:1907.11692. [Google Scholar]
  31. Young, T.; Hazarika, D.; Poria, S.; Cambria, E. Recent trends in deep learning based natural language processing. IEEE Comput. Intell. Mag. 2018, 13, 55–75. [Google Scholar] [CrossRef] [Scilit]
  32. Ubale, K.S.; Shirsath, K.A. SpamNet: A hybrid deep learning framework for robust spam email detection using multi-modal features. Int. J. Appl. Math. 2025, 38, 670–693. [Google Scholar] [CrossRef] [Scilit]
  33. Hijji, M.; Alam, G. A Multivocal Literature Review on Growing Social Engineering Based Cyber-Attacks/Threats During the COVID-19 Pandemic: Challenges and Prospective Solutions. IEEE Access 2021, 9, 7152–7169. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Lundberg, S.M.; Lee, S.I. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems; Curran Associates Inc.: Red Hook, NY, USA, 2017; Volume 30, pp. 4765–4774. [Google Scholar]
  35. Ribeiro, M.T.; Singh, S.; Guestrin, C. “Why should I trust you?”: Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 1135–1144. [Google Scholar] [CrossRef] [Scilit]
  36. Arrieta, A.B.; Díaz-Rodríguez, N.; Del Ser, J.; Bennetot, A.; Herrera, F. Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Inf. Fusion 2020, 58, 82–115. [Google Scholar] [CrossRef] [Scilit]
  37. Jain, S.; Wallace, B.C. Attention is not explanation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), Minneapolis, MN, USA, 2–7 June 2019; pp. 3543–3556. [Google Scholar] [CrossRef] [Scilit]
  38. EPSRC. Framework for Responsible Innovation; Engineering and Physical Sciences Research Council: Swindon, UK, 2020.
  39. Gorman, K.; Bedrick, S. We need to talk about standard splits. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), Florence, Italy, 28 July–2 August 2019; pp. 2786–2791. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  40. Torralba, A.; Efros, A.A. Unbiased look at dataset bias. In Proceedings of the 2011 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Colorado Springs, CO, USA, 20–25 June 2011; pp. 1521–1528. [Google Scholar] [CrossRef] [Scilit]
  41. Geirhos, R.; Jacobsen, J.H.; Michaelis, C.; Zemel, R.; Brendel, W.; Bethge, M.; Wichmann, F.A. Shortcut learning in deep neural networks. Nat. Mach. Intell. 2020, 2, 665–673. [Google Scholar] [CrossRef] [Scilit]
  42. Recht, B.; Roelofs, R.; Schmidt, L.; Shankar, V. Do ImageNet classifiers generalize to ImageNet? In Proceedings of the 36th International Conference on Machine Learning (ICML), Long Beach, CA, USA, 9–15 June 2019; pp. 5389–5400. [Google Scholar]
  43. Cormack, G.V.; Lynam, T.R. TREC 2005 Spam Track Overview. In Proceedings of the Fourteenth Text REtrieval Conference (TREC 2005); Voorhees, E.M., Buckland, L.P., Eds.; NIST Special Publication 500-266; NIST: Gaithersburg, MD, USA, 2005. Available online: https://trec.nist.gov/pubs/trec14/papers/SPAM.OVERVIEW.pdf (accessed on 18 August 2026).
  44. Cormack, G.V. TREC 2006 Spam Track Overview. In Proceedings of the Fifteenth Text REtrieval Conference (TREC 2006); Voorhees, E.M., Buckland, L.P., Eds.; NIST Special Publication 500-272; NIST: Gaithersburg, MD, USA, 2006. Available online: https://trec.nist.gov/pubs/trec15/papers/SPAM06.OVERVIEW.pdf (accessed on 18 August 2026).
  45. Cormack, G.V. TREC 2007 Spam Track Overview. In Proceedings of the Sixteenth Text REtrieval Conference (TREC 2007); Voorhees, E.M., Buckland, L.P., Eds.; NIST Special Publication 500-274; NIST: Gaithersburg, MD, USA, 2007. Available online: https://trec.nist.gov/pubs/trec16/papers/SPAM.OVERVIEW16.pdf (accessed on 18 August 2026).
  46. Peixoto, R.F. phishing_pot: A Collection of Phishing Samples. GitHub Repository. 2024. Available online: https://github.com/rf-peixoto/phishing_pot (accessed on 18 August 2026).
  47. Baltrušaitis, T.; Ahuja, C.; Morency, L.P. Multimodal machine learning: A survey and taxonomy. IEEE Trans. Pattern Anal. Mach. Intell. 2019, 41, 423–443. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  48. Ramachandram, D.; Taylor, G.W. Deep multimodal learning: A survey on recent advances and trends. IEEE Signal Process. Mag. 2017, 34, 96–108. [Google Scholar] [CrossRef] [Scilit]
  49. Loshchilov, I.; Hutter, F. Decoupled weight decay regularization. In Proceedings of the 7th International Conference on Learning Representations (ICLR), New Orleans, LA, USA, 6–9 May 2019. [Google Scholar]
  50. Lipton, Z.C. The mythos of model interpretability: In supervised learning, the highest-performing models are often the least transparent. Commun. ACM 2018, 61, 36–43. [Google Scholar] [CrossRef] [Scilit]
  51. Goodman, B.; Flaxman, S. EU regulations on algorithmic decision-making and a “right to explanation”. AI Mag. 2017, 38, 50–57. [Google Scholar] [CrossRef] [Scilit]
  52. Mireshghallah, F.; Taram, M.; Vepakomma, P.; Singh, A.; Raskar, R.; Esmaeilzadeh, H. Privacy in deep learning: A survey. arXiv 2020, arXiv:2004.12254. [Google Scholar]
Figure 1. Framework overview. Upper band: corpus assembly, deduplication and leakage auditing, and the two split designs. Lower band: the six baseline models and the three Transformer arms—text-only, embeddings with an MLP head, and early fusion of embeddings with the structural block—all evaluated under both protocols. The second and third arms differ only in input width (768 against 781 dimensions), isolating the structural contribution with head architecture held constant.
Figure 1. Framework overview. Upper band: corpus assembly, deduplication and leakage auditing, and the two split designs. Lower band: the six baseline models and the three Transformer arms—text-only, embeddings with an MLP head, and early fusion of embeddings with the structural block—all evaluated under both protocols. The second and third arms differ only in input width (768 against 781 dimensions), isolating the structural contribution with head architecture held constant.
Futureinternet 18 00542 g001
Figure 2. Class distribution after cleaning and deduplication.
Figure 2. Class distribution after cleaning and deduplication.
Futureinternet 18 00542 g002
Figure 3. Linguistic coverage. (Left): The twelve most frequent primary languages, log-scale. (Right): Class composition of the non-English subset. Language codes: en English, de German, ru Russian, pt Portuguese, ja Japanese, fr French, es Spanish, zh Chinese, nl Dutch, pl Polish, uk Ukrainian, it Italian.
Figure 3. Linguistic coverage. (Left): The twelve most frequent primary languages, log-scale. (Right): Class composition of the non-English subset. Language codes: en English, de German, ru Russian, pt Portuguese, ja Japanese, fr French, es Spanish, zh Chinese, nl Dutch, pl Polish, uk Ukrainian, it Italian.
Futureinternet 18 00542 g003
Figure 4. Direction-adjusted univariate AUC for each structural feature over the full corpus. A value of 0.50 indicates no signal; bars are coloured by the direction in which higher values point. Only the HTML flag and the URL family separate the classes appreciably.
Figure 4. Direction-adjusted univariate AUC for each structural feature over the full corpus. A value of 0.50 indicates no signal; bars are coloured by the direction in which higher values point. Only the HTML flag and the URL family separate the classes appreciably.
Futureinternet 18 00542 g004
Figure 5. Source–label association in corpus metadata. Fields shown here are excluded from the modelling feature set.
Figure 5. Source–label association in corpus metadata. Fields shown here are excluded from the modelling feature set.
Futureinternet 18 00542 g005
Figure 6. Campaign group structure and the effect of two-stage deduplication. All retained groups are of size one, confirming that deduplication eliminated cross-partition campaign exposure before splitting. (Left): Distribution of campaign group sizes (log–log). (Middle): The twelve largest groups, coloured by class (green, legitimate; red, phishing); each label gives the source corpus followed by the group’s rank by size (#1 = largest). (Right): Emails retained (amber) and removed as near-duplicates (grey).
Figure 6. Campaign group structure and the effect of two-stage deduplication. All retained groups are of size one, confirming that deduplication eliminated cross-partition campaign exposure before splitting. (Left): Distribution of campaign group sizes (log–log). (Middle): The twelve largest groups, coloured by class (green, legitimate; red, phishing); each label gives the source corpus followed by the group’s rank by size (#1 = largest). (Right): Emails retained (amber) and removed as near-duplicates (grey).
Futureinternet 18 00542 g006
Figure 7. The two split designs. Split A measures in-distribution performance; Split B withholds an entire source from training.
Figure 7. The two split designs. Split A measures in-distribution performance; Split B withholds an entire source from training.
Futureinternet 18 00542 g007
Figure 8. Cross-source distribution shift for each candidate held-out source. (Left): Population Stability Index per structural feature. (Right): Count of features exhibiting major shift. The three TREC folds retain both classes; the honeypot source does not.
Figure 8. Cross-source distribution shift for each candidate held-out source. (Left): Population Stability Index per structural feature. (Right): Count of features exhibiting major shift. The three TREC folds retain both classes; the honeypot source does not.
Futureinternet 18 00542 g008
Figure 9. DistilBERT fine-tuning under both split designs, regularised configuration (Table 2). Training loss is measured in evaluation mode for direct comparability with validation loss. Under Split A, training ran 11 epochs with the checkpoint selected at epoch 8 (validation-loss minimum); under Split B, training ran 7 epochs with the checkpoint at epoch 4. The dotted line marks the selected epoch in each case; the bottom row reports validation-expected calibration error, which stays flat (0.039–0.046) across training under both designs.
Figure 9. DistilBERT fine-tuning under both split designs, regularised configuration (Table 2). Training loss is measured in evaluation mode for direct comparability with validation loss. Under Split A, training ran 11 epochs with the checkpoint selected at epoch 8 (validation-loss minimum); under Split B, training ran 7 epochs with the checkpoint at epoch 4. The dotted line marks the selected epoch in each case; the bottom row reports validation-expected calibration error, which stays flat (0.039–0.046) across training under both designs.
Futureinternet 18 00542 g009
Figure 10. Phishing-class F1 for all nine models under both protocols. In-distribution scores cluster tightly; cross-source scores separate the models.
Figure 10. Phishing-class F1 for all nine models under both protocols. In-distribution scores cluster tightly; cross-source scores separate the models.
Futureinternet 18 00542 g010
Figure 11. DistilBERT (text-only) confusion matrices under both split designs. Cell shading encodes the share of each true class (row-normalised); darker cells indicate a larger share.
Figure 11. DistilBERT (text-only) confusion matrices under both split designs. Cell shading encodes the share of each true class (row-normalised); darker cells indicate a larger share.
Futureinternet 18 00542 g011
Figure 12. Multimodal (fusion) confusion matrices under both split designs. Missed phishing rises from 1.23% to 8.51% of phishing instances, against 1.67% to 9.18% for the text-only model in Figure 11. Cell shading encodes the share of each true class (row-normalised); darker cells indicate a larger share.
Figure 12. Multimodal (fusion) confusion matrices under both split designs. Missed phishing rises from 1.23% to 8.51% of phishing instances, against 1.67% to 9.18% for the text-only model in Figure 11. Cell shading encodes the share of each true class (row-normalised); darker cells indicate a larger share.
Futureinternet 18 00542 g012
Figure 13. Retraining-based ablation across four structural feature configurations. All four share an identical head and differ only in structural input width. Each point is a separate model fit; false-negative counts annotated.
Figure 13. Retraining-based ablation across four structural feature configurations. All four share an identical head and differ only in structural input width. Each point is a separate model fit; false-negative counts annotated.
Futureinternet 18 00542 g013
Figure 14. Threshold sensitivity on validation data and achievable recall under a false-positive budget.
Figure 14. Threshold sensitivity on validation data and achievable recall under a false-positive budget.
Futureinternet 18 00542 g014
Figure 15. Class composition by language. Only English, German, French, and Spanish contain sufficient instances of both classes at corpus level to support full classification metrics; after partitioning, only English retains sufficient instances of both classes within the test partition.
Figure 15. Class composition by language. Only English, German, French, and Spanish contain sufficient instances of both classes at corpus level to support full classification metrics; after partitioning, only English retains sufficient instances of both classes within the test partition.
Futureinternet 18 00542 g015
Figure 16. Global importance: text embedding block against individual structural features. The dark bar shows the aggregated text embedding block; amber bars show individual structural features, which are small at this scale.
Figure 16. Global importance: text embedding block against individual structural features. The dark bar shows the aggregated text embedding block; amber bars show individual structural features, which are small at this scale.
Futureinternet 18 00542 g016
Figure 17. Local attributions for the same two instances under both explainers. (Top): SHAP. (Bottom): LIME. (Left): Correctly classified phishing instance. (Right): Misclassified instance. The two agree on the dominant text-embedding term and diverge on the structural block. Amber bars push the prediction towards phishing and red bars towards legitimate.
Figure 17. Local attributions for the same two instances under both explainers. (Top): SHAP. (Bottom): LIME. (Left): Correctly classified phishing instance. (Right): Misclassified instance. The two agree on the dominant text-embedding term and diverge on the structural block. Amber bars push the prediction towards phishing and red bars towards legitimate.
Futureinternet 18 00542 g017
Table 1. Representative spam detection studies and gaps addressed by the proposed framework.
Table 1. Representative spam detection studies and gaps addressed by the proposed framework.
StudyMethodLangs.ExplainabilityCross-Source Eval.
Kontsewaya et al. [23]SVM, NB, LR1NoneNo
Kim [24]CNN1NoneNo
Lai et al. [25]Recurrent CNN1NoneNo
Jamal & Wimmer [13]Fine-tuned BERT1AttentionNo
Asliyuksek et al. [10]DistilBERT + structural1NoneNo
Ubale & Shirsath [32]SpamNet (DL)1NoneNo
Shirvani et al. [14]Zero-shot FLAN-T51PartialNo
This paperDistilBERT + MLP91 (4 eval.)SHAP + LIMEYes
“Langs.” gives the number of primary languages present in the corpus, with the number supporting full classification metrics in parentheses.
Table 2. Training configuration for all model families.
Table 2. Training configuration for all model families.
SettingValue
Transformer encoder
Checkpointdistilbert-base-uncased (67.0 M parameters per encoder)
Max sequence length256 tokens
Batch size32
Learning rate 2 × 10 − 5
Epoch budget policyStart 10, extend by 5 to ceiling 25; stop on early stopping (patience 3 on validation loss) or plateau (range < 0.002 over last 4 epochs); checkpoint = epoch of lowest validation loss
Epochs run (A, B)11, 7 (best epoch: 8, 4)—both terminated by early stopping
Regularisation (weight decay/
label smoothing/dropout)
0.05/0.1/0.2 (hidden, attn), 0.3 (classifier)
Frozen parametersEmbeddings + lowest 2 of 6 blocks (28.9 M trainable)
Embedding poolingMean over non-padding tokens of the final layer
Shared optimisation
OptimiserAdamW, decoupled weight decay (0.05 DistilBERT, 0.01 elsewhere)
LR scheduleLinear warmup over 10% of steps, then linear decay to zero
Gradient clippingMax-norm 1.0
Mixed precisionfp16
LossCross-entropy with class weights n / ( 2 n c ) from each design’s training partition
Imbalance handlingClass weighting only; no oversampling, undersampling, or synthetic data
Classification heads
MLP heads(256, 64) ReLU, dropout 0.2, AdamW lr 1 × 10 − 3 , batch 256, early stopping patience 10 (max 100 epochs)
CNNEmbedding dim 128 trained from scratch, filters 3/4/5 × 100, max-pool, dropout 0.5, lr 1 × 10 − 3 , unchanged val-F1-based adaptive budget policy
Classical baselines
TF-IDF1–2 g, max 50,000 features, min_df 3, max_df 0.95, sublinear tf, first 5000 characters
Logistic regression/SVM C = 1.0 , class_weight balanced
Naive BayesMultinomial, α = 0.1
Random forest300 trees, min_samples_leaf 2, class_weight balanced, 13 features; refit per seed
Environment
Seeds42, 43, 44, 45, 46
HardwareNVIDIA Tesla T4
SoftwarePython 3.13.15, PyTorch 2.11.0+cu128, Transformers 5.16.1, scikit-learn 1.6.1, NumPy 2.1.3, pandas 2.2.3, SciPy 1.16.3
Table 3. The conditional contribution of structural fusion. Paired bootstrap differences with 95% confidence intervals and McNemar tests and the fusion model against each text-only variant under both protocols.
Table 3. The conditional contribution of structural fusion. Paired bootstrap differences with 95% confidence intervals and McNemar tests and the fusion model against each text-only variant under both protocols.
ComparisonSplit Δ F195% CI Δ RecallMcNemar p
Fusion − ablated headA+0.0014[0.0005, 0.0023]+0.00300.0037
Fusion − text-onlyA+0.0001[−0.0012, 0.0012]+0.00451.000
Fusion − ablated headB+0.0078[0.0036, 0.0120]+0.0107<0.001
Fusion − text-onlyB+0.0044[−0.0001, 0.0092]+0.00670.065
Table 4. Generalisation gap by model. F1 is for the phishing class; the gap is the difference between the in-distribution and cross-source protocols. Rows are ordered by gap, with the smallest first.
Table 4. Generalisation gap by model. F1 is for the phishing class; the gap is the difference between the in-distribution and cross-source protocols. Rows are ordered by gap, with the smallest first.
ModelF1 (A)F1 (B)GapRecall (B)AUC (B)
Structural only (RF)0.92560.91950.00610.90670.9743
Structural + lexical (LR)0.98030.95100.02930.92920.9925
Lexical only (TF-IDF + LR)0.97970.94390.03580.91760.9924
DistilBERT + Structural (Fusion)0.98780.94880.03900.91490.9844
DistilBERT (text-only)0.98770.94440.04330.90820.9904
DistilBERT embeddings + MLP0.98640.94100.04540.90420.9934
CNN (word-embedding, from scratch)0.97660.91890.05770.87020.9907
SVM (TF-IDF, linear)0.98570.91880.06690.86550.9910
Naive Bayes (TF-IDF)0.96130.86410.09720.77390.9862
Table 5. Phishing-class F1 for every model across five seeds under both split designs. Models marked † are deterministic given the data or were fine-tuned once with predictions reused across seeds; their zero standard deviation reflects that determinism or reuse, not demonstrated stability under reinitialisation. Only the random forest and the two neural heads were genuinely refit per seed.
Table 5. Phishing-class F1 for every model across five seeds under both split designs. Models marked † are deterministic given the data or were fine-tuned once with predictions reused across seeds; their zero standard deviation reflects that determinism or reuse, not demonstrated stability under reinitialisation. Only the random forest and the two neural heads were genuinely refit per seed.
ModelSplit4243444546MeanSD
Structural only (RF)A0.92560.92640.92610.92560.92680.92610.0005
Structural only (RF)B0.91950.91910.91900.91720.91940.91880.0009
Naive Bayes †A0.96130.96130.96130.96130.96130.96130.000
Naive Bayes †B0.86410.86410.86410.86410.86410.86410.000
SVM †A0.98570.98570.98570.98570.98570.98570.000
SVM †B0.91880.91880.91880.91880.91880.91880.000
Lexical only †A0.97970.97970.97970.97970.97970.97970.000
Lexical only †B0.94390.94390.94390.94390.94390.94390.000
Structural + lexical †A0.98030.98030.98030.98030.98030.98030.000
Structural + lexical †B0.95100.95100.95100.95100.95100.95100.000
CNN †A0.97660.97660.97660.97660.97660.97660.000
CNN †B0.91890.91890.91890.91890.91890.91890.000
DistilBERT (text-only) †A0.98770.98770.98770.98770.98770.98770.000
DistilBERT (text-only) †B0.94440.94440.94440.94440.94440.94440.000
DistilBERT + MLPA0.98640.98670.98640.98630.98660.98650.0002
DistilBERT + MLPB0.94100.94330.94070.95070.94710.94450.0043
FusionA0.98780.98750.98810.98850.98760.98790.0004
FusionB0.94880.94240.95130.93230.94820.94460.0076
Table 6. Per-language evaluation on the Split A test partition. Full metrics are reportable for English only; all other languages are reported as phishing detection rate against a majority-class baseline. Slovenian, Finnish, Persian, and Turkish are pooled with the sub-threshold languages.
Table 6. Per-language evaluation on the Split A test partition. Full metrics are reportable for English only; all other languages are reported as phishing detection rate against a majority-class baseline. Slovenian, Finnish, Persian, and Turkish are pooled with the sub-threshold languages.
LanguagenLegit.Phish.Det. RateDet. RateMajority
(Text-Only)(Fusion)Baseline
English18,75711,18075770.98170.98640.5960
German285122730.98530.99270.9579
Portuguese13921371.00001.00000.9856
Russian13811371.00001.00000.9928
French958870.98851.00000.9158
Japanese832811.00001.00000.9759
Spanish793761.00001.00000.9620
Chinese741731.00001.00000.9865
Dutch460461.00001.00001.0000
Polish270271.00001.00001.0000
Ukrainian232211.00001.00000.9130
Italian211201.00001.00000.9524
Other (pooled)17571680.98800.98800.9600
For English, full metrics are text-only accuracy 0.9892, F1 0.9866; fusion accuracy 0.9893, F1 0.9867.
Table 7. Comparison with representative spam and phishing detection systems.
Table 7. Comparison with representative spam and phishing detection systems.
StudyMethodCorpusAcc.Explain.Cross-Source
Jamal & Wimmer [13]Fine-tuned BERTCustom0.9892AttentionNo
Asliyuksek et al. [10]DistilBERT + structuralCombined (81,586)0.9962NoneNo
Kontsewaya et al. [23]SVM, NB, LRMultiple∼0.990NoneNo
Ubale & Shirsath [32]SpamNet (multimodal DL)Balanced (6000)0.9881NoneNo
Shirvani et al. [14]Zero-shot FLAN-T5Customn/rPartialNo
This paperDistilBERT + MLPMEPC (99,707)0.9894SHAP + LIMEYes
Accuracies are in-distribution and derive from different corpora and partitioning schemes; the column is reported for completeness rather than ranking. Cross-source results for this paper are given in Section 4.3.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Xhafaj, G.; Chaudhry, U.B.; Jahankhani, H. Hybrid Transformer Architecture for Context-Aware Spam Email Classification. Future Internet 2026, 18, 542. https://doi.org/10.3390/fi18100542

AMA Style

Xhafaj G, Chaudhry UB, Jahankhani H. Hybrid Transformer Architecture for Context-Aware Spam Email Classification. Future Internet. 2026; 18(10):542. https://doi.org/10.3390/fi18100542

Chicago/Turabian Style

Xhafaj, Geldi, Umair B. Chaudhry, and Hamid Jahankhani. 2026. "Hybrid Transformer Architecture for Context-Aware Spam Email Classification" Future Internet 18, no. 10: 542. https://doi.org/10.3390/fi18100542

APA Style

Xhafaj, G., Chaudhry, U. B., & Jahankhani, H. (2026). Hybrid Transformer Architecture for Context-Aware Spam Email Classification. Future Internet, 18(10), 542. https://doi.org/10.3390/fi18100542

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Article metric data becomes available approximately 24 hours after publication online.
Back to TopTop