Next Article in Journal
Multi-Agent GenAI for Harm Reduction: A Psychologist’s (or Social Scientist’s) Tutorial on Classifying Harmful Workplace Language
Previous Article in Journal
Distribution of Cognitive Load and Related Debates in Students Interacting with an Augmented Reality Application for History Learning
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Adversarial Training and Differential Privacy-Style Noise Injection for Privacy-Preserving Vertical Federated Learning

by
Nureni Ayofe Azeez
1,*,
Oluwatobi Sunday Malomo
1,
Omotolani Mary Okerinde
1,
Abdullateef Akorede Ademoye
1,
Damilola Seun Aaron
1,
Charles Van Der Vyver
2 and
Chijioke Erasmus Ogbonna
3
1
Department of Cybersecurity and Software Engineering, Faculty of Computing and Informatics, University of Lagos, Lagos 101017, Nigeria
2
School of Computer Science and Information Systems, Faculty of Natural and Agricultural Sciences, Vaal Triangle Campus, North-West University, Vanderbijlpark 1900, South Africa
3
Department of Mechatronics Engineering, Faculty of Engineering, Federal University of Technology, Owerri 460114, Nigeria
*
Author to whom correspondence should be addressed.
Informatics 2026, 13(8), 127; https://doi.org/10.3390/informatics13080127
Submission received: 3 April 2026 / Revised: 10 July 2026 / Accepted: 20 July 2026 / Published: 9 August 2026

Abstract

The adoption of federated learning (FL) has been on the rise in recent years due to the decentralized approach to data handling. Vertical federated learning is a type of FL that allows different parties to train shared models on complementary feature spaces without the direct exchange of data. However, the gradients these parties exchange can inadvertently carry sensitive information. Adversaries exploit this leakage to mount label inference attacks (LIAs) and adversarial attacks. To curb this, defense mechanisms have been deployed, but most of them either trade robustness for privacy and model utility or vice versa. This study addresses this gap by introducing an improved defense mechanism that combines adversarial training (to harden the model against adversarial perturbations) and differential-privacy-style noise injection (aimed at restoring the label privacy weakened by adversarial training) to collectively enhance the robustness of the existing KD k defense mechanism with marginal model utility trade-off. Instead of relying on heavy encryption or post-processing techniques, it builds privacy directly into the learning dynamics of the model. It was evaluated using five publicly available datasets spanning three data modalities with the proposed mechanism achieving competitive near-baseline accuracy while significantly reducing label-inference success. Under FGSM-based adversarial evaluation, the robustness gap of this mechanism was found to be approximately 1% compared to the 36% robustness gap of the existing KD k mechanism. The Privacy Leakage Index (PLI) reached 81.32%, 96.08%, 82.41%, 86.68% and 73.88% for CIFAR-10, CIFAR-100, CINIC-10, Yahoo! Answers and Criteo datasets, respectively. The results suggest that robustness and privacy security objectives can coexist to secure VFL with minimal effect on model accuracy.

1. Introduction

The integration of artificial intelligence (AI) techniques across critical sectors, such as cybersecurity, finance, healthcare, education, e-commerce, real estate, IoTs, government, and media has raised data privacy and security concerns in recent years [1,2,3]. Traditionally, machine learning (ML) techniques require having the training data in centralized storage; however, this approach has proven to come with a high level of risk, especially concerning data exfiltration, unauthorized access, and regulatory violations under data protection laws like the General Data Protection Regulation (GDPR) [4,5].
To address these privacy issues, federated learning (FL) was introduced by Google researchers in 2016 as a machine learning approach in which the clients (devices) train a model collaboratively under a central server without the need to share their training data [6]. The overall idea was to develop an approach whereby, instead of bringing data to code as we have in traditional ML, the code is brought to the data in a privacy-preserving learning manner [7]. This approach has been widely adopted in industries that operate on sensitive data, such as healthcare (disease prediction) and finance (fraud detection), among others [8,9].
In terms of operation Chakraborty, et al., (2025) [10] submitted that an FL system can either be centralized or decentralized (as proposed in [11]). While the former possesses an aggregator that oversees the coordination of the entire learning process before aggregating them to form a global model, the latter gives room for direct gradient exchange between the contributing devices. This ensures that sensitive user data remains localized, thereby reducing the risk of exposure to external threats.
When considering the data partition strategy, there are two major categories of FL. Horizontal federated learning (HFL) is deployed when an organization has the same feature space but different users, such as a chain of hospitals training a model on patient health records across their hospital branches, state or country-wide. The VFL, on the other hand, is used when building an FL system with the same users but across different feature spaces. For instance, a telco and a bank can team up to build a fraud detection model whereby the bank presents the financial transaction data of the users, while the telco presents their call history and device information. Since regulations prevent direct data sharing between these parties, VFL enables them to jointly train ML models in a privacy-preserving environment, promoting collaboration between non-competing organizations by partitioning training data vertically. Hence, the goal is to ensure that one of the organizations possesses the labels with some of the features of samples, while the remaining participants provide additional feature information of the same sample space without disclosing data directly to other participants.
However, the advancement in AI technologies has further equipped adversaries with a sophisticated means of sabotaging VFL environments, such as the label inference attack, which aims to reconstruct private labels by analyzing gradient updates exchanged by the involved parties during training [12]. For industries where sensitive data is exchanged, this challenge ultimately leaves VFL systems at risk of data breach. Over the years, some privacy-preserving defense mechanisms based on differential privacy, homomorphic encryption, and secure multi-party computation have been proposed to mitigate LIAs in VFL environments. These mechanisms have focused on gaining either improved privacy or robustness but not both at the same time, representing a difficult security trade-off.

1.1. Problem Statement

The accuracy with which adversaries extract private label information in VFL systems is a major challenge today [12]. By exploiting the mathematical properties of gradient descent, attackers can infer sensitive labels with great precision without being noticed. The use of privacy-preserving mechanisms such as knowledge distillation and data anonymization provide partial baselines for protecting VFL systems against adversarial attacks and LIAs. However, these defense mechanisms are static, failing to address an emerging but even more deadly threat model, adversarial attacks. Adversarial training (AT) presents a standard defense against adversarial attacks albeit with a private label leakage problem. Therefore, VFL systems defended with these standard defense mechanisms independently are often subjected to the difficult privacy–robustness or privacy–utility trade-off whereby the robustness of a model against adversarial attacks is at the cost of its privacy or otherwise. Hence, there is a need for an improved defense mechanism that will simultaneously offer improved robustness against adversarial and privacy threat models without sacrificing label privacy gains.

1.2. Objectives of the Study

This study seeks to improve the robustness of current defense mechanisms against evolving adversarial and privacy threat models.
The major contributions of the proposed defense mechanism are as follows:
  • To analyze the vulnerabilities of standard VFL systems to both adversarial attacks and label inference attacks.
  • To show how adversarial training (AT) improves adversarial robustness but measurably weakens label privacy, quantifying this trade-off across modalities.
  • To introduce a differential-privacy-style calibrated noise stage that restores the label privacy weakened by AT without sacrificing the robustness gains.
  • To demonstrate empirically, across five datasets spanning three modalities, that adversarial robustness and label privacy are complementary rather than competing objectives in VFL environments.
The remaining part of this paper is organized as follows; in Section 2, related works are reviewed to identify the research gap. Section 3 and Section 4 present the methodology and results, respectively, while the conclusion and recommendations for further studies are highlighted in Section 5.

2. Related Works

Traditional machine learning techniques involve the collection of data at a central server; as such, there exists a singular point of vulnerability, and once exploited, the model can be deployed for malicious intentions. This major privacy risk is what the federated learning method seeks to address by training machine learning models in a distributed manner [12] such that the central server, also called an aggregator, does not hold all of the data; instead, it is responsible for aggregating the model gradients from all of the parties in the FL system.
This approach offers a secure alternative for privacy-preserving ML with respect to data sharing. Its relevance is particularly notable in highly regulated sectors such as finance, healthcare, and security, where direct data sharing is restricted due to stringent legal regulations [8]. This underscores the significant potential of FL in revolutionizing data-driven industries without compromising user privacy. The following subsections will provide a detailed review of literature covering background information about the concept of FL, its classification and privacy-preserving solutions.

2.1. Vertical Federated Learning

Vertical federated learning is a type of FL in which the data is partitioned vertically, as previously discussed in the opening section of this study. In principle, it is adopted when parties from diverse feature spaces collaborate to work on an ML project involving the same set of users without direct training data exchange [13]. VFL has continued to gain widespread popularity as a practical solution to real-world problems for organizations seeking to work collaboratively on mutual user data while preserving privacy [11,14].

2.2. Challenges in the VFL Environment

Privacy has remained a source of concern in VFL systems due to the possibility of inference leaks that expose shared samples to unintentional and/or intentional inferences as embeddings and gradients are exchanged between parties [15,16].
In the absence of feedback for model convergence, ground-truth labels become inaccessible to feature-holding parties in the VFL environment. Limited label accessibility is an issue that can degrade the efficiency of a VFL system; therefore, methods such as label propagation and gradient label sharing have been deployed to ensure all parties have label access when needed.
Furthermore, the advancements made with generative adversarial networks (GANs) today have increased the possibilities for data inference on each FL party [17]. Hence, protecting the privacy of the labels owned by each client in a VFL is important because they are highly sensitive [18]; hence, extra security measures are needed to protect the labels and sensitive data owned by each party in the system. In [19], it is suggested that unintentional data leakage and reconstruction through inference occur when gradients transmitted by clients inadvertently reveal sensitive information at the central server, where the gradient aggregation takes place.

2.3. Label Inference Attacks in VFL

Label inference attacks are threat models based on the output labels of a VFL system, thus allowing adversaries to infer important features by simply observing them. This connotes that an attacker requires direct access to the FL system’s gradient information to perpetrate their malicious intents. Adversaries make inferences by intent, through the observation of the corresponding output labels of different samples [20]. The goal of LIAs is to exploit the mutual information among labels, gradients, and embeddings. For example, the diagnostic results of a patient or the loan default records of an individual are sensitive information that only authorized institutions should have access to [21,22]. Sometimes, even an honest-but-curious party can expose data to an attacker who then infers important labels held by an active party using the information they accumulate during training or inference. This is often achieved by following the protocol passively under the honest-but-curious security assumptions or actively by tampering with the protocol under the malicious assumptions [13].
Fu et al., [5] described three categories of LIAs. Passive attack involves the fine-tuning of the bottom model to infer labels in a semi-supervised manner. This is made possible with the help of some auxiliary labelled data. In an active attack, the attacker attempts to scale up the learning rate of the bottom models during the training phase. This enforces the overreliance of the top model on the bottom model, which in turn boosts the accuracy of label inference. Direct attack involves the analysis of gradient signs from the server. Lastly, perturbed attack involves adversarially perturbing embeddings to deliberately push model outputs towards a specific label. In any of these cases, the attacker extracts information on labels by exploiting the feedback received from the loss function estimated on the top model, before performing the backpropagation on the local model.

2.4. Adversarial Machine Learning and Adaptive Defenses

The significant growth in the field of artificial intelligence and its increasing fusion with cybersecurity has drastically changed cyber operations in terms of the sophistication of both offense and defense [23]. Furthermore, [14] posited that the adoption of ML-based malware detection systems can help security teams to efficiently classify unknown samples by learning from known malware characteristics [24]. The goal of this threat model is to exploit the injection of small, targeted perturbations into training data to influence model classification or predictions. According to [25,26], adversarial attacks not only threaten the effectiveness of ML-based cybersecurity systems but can also lead to severe security breakdowns once the vulnerabilities they create are exploited. Adversarial machine learning (AML) examines how malicious third-party activities impact machine learning models, resulting in the misclassification of input, distortion of model accuracy, or data exfiltration. Hence, a defense mechanism is said to be adversarial when it explicitly models an attack, has an adversarial objective, that is, minimizing classification accuracy while maximizing resistance to LIAs, and its training is adaptive (involves a feedback loop).

2.5. Review of Related Works

2.5.1. Privacy and Label Inference Defenses in VFL

Arazzi et al., [12] addressed the issue of LIAs in vertical federated learning environments with respect to the reconstruction of private labels from leaked gradient updates. The authors proposed a novel KD k framework by combining knowledge distillation with k-anonymity (KDk) on the softened labels exchanged during training. This modifies the active party by integrating the distilled soft labels with the k-anonymity processing step to obtain a group of k -most probable soft labels for each item in place of a single hard label. This approach introduces a level of uncertainty that further makes it difficult for attackers to perform label inference in the VFL system. In experiments on the CIFAR-10, CIFAR-100, and CINIC-10 datasets, the KD k framework proved efficient, reducing attack accuracy by over 60%. In terms of utility, the model remained largely unaltered; however, it lacked the robustness needed to mitigate adversarial attacks [27].
Yan et al., [28] proposed label-anonymized defense with substitution gradient (LADSG), integrating anomaly detection, gradient substitution, and label anonymization together to mitigate label leakage. The mechanism was evaluated on six real-world datasets from healthcare, finance, and language, among other sectors. The results demonstrated a reduction in the success rates of LIAs by approximately 60% without incurring prohibitive computational overheads. LADSG performed well in untrusted or partially adversarial settings, making it deployable in practical FL systems due to its low integration cost and component-wise modularity. However, it proved to be vulnerable to adaptive adversaries since an attacker can use auxiliary data to construct inference models to weaken anonymization and distillation.
Wang et al., [29] proposed a dispersed training mechanism that breaks the correlations between training data and the bottom model using secret sharing. This prevents attackers from deducing the feature representation of labels from the bottom model, even if the gradients are exposed. This is achieved by integrating a specific method for model aggregation with the linearity of secret sharing schemes to preserve the accuracy of the training. The CIFAR-10, CINIC-10, and BCW datasets were used to evaluate the framework, with Top-1 Accuracy selected as the performance indicator. The experiments show that dispersed training can effectively prevent label inference attacks in VFL systems; however, this comes at a cost, as the experiment revealed a discrepancy in the accuracy and integrity of the original architecture [30].
Ding et al., [31] developed a threshold-based detection system that monitors deviation in gradient distributions and flags inference attacks. The mechanism compares the difference in features amongst training rounds and raises a notification when a given parameter is exceeded. Moreover, the authors created six taxonomies of threat models to examine the cases of the local inference attacks (LIAs) and membership inference attacks (MIAs) in different a priori conditions. These classifications help detect the attacks and allow the measurement of the mechanism’s effectiveness in terms of the collective effect on the FL system. In this regard, the authors evaluated the variation in such metrics as the attack success rate (ASR), F1-score and robust accuracy. However, the system was not designed to identify the threats aimed at compromising the functional integrity of the model, such as the injection of adversarial examples.
Across these approaches, the emphasis is uniformly on reducing label or membership leakage; none incorporates a defense against adversarial perturbation of the model.

2.5.2. Adversarial Robustness Techniques in VFL

The application of adversarial training and robustness technqiues in VFL settings has remained largely unexplored compared to centralized machine learning as established earlier in Section 2.4. Existing VFL defenses concentrate on label and gradient privacy and do not harden the model against adversarial perturbation, leaving adversarial robustness and its interaction with label privacy as an open problem in the VFL setting. This study addresses this gap.

2.5.3. Property-Level and Other Threats

Despite their prevalence in recent years, studies have shown that record-level privacy concerns are not the only variant of attacks on VFL systems. Bai et al., [32] proved this by examining how adversarial parties exploit the training set of a victim party to identify the global distribution information of a target property. The study revealed that the observation of a target property’s L p -norm can implicitly reflect its fraction in a training set. To tackle this problem, the authors proposed the novel ProVFL framework that involves learning the implied relationship between L p -norm distributions and their fractions through the introduction of a distribution comparison module that involves creating intermediate-result populations across different proportions. The other part of the framework involves the theoretical analysis of factors contributing to the effectiveness of an attack before developing a correlation augmentation module by amplifying property information leakage with the use of label replacement and model refinement. Through various experiments, the attack was found to achieve inferences with estimation errors as low as 1%, further highlighting the immediate threat of property information leakage in VFL systems.
With advancements in technology, particularly AI, increasingly sophisticated techniques are used by adversaries to degrade defense systems. It can be deduced from the above that most of the existing defense mechanisms trade off robustness for privacy and model utility or vice versa. However, these security objectives should not be seen as competing objectives [27]. To address this limitation, this paper presents an integration of adversarial training and differential privacy techniques to achieve both robustness and privacy security objectives concurrently with minimal model utility gap. However, recent advances in adversarial robustness and differential privacy remain largely unexplored in privacy-preserving VFL. As highlighted in Table 1, addressing both the privacy and robustness needs of privacy-preserving VFL systems remains an open research gap addressed in this study.

3. Methodology

This section explores the approach employed to address the problem, presents the algorithm for the improved defense mechanism, and presents the implementation procedure.

3.1. Framework for the Improved Defense Mechanism

The defense mechanism was designed in a two-party VFL environment. The active party holds the private labels and a subset of the features, while the passive party holds another subset of the features. For model generalizability, the benchmark datasets were selected from three domains, namely, image classification, textual/natural language processing, and binary classification. The implementation is a layered VFL pipeline (Figure 1), with each stage built on the previous one to achieve the robustness and privacy security objectives. This approach allows for a fair comparative analysis of the trade-offs that occur at each step of the implementation.
The baseline or original architecture (OA) is trained on clean data to establish a baseline for model utility and initial vulnerability to both adversarial and privacy attacks. KD k added two foundational defense techniques: knowledge distillation, to enhance performance as the model learns directly from OA’s soft labels, and a data anonymization technique ( k -anonymity) to obfuscate labels. The introduction of adversarial training (AT) hardens the model against adversarial attacks.
The last step of the multi-stage implementation represents the novel contribution of this research regarding the popular privacy–robustness trade-off in privacy-preserving FL research. By analyzing the privacy vulnerabilities created by the KD k +AT model, differential privacy (DP) emerges as a targeted countermeasure for restoring label privacy while preserving the robustness gains. This is achieved through the integration of calibrated Gaussian noise applied to the exchanged embeddings (forward pass) and their gradients (backward pass).
The robustness and privacy gains of the defense mechanism are evaluated by exposing it to adversarial attack during training and label inference attacks (LIAs) post-training. This ensures the objectives of the work are not merely assumed but empirically validated, with the results converging into a comparative evaluation framework where robustness to adversarial attacks (robust accuracy), model utility (clean accuracy), and privacy (ASR and PLI) are jointly analyzed using the five datasets. This structure allows a fair comparison of the improved defense mechanism, highlighting the fact that robustness and privacy are, in fact, complementary security objectives in VFL.

3.2. Dataset Description and Preprocessing

The benchmark datasets were selected across image classification, textual/NLP, and binary classification domains to improve the generalizability of the improved defense mechanism. This is a common practice in federated learning literature, as it ensures cross-domain robustness can be properly evaluated. All of the datasets were preprocessed to ensure correctness and suitability for the VFL setting, an important step for cross-modality comparison since the datasets are of varying types. The structure and statistical model convergence and robust evaluation of neural networks are quite sensitive to input distributions; hence, any deviation shifts can be problematic. Table 2 presents an overview of the datasets used in terms of their task type, domain, size and evaluation focus.

3.2.1. Vision Datasets

The standard computer vision dataset preprocessing was followed for the image datasets (CIFAR-10, CIFAR-100, and CINIC-10). The images were converted into tensors, which were further normalized using per-channel mean and a 0.5 standard deviation. This process helps to facilitate stable propagation of gradients during model training and increases the rate of convergence as it maps pixel intensities to the range of [−1. 1]. Since the defense mechanism is to be built in a VFL setting, there is a need to split the vision datasets along the channel dimension. Hence, while the passive party received both the green and blue (GB) channels, the active party retained the red (R) channel. Each image remained incomplete individually since the tensors only preserved their complementary information. This measure was necessary to prevent the independent reconstruction of the full input by either of the active or passive parties. Lastly, there was a need to have a data structure that is consistent for all vision, text, and tabular modalities; hence, data loaders were defined to return batches in the form X a c t i v e ,   X p a s s i v e , y .

3.2.2. Textual Dataset

The Yahoo! Answers corpus has an entirely different structure as it contains raw natural language texts rather than numeric features. Since the records are made up of the pairing of categorical topic labels with question titles and question content, the preprocessing followed the normal text cleaning approach by normalizing whitespace, removing special characters, and lowering the cases of all records, after which, the pre-trained a l l − m p n e t − b a s e − v 2 was used to convert both title and content into dense vector representations within a 768-dimensional embedding space. To ensure suitability for a real-world VFL configuration, where two distinguishable organizations can hold complementary textual views of an entity (a news headline versus the news in detail), the active party received the title embeddings while the content embeddings were assigned to the passive party. Lastly, to mitigate magnitude imbalance, the embeddings were L 2 -normalized, to ensure smoother optimization during training.

3.2.3. Tabular Dataset

The Criteo Click-Through-Rate (CTR) dataset was initially downsampled to its dense (continuous) features because including categorical features would lead to high-cardinality encoding, and the size of its raw form is quite enormous. The parquet downsample (train/test split) was standardized using z-score normalization, computed from training-set statistics. To prevent larger-scale features from dominating the optimization process, the data is centered around zero with unit variance. The normalized features are then vertically split, with the active party assigned six features and the passive party obtaining seven, creating non-overlapping yet correlated views of the same samples.
Lastly, with CIFAR-10/100, CINIC-10, Yahoo! Answers, and Criteo CTR, the benchmark dataset used in this study perfectly balances reproducibility, scalability, and domain transferability. This is very crucial for the evaluation of the defense mechanism, ensuring it does not overfit to a single modality.

3.3. Model Architecture Design

The experimental defense mechanism used a series of modular neural architectures, whereby each of the parties in the system holds a fraction, or subset of the overall feature space. These parties then process their private labels through their models, known as the bottom model. The top model is responsible for aggregating intermediate embeddings from the bottom models to produce the final prediction. It ensures that all parties maintain their data locally while collaborating in end-to-end learning. On the other hand, the bottom model serves as the representation extractor for each party. It is responsible for independently processing and transforming each party’s local features into intermediate embeddings, which are then passed to the top model.
The bottom models serve as feature extractors for each party before forwarding their output signals to the top model for subsequent aggregation. In the process of backpropagation, the gradients flow from the top model back to each bottom model for joint optimization without the data exchange. An overview of network architecture used for each dataset is presented in Table 3, with the rationale for each explained below.
Vision Datasets: To balance data representation with the computational efficiency needed in a VFL environment, a custom, lightweight architecture was deployed. The bottom model is a shallow VGG-style convolutional neural network, which involves stacking simple 3 × 3 convolutional layers to extract features. This was opted for to minimize communication overhead and computation time, which are critical constraints in a federated environment. The architecture consists of sequential blocks of convolutional, BatchNorm, and ReLU layers, followed by MaxPool layers, an adaptive average pooling layer, and a fully connected embedding head. It captures the spatial features from the image channels without the parameter overhead of deeper networks. The top model is a three-layer multi-layer perceptron (MLP), which receives the processed feature embeddings from the two bottom models and fuses them to perform the final classification. The MLP can model complex, non-linear interactions between feature sets provided by both parties, thereby making it the ideal model for the feature fusion task.

3.3.1. Training Strategy

The model training process was designed to allow some level of control over experimental variables with all models trained under identical optimization and hardware conditions and data partitioning (Table 4). This gives the leverage to attribute noticeable differences in both performance and privacy strictly to the learning framework adopted for each model [33].
Once the models have been defined, actual experimentation was conducted on Google Colab, which gives access to the much faster A100 (40 GB) GPUs. To promote smooth convergence during the implementation runs, the Adam optimizer ( l r =   1   ×   10 − 3 ) and a cosine-annealing schedule were adopted. To mitigate overconfidence, the label smoothing ( ε   =   0.05 for vision; 0.05 − 0.1 for text/tabular) was set while the gradient clipping ( ‖ g ‖ 2   ≤   1.0 ) prevented numerical instability during adversarial training. The batch sizes were set to 64 and 128 for text/tabular and vision tasks, respectively, while the epoch limit was predefined at 10 to 50 in the CONFIG declaration.

3.3.2. Training Loop for VFL

The training is sectioned into three stages per epoch:
  • Forward Pass: The active and passive parties both process their local inputs and produce embeddings through their respective bottom models.
  • Fusion and Prediction: The embeddings are collected and concatenated, then passed to the top model to generate global logits.
  • Backward Pass: The gradients of the global loss are backpropagated to each party to update their local parameter.
This structure ensures that parties only exchange gradients and embeddings to preserve data sovereignty. Since the tensors exchanged between parties have fixed dimensionality, per-round communication cost is determined analytically from the passive-party embedding width, the number of training samples, and the number of epochs, rather than instrumented at runtime. As AT and noise injection are applied locally, they leave the exchanged tensor shapes unchanged and add no communication overhead relative to the baseline. Per-stage communication volumes and measured training times are reported in Table 11.

3.3.3. Knowledge Distillation + k -Anonymity (KD k )

Once the OA training is completed, the logits are softened using a temperature-scaled s o f t m a x   ( τ   ∈   { 3 ,   4 } ) . Hence, to achieve k -anonymity, the top k probabilities ( k   ∈   { 3 ,   5 ,   8 ,   10 } ) were retained and are subjected to another round of renormalization. The KD k model optimized a hybrid loss combining cross-entropy and Kullback–Leibler divergence: where λ = 0.7 balances both distillation and accuracy strength. This configuration allows the student to approximate the teacher’s decision boundaries while reducing label leakage.
The pseudocode for the knowledge distillation with k-anonymity (KDk) implementation is shown in Algorithm 1.
Algorithm 1. KD k (Knowledge Distillation + k -Anonymity, Student Training).
Input: Training data D_train, Test data D_test, Teacher_Soft_Labels
Output: Trained student model M_student, Accuracy ACC_KDk
1. Initialize Model:  M _ s t u d e n t   =   ( B o t t o m ,   T o p )
2. For each epoch in EPOCHS do:
      For each batch  ( x ,   y ) in D_train do:
          a. Student forward pass:  s _ l o g i t s   =   M _ s t u d e n t ( x )
          b. Teacher forward pass (frozen):  t _ l o g i t s   =   M _ t e a c h e r ( x )
          c. Apply temperature scaling: t _ s o f t   =   S o f t m a x ( t _ l o g i t s   /   T )
          d. Apply k -anonymity:  t _ s o f t _ k   =   K A n o n y m i t y ( t _ s o f t ,   k )
          e. Compute KD loss = K L D i v ( s _ l o g i t s / T   | |   t _ s o f t _ k )   ∗   T 2
          f. Compute CE loss = C r o s s E n t r o p y ( s _ l o g i t s ,   y )
          g. Total loss = λ   ∗   K D   l o s s   +   ( 1   −   λ )   ∗   C E   l o s s
          h. Backpropagate and update student parameters
3. Evaluate M_student on D_test to compute ACC_KDk
4. Save trained M_student
5. Return M_student, ACC_KDk

3.3.4. Adversarial Training (KD k + AT)

This was necessary to enhance the robustness of the model against adversarial attacks, as it involves introducing adversarial examples during the model training phase. The magnitude of perturbation is dataset-dependent due to domain differences. The specifics are itemized below:
  • Vision datasets: The adversarial examples were crafted using the fast gradient sign method (FGSM) calculated as ε · s i g n ( ∇ x L ) , where ε represents the magnitude of the perturbation set to 8 / 255 .
  • Textual datasets: A normalized gradient ascent was used for feature vectors, with the perturbation calculated as ϵ · ( ∇ x L / ∇ x L | | 2 ) , with ϵ set to 0.1 and 0.05 for Criteo and Yahoo! Answers, respectively.
The pseudocode for the KDk+AT implementation is shown in Algorithm 2.
Algorithm 2. KD k +AT Implementation
Input: D_train, D_test, Teacher model M_teacher, Parameters ( T ,   k ,   λ ,   ε )
Output: Trained adversarial student model M_AT, Accuracy ACC_KDk_AT
1. Initialize model  M _ A T   =   ( B o t t o m ,   T o p )
2. For each epoch in EPOCHS do:
      For each batch (x, y) in D_train do: Generate adversarial examples:
              i. Compute gradient g = ∇ x   C r o s s E n t r o p y ( M _ A T ( x ) ,   y )
              ii. Perturb input:  x _ a d v   =   x   +   ε   ∗   s i g n ( g )
         b. Forward pass with x_adv: s_logits = M_AT(x_adv)
         c. Teacher forward pass:  t _ l o g i t s   =   M _ t e a c h e r ( x )
         d. Apply temperature scaling + k-anonymity: t_soft_k
         e. Compute KD loss = K L D i v ( s _ l o g i t s / T   | |   t _ s o f t _ k )   ∗   T 2
         f. Compute CE loss = CrossEntropy(s_logits, y)
         g. Total loss = λ   ∗   K D   l o s s   +   ( 1   −   λ )   ∗   C E   l o s s
         h. Backpropagate and update student parameters
3. Evaluate M_AT on D_test to compute ACC_KDk_AT
4. Save trained M_AT
Return M_AT, ACC_KDk_AT

3.3.5. Differential Privacy

Although AT strengthens robustness, it does so by training on perturbed inputs that sharpen the model’s decision boundaries, thereby making its gradients and embeddings more discriminative of true labels. Since label-inference attacks exploit precisely the mutual information between these exchanged tensors and the labels, the more label-aligned gradients produced by AT are also more informative to an attacker; hence, the same mechanism that improves robustness increases label leakage. This is visible in our results, where introducing AT raises attack success before the noise stage restores privacy. A dual-layer defense is therefore required: robustness from AT, and restored privacy from calibrated noise. With differential privacy, calibrated embedding and gradient noise are used to obscure information during the forward and backward passes of the VFL system. We provide a threat-model-level analysis for each LIA. A passive attacker observes the backward gradients only; injecting noise into the passive gradient directly degrades the gradient–label correlation it exploits. Direct and active attackers probe or query the forward embedding; embedding noise perturbs exactly these forward representations, reducing the information available for reconstruction and query-based inference. A perturbed attacker relies on small input variations reflected in the exchanged tensors; the added stochasticity dominates these variations, making label recovery statistically unreliable. The corresponding ASR reductions in Section 4 are consistent with this analysis.
The noise mechanism used here injects calibrated Gaussian noise into the exchanged embeddings and gradients. It does not perform per-example gradient clipping or privacy accounting, and therefore does not produce a formal (ε, δ)-differential-privacy budget; we consequently report privacy empirically, through attack success rate across four LIA types and the Privacy Leakage Index, rather than through a stated ε. A formal (ε, δ) treatment with gradient clipping and a privacy accountant (e.g., an RDP/PRV accountant) is a direct and well-defined extension left for future work.
The pseudocode for the KDk+AT+DP implementation is shown in Algorithm 3.
Algorithm 3. KD k +AT with Differential Privacy
Input: D_train, D_test, Teacher model M_teacher, Parameters (T, k, λ, ε, noise_multiplier)
Output: Trained private student model M_DP, Accuracy ACC_KDk_AT_DP
  • Initialize model  M _ D P   =   ( B o t t o m ,   T o p )
  • For each epoch in EPOCHS do:
  • For each batch  ( x ,   y ) in D_train do:
  • Generate adversarial examples:  x _ a d v   =   x   +   ε   ∗   s i g n ( ∇ x   C r o s s E n t r o p y ( M _ D P ( x ) ,   y ) )
  • Forward pass student:  h _ a   =   B o t t o m _ a ( x _ a d v _ a ) , h _ p   =   B o t t o m _ p ( x _ a d v _ p )
  • Register noise hook on h_p:  h _ p . r e g i s t e r _ h o o k ( l a m b d a   g r a d :   g r a d   +   n o i s e )
  • s _ l o g i t s   =   T o p ( h _ a ,   h _ p )
  • Compute Total loss (including KD and CE loss) as in Algorithm 2
  • Backpropagate and update student parameters (hook will add noise automatically)
  • Evaluate, Save, and Return  M _ D P

3.3.6. Training Objectives for the Pipeline

Each of the models in this study serves a distinguishable purpose. Table 5 below provides an overview of the training objectives for each model.

3.4. Evaluation Metrics

Adversarial robustness in this study is evaluated using the FGSM attack; evaluation under stronger multi-step attacks (PGD, CW, and AutoAttack) is identified as future work. The performance of the three models was evaluated under clean and adversarial conditions using the metrics below:
  • Top-1 Accuracy: A metric is used to determine the frequency with which a model predicts the correct label. It is the ratio of correct sample predictions to the total number of samples in a dataset. An interpretation summary of the metrics is presented later in Table 5 for clarity.
  • Top-5 Accuracy: This is used to indicate the number of times a correct label shows up in the top 5 predicted classes of a model. It was included particularly for the CIFAR-100 dataset due to its large number of classes.
  • F1-Score and AUC: These provide a balanced model performance evaluation for binary tabular and multi-class text datasets.
  • Robust Accuracy: This is measured as the ratio of resilience to perturbation. It is the accuracy of the model in the face of adversarial attack.
  • Attack Success Rate (Top-1 ASR, Top-5 ASR): This is used to determine the level of privacy leakage by measuring the accuracy of a label inference attempt on a model.
  • Privacy Leakage Index: This is a composite metric used for observing privacy–utility balance during model performance evaluation.
Where ASR is the passive-attack success rate; PLI is thus a composite of privacy and the robustness-to-utility ratio, not a pure privacy measure.
L K D k = 1 − λ L C E + λ T 2 L K L
A C C T o p − 1 = 1 N = ∑ i = 1 N ( a r g m a x c p ^ i , c = y i )
A C C T o p − 5 = 1 N = ∑ i = 1 N ( y i ∈ T o p 5 ( p ^ i ) )
F 1   S c o r e = 2   ×   P r e c i s i o n   ×   R e c a l l P r e c i s i o n + R e c a l l
A S R T o p − 1 = 1 N a t t ∑ i = 1 N 1 ( arg m a x c p ^ i , c = y i )
P L I = 1 − A S R 100   ×   R o b u s t   A c c u r a c y C l e a n   A c c u r a c y   ×   100
PLI is a composite index that rewards a model for simultaneously achieving low attack success and preserving accuracy under adversarial conditions; it therefore assumes a functioning model, i.e., clean accuracy > 0 and robust accuracy > 0. The degenerate case robust accuracy = 0 drives PLI to zero regardless of attack success: this is intentional, as a model with no adversarial robustness is not considered a valid privacy-preserving defense under our joint criterion. This condition does not arise in our experiments, where robust accuracy remains high across all datasets.

4. Experiments

This subsection highlights the performance of the models across the five datasets used for performance evaluation, maintaining the grouping order such that the vision datasets are first evaluated, followed by the text/tabular datasets.

4.1. Analysis of the Vision Datasets

CIFAR-10
The result for CIFAR-10 (Table 6) shows the strength of the defense mechanism without trading off utility. It sustained relatively stable accuracy across all model variants, moving from KD k ’s 87.59% to 86.37% (Figure 2a) after the introduction of both adversarial training and differential privacy. The drop in accuracy (1.22%) supports the notion that neither measure had a significant impact on the model’s utility. Furthermore, the random perturbations introduced by the DP mechanism acted below the gradient-signal threshold that would have distorted convergence. This implies that the model training proceeded as if it was not regularized, yet it still embedded privacy at each communication stage.
Adversarial examples drove structural improvement, achieving a significant 23.92% robustness gain over KD k (60.72%) as seen in Figure 2b. The accuracy after an adversarial attack is almost identical to the clean accuracy. The robustness gap decreased to just 0.56% after integrating the DP mechanism, which moves the robustness accuracy to 85.81%. This suggests that adversarial exposure pushes the model towards features that generalize across perturbations, while the injected noise further flattens sharp curvature in the loss surface. The near-overlap of clean and robust curves in Figure 2a implies that the dual-defense mechanism’s local predictions are smoother and less sensitive to pixel-level variations. It can be concluded that the model achieved a robust operating point without compromising the high-fidelity decision boundaries, which historically make maintaining CIFAR-10 accuracy relatively difficult.
Conclusively, the privacy trajectory in Figure 3a highlights the impressive improvement in the Privacy Leakage Index across the models, moving from an average standpoint (51.6%) to 81.3%. This demonstrates the combined effect of stronger robustness and reduced gradient-label correlation, as this approach ensures the model only transmits information required for collaborative optimization. On CIFAR-10’s low-resolution images, the stochasticity behaves like an auxiliary augmentation to marginally improve the model’s generalization. Furthermore, AT and DP helped the model to preserve the utility of the KDk baseline while markedly reducing the FGSM robustness gap. In the context of Vertical Federated Learning, this balance marks a significant step forward, proving that privacy and adversarial robustness can be achieved together without the usual loss in model utility.
CIFAR-100
The CIFAR-100 dataset represents a harder generalization problem than CIFAR-10 due to its 100 fine-grained categories, therefore serving as a good test of the improved defense mechanism’s performance in a label-dense environment. As depicted in Figure 4a, clean accuracy was within the ~60–65% range, indicating that both adversarial training and differential privacy affect model utility. The KD k (64.64%) slightly improved clean accuracy over OA, while the improved KD k +AT+DP settled at a competitive 60.09%, an acceptable trade-off for the security gains, especially in a VFL environment. From Table 7 below, the stability of both F1-score and AUC values across all models proves that fundamental class discrimination and ranking integrity are preserved.
The security robustness gains are quite significant, with a robust accuracy of only 39.71%, the OA model is highly fragile and susceptible to attacks. KD k +AT (57.20%) closed the robustness gap to 6.32% and the addition of differential-privacy noise further boosts robustness to 59.41%, bringing the margin between the model’s performance pre and post adversarial attack down to just 0.68%. This is an important finding of this study, as the injected gradient noise, originally intended for privacy, also served as a powerful regularizer, making the model’s predictions almost invariant to small-norm perturbations.
Privacy gains stem entirely from the altered statistics of the backward-propagated tensors, as only gradient-level noise is applied to the vision task. The complexity of the CIFAR-100 already made it a difficult task for attackers, hence the passive and direct ASR are quite low, further dropping to the 2–8% range for KD k . From Figure 4b, the trend is noticeable as the ASRs dropped further to 2.52% and 6.77%, respectively. These extremely small success rates indicate that gradients transmitted from the top model to the passive party are now largely label independent. The same trend is noticeable with the active and perturbed LIAs, which dropped to 12.38% and 9.84% from 20.57% and 14.73% recorded for the baseline OA model. In practical terms, gradient noise injects just enough variance to make label inference statistically infeasible without corrupting the semantic direction of the update. CIFAR-100 contains many visually similar classes; hence, the small overlaps in feature space magnify any residual correlation between gradients and labels; randomizing those gradients thus yields a substantial privacy payoff. As seen in Figure 3a, the steady compression of the threat model curves towards the lower bound confirms that the model leaks minimal information through its shared updates.
From Figure 3b, the Privacy Leakage Index moved from the near average 59.34% (OA) to a remarkable 96.08% with the improved KD k +AT+DP defense mechanism. This near-saturation of the privacy metric perfectly demonstrates the ability to balance utility, robustness, and privacy. The model’s robustness against evasion attacks and its highly competitive accuracy post-attack are essential properties for achieving a robust, privacy-preserving VFL, especially in privacy-critical real-world scenarios.
CINIC-10
CINIC-10 combines multiple data sources into a heterogeneous domain, combining CIFAR-10 and ImageNet for visual diversity and scale. For optimal performance, CINIC-10 requires a model that is adaptable to distribution shifts while maintaining stable feature geometry. The similarity of Clean Accuracy across the four models demonstrates the strong balance of the dual-defense mechanism between accuracy and generalization. With 76.6% compared to KD k ’s 79.5% clean accuracy, it can be argued that the model’s feature extractor preserves semantic integrity in the face of gradient perturbation. The minor decline reflects normal stochastic regularization, further proving that the mechanism operates in a region of controlled noise that is small enough to maintain the fidelity of the learned visual representations but sufficient to conceal label information. This confirms that adversarial robustness can be achieved without eroding a model’s ability to classify clean, unseen images correctly.
As presented in Table 8 below, adversarial training plays a decisive role in improving robustness. In Figure 5a, the KD k model achieved a 43.09% robust accuracy, slightly higher than the OA architecture (33.80%). This is way below the average mark, further proving that k -anonymity smoothing is not sufficient for feature boundaries in a VFL environment. With the KD k +AT+DP model, robustness improved to an impressive 74.36%, aligning both clean and perturbed accuracies. The overlap of the curves in Figure 4b indicates the importance of adversarial training in reshaping the model’s gradients as DP regularization dampens minor oscillations that cause sensitivity to inter-domain features. The improved defense mechanism, therefore, attains a stable operating point in which robustness and accuracy coexist, even under the CINIC-10’s complex domain blend.
It was evident that gradients no longer carry exploitable signals, with passive ASR falling from 29.86% to 16.18%, and direct ASR from 26.39% to 19.93%. This proves that gradient noise translates directly into measurable confidentiality gains. Active ASR decreased to 33.51%, reflecting the fact that crafted-input probing elicits more homogeneous responses after adversarial training. The perturbed ASR (12.31) is the most stable, confirming that small random variations cannot recover hidden labels once a model is noise-regularized. Figure 5b shows the four curves converging downward, indicating how the models uniformly mitigate both passive observation and active querying. The privacy gains can be attributed to the reduction of label correlation in the transmitted gradients since noise was added only in the backward pass.
The Privacy Leakage Index curve presented in Figure 3c consistently rose from the OA model (23.37%) to 82.41%. This privacy–utility gain perfectly demonstrates how the defense structure evolved from a model that exposes half of its labels through gradients and embeddings to one that no longer exposes them. The massive defensive gain is partly driven by the mixed-domain nature of the CINIC-10 dataset, where gradient noise acts as an implicit alignment mechanism, suppressing domain-specific biases that would otherwise appear as leakage channels. With this outcome, the KD k +AT+DP model achieved three important goals simultaneously by competitively maintaining KD k ’s high accuracy in addition to its adversarial-level robustness, and the best privacy index among all tested vision datasets. More succinctly, these outcomes demonstrate that the improved defense mechanism is not only effective under uniform data distributions but also resilient to real-world dataset heterogeneity, an important achievement towards deployable privacy-preserving federated learning in the visual domains.

4.2. Analysis of the Textual Dataset

The Yahoo! Answers dataset represents a different modality, as it is based entirely on linguistic content. Therefore, the models must infer class semantics from contextual embeddings, making linguistic generalization and label privacy both dependent on how the frozen all-mpnet-base-v2 sentence embeddings are processed by the bottom and federated top models. The clean-accuracy pattern (Figure 6a) is noticeably stable, with 73.1% and 72.8% for KD k and KD k +AT+DP, respectively. The slight reduction in accuracy (0.3%) indicates that the injection of gradient noise into the backward pass does not degrade semantic alignment between label representations and embeddings. The F1-score (72%), which is a better metric for measuring balanced precision–recall in textual data, validates the model’s ability to preserve linguistic structures distilled from OA. DP noise acted like a gradient stabilizer, helping to prevent overfitting to frequent lexical patterns while maintaining the model’s linguistic discriminability.
As predicted, adversarial training improved robustness against input data perturbations, with robust ACC rising from KD k (64.46%) to 70.12%. This gain shrank the vulnerability window that adversarial token substitutions typically exploit for malicious intents. Furthermore, the improved defense mechanism sustained 69.95% robustness, indicating that gradient noise does not interfere with utility (Table 9).
To put it more succinctly, this highlights the alignment of the model’s decision boundaries with topic-level meaning rather than relying on lexical cues, which is important for real-world text classification under noisy or adversarial conditions. The convergence of clean and robust curves in Figure 6a indicates that the model learned perturbation-invariant linguistic features, likely relying on contextual semantics rather than individual word embeddings.
The backward pass perturbation obscured gradient–label correlations without disrupting the model’s forward semantic flow. This is supported by the privacy gains across the text pipeline, with passive ASR dropping to 9.75% after the integration of AT nearly doubled it, while direct ASR stabilized at 18% (Figure 6b). However, this led to more predictable behavior under query-based probing as the active ASR rose from 18.27% to 26.73%. The absolute success rates remained low, and perturbed ASR stabilized near 10%, indicating that the model’s gradients no longer encode deterministic label traces. The Privacy Leakage Index (Figure 3d), increasing from 79.33% in KDk to 86.68% under KD k +AT+DP, represents an important privacy gain.
Overall, the Yahoo! Answers results demonstrate that the defense mechanism can generalize well beyond the vision domain by preserving both predictive performance and privacy integrity. This adaptability highlights the robustness of the approach for real-world text-classification deployments in VFL environments, where the coexistence of label privacy and communication efficiency is necessary.

4.3. Analysis of the Tabular Dataset

The Criteo CTR is a large-scale binary classification dataset, presenting unique challenges for privacy-preserving learning. Its dense numerical features and class imbalance make AUC (within the narrow 69–71% range) a more reliable indicator of predictive utility than raw accuracy as presented in Table 10 below. This implies that neither adversarial exposure nor DP noise undermines the model’s ability to discriminate between positive and negative samples. Furthermore, clean accuracy also remained steady, hovering around 75% across variants, with the improved defense model slightly exceeding KDk in clean performance (75.10% vs. 75.67%) as seen in Figure 7a. This consistency indicates that gradient-level noise is well-calibrated, effectively regularizing weight updates without distorting the signal-to-noise ratio in the dense feature space.
Adversarial training enhanced the model’s resilience to distributional noise that is consistent with feature perturbations in tabular data. The robust ACC improved to nearly match the clean accuracy, moving from 72.82% (KD k ) to 74.66% (KD k +AT), thereby shrinking the robustness gap to 0.06%. This convergence implies that adversarial examples did not distort the learned decision surface, with confidence remaining stable even when input distributions shifted.
Furthermore, DP noise stabilized robust accuracy at 74.52%, suggesting that the gradient noise introduced minimal degradation after AT. The convergence of both curves (Figure 7a) demonstrates that robustness was achieved without sacrificing accuracy, which is a rare outcome in noisy, privacy-sensitive tabular settings. Lastly, since features are continuous and correlated, Gaussian perturbations at the gradient level primarily act as a smoothness constraint, preventing sharp curvature in the loss landscape.
The Criteo CTR results demonstrate a clear downward trajectory in all attack success rates once the defense is applied. AT reshapes gradient magnitudes, while DP noise masks the residual label correlation within them. As a result, passive ASR drops from 73.08% to 25.54%, and active ASR drops from 74.47% to 63.21%. Direct ASR also decreased from 74.10% to 67.30%, while perturbed ASR dropped from 74.47% to 47.83%, confirming that adversarially stable gradients are less susceptible to reconstruction-based inference. This consistent downward trend of attack success across the four privacy threat models is captured in Figure 7b.
The improvement in the Privacy Leakage Index (Figure 3e) from KD k ’s 24.28% to 73.88%, depicts a three-fold improvement in overall label protection. It can therefore be concluded that the gradient-level differential-privacy noise mitigated attack types, achieving strong privacy reinforcement without impairing task performance.

4.4. Summary of Results

It is evident that the improved defense mechanism, KD k +AT+DP, attained privacy and adversarial robustness while preserving model utility. In so doing, it has successfully addressed this study’s research problem, unlike existing VFL defense mechanisms, which often trade privacy for robustness.

4.4.1. Model Utility

The improved defense mechanism maintained competitive clean accuracy, close to that of the OA baseline and KDk, across the benchmark dataset suite. For the image classification tasks, accuracy differences stayed within ~2–3%, even after DP noise was injected, while the AUC and F1-Score remained relatively unchanged for the text and tabular tasks. By obtaining this level of stability, the mechanism created a training approach that is resilient to both perturbation and over-fitting, explaining the uniform preservation of predictive performance across modalities.

4.4.2. Adversarial Robustness

The inclusion of adversarial examples strengthens the model’s resilience to small-norm or semantic perturbations. The robustness gap, which was often above 30% in the traditional VFL (OA), collapsed to less than 1% across some of the datasets due to the improved defense mechanism. For the visual tasks, this translates to stable predictions under pixel-level noise, consistent topic recognition under token substitutions in textual data, and unchanged probability rankings when feature noise is introduced in tabular data. More importantly, the addition of differential privacy did not undo these gains, with the injected gradient noise aligning with adversarial regularization.
Communication per round is identical across all four stages, confirming that neither adversarial training nor noise injection introduces communication overhead. The only measurable computational increase arises at the adversarial-training stage, from adversarial-example generation; the noise stage adds a negligible per-epoch increment (≈0.2–0.3 s) and no communication cost. These overheads are confined to training. Both DP-style noise hooks and the adversarial-example generation operate only during the training loop; at inference, the model reduces to a standard forward pass that is identical across all four configurations, so no additional inference-time or communication overhead is introduced at deployment.
Table 11. Per-step communication and computational cost.
Table 11. Per-step communication and computational cost.
DatasetStageCommunication/EpochTime/Epoch (Measured)
Yahoo! AnswersOA204.7 MB6.3 s
KD k 204.7 MB5.7 s
KD k +AT204.7 MB8.6 s
KD k +AT+DP204.7 MB8.8 s
Criteo CTROA41.0 MB7.6 s
KD k 41.0 MB8.2 s
KD k +AT41.0 MB11.0 s
KD k +AT+DP41.0 MB11.3 s
CIFAR-10All stages51.2 MBNot retained
CIFAR-100All stages51.2 MBNot retained
CINIC-10All stages184.3 MBNot retained
The stages differ only in local computation, not communication. KDk introduces one frozen-teacher forward pass per batch; adversarial training adds one forward–backward pass per batch to generate the FGSM example, approximately doubling the base per-batch cost; and noise injection adds a single Gaussian draw of the embedding’s dimension, which is negligible. These expectations are consistent with the measured per-epoch times, where the transition to adversarial training produces the largest increase (e.g., Yahoo 5.7 → 8.6 s; Criteo 8.2 → 11.0 s), while the noise stage adds only ~2–3%.

4.4.3. Label Privacy

The analysis of all attacks under both adversarial and privacy threat models confirmed that the improved defense mechanism lowered LIA success rates. Gradient noise proves most effective against passive and direct LAs, reducing their success to single-digit percentages in most of the datasets. Adversarial training played an important role in flattening gradient distributions, indirectly reducing leakage channels even before noise injection. The marginally higher active LIA success rates, notably in the Yahoo! Answers dataset, were due to the smoother response surfaces that made model outputs more predictable under crafted queries, yet those surfaces carry little usable label information. The upward trajectory of the Privacy Leakage Index across all datasets indicates that privacy protection improves steadily without cost to task accuracy, confirming that the system resists both observation and query-based LIAs.

4.4.4. Training Stability

Across all stages, the training-loss and accuracy trajectories converge smoothly without oscillation (see logs). The noise stage was set to a small learning rate deliberately to inject privacy noise without displacing the converged robust operating point; the resulting flat trajectories therefore reflect both the well-behaved objective and this conservative step size. A formal convergence proof for the joint distillation–adversarial–noise objective is beyond the present scope and is left for future work.

4.4.5. Broader Implications

The results of this study have demonstrated that privacy-preserving VFL environments do not need to depend on disruptive model redesign and/or heavy cryptographic protocols for effectiveness. By targeting the learning dynamics, such as statistical signatures that carry label information, this defense mechanism achieved privacy through controlled regularization rather than external encryption. Additionally, the multi-modal relevance of this mechanism emphasizes its generalizability, as the entire experiment was conducted under the same configuration.
Lastly, by addressing the research problem, this study showed that privacy and robustness are not opposing objectives, but they can reinforce each other when jointly calibrated. This insight contributes to the broader discourse on privacy-preserving FL with the empirically validated approach for maintaining robustness and utility under the most stringent privacy constraints.

5. Conclusions

The continued interconnection of devices globally has projected adversarial and privacy attacks as one of the most persistent threats in VFL environments; hence, addressing them has typically forced an unwanted privacy–utility or privacy–robustness trade-off. Since the embeddings and gradients exchanged between parties can be exploited to recover hidden label information, defenses that strengthen robustness have often done so at the expense of privacy. This study set out to enhance robustness without weakening privacy, and presented a defense mechanism that treats the two as complementary rather than competing security objectives in VFL environments.
The knowledge-distillation-based KDk mechanism achieved substantial privacy gains but did not improve robustness. To secure both, this study integrated adversarial training (AT) with a differential-privacy-style stage based on calibrated noise injection. This helps to restore the label privacy weakened by AT while preserving the robustness gains, with minimal impact on model utility. The mechanism was evaluated across five datasets spanning vision, textual, and tabular classification, thus providing cross-domain coverage and evidence of generalizability. Privacy was assessed empirically through the success rate of four label-inference attacks and the Privacy Leakage Index, which improved consistently across datasets, indicating that the models retained useful predictive patterns while exposing little recognizable label information during embedding and gradient exchange. Under FGSM-based evaluation, robust accuracy was maintained close to clean accuracy, and—as the cost analysis showed—these gains were obtained without additional communication or cryptographic overhead. Taken together, the results indicate that privacy-preserving and robust VFL is achievable through targeted training design, with privacy and robustness acting as complementary objectives.
This study has three limitations that define a path for future work. Adversarial robustness is evaluated only under FGSM, and single-step evaluation can overstate robustness, so assessment against stronger attacks such as PGD, CW, and AutoAttack will give clearer indications of the robustness gains. The differential-privacy component is applied as calibrated noise injection with privacy shown empirically, rather than through a formal (ε, δ) budget. An improvement on this study will involve the use of clipping and privacy accounting. Lastly, a broader benchmarking against recent VFL defenses and a formal convergence analysis of the joint objective would further strengthen the work.

5.1. Recommendations

Based on the outcomes of this study, the following are recommended:
  • Privacy-preserving VFL implementations should embrace adaptive noise control mechanisms to ensure noise variance scales as gradient sensitivity increases.
  • The availability of benchmark datasets and standardized scripts for threat model evaluations would aid reproducibility across the research community. This will be helpful in accelerating the progress made towards achieving privacy-preserving FL at scale.
  • Finance, healthcare, marketing and other domains holding sensitive user information should embrace multi-stage defense mechanisms for an added layer of security. The organizations should also embrace integrating training-based defense mechanisms with homomorphic masking techniques.

5.2. Future Work

The improved defense mechanism has brought about important robustness and privacy gains but also pointed out research avenues that should be explored in the future. The use of adaptive and context-aware noise budgets will make privacy protection more intelligent and data-driven. The defense mechanism can also be extended to multi-party and hierarchical federated structures to explore the dynamics of gradient perturbation in more complex network structures.

Author Contributions

Conceptualization, N.A.A., O.S.M. and O.M.O.; methodology, N.A.A., O.S.M. and O.M.O.; software, O.S.M., O.M.O. and A.A.A.; validation, N.A.A., O.S.M., O.M.O., A.A.A. and D.S.A.; formal analysis, D.S.A., C.V.D.V. and C.E.O.; investigation, N.A.A., O.S.M., O.M.O. and A.A.A.; resources, N.A.A., O.S.M., O.M.O., A.A.A., D.S.A., C.V.D.V. and C.E.O.; data curation, N.A.A., O.S.M. and O.M.O.; writing—original draft preparation, N.A.A., O.S.M., O.M.O., and A.A.A.; writing—review and editing, A.A.A., D.S.A., C.V.D.V., and C.E.O.; visualization, N.A.A., O.S.M., and O.M.O.; supervision, N.A.A.; project administration, N.A.A., O.S.M., O.M.O., A.A.A., D.S.A., C.V.D.V., and C.E.O.; funding acquisition, C.V.D.V. All authors have read and agreed to the published version of the manuscript.

Funding

The authors wish to acknowledge the financial support received from the School of Computer Science and Information Systems, Faculty of Natural and Agricultural Sciences, Vaal Triangle Campus, North-West University, Vanderbijlpark 1900, South Africa.

Data Availability Statement

The data presented in this study are openly available in https://github.com/malomodaniels/KDk-AT-DP-Implementation.

Conflicts of Interest

There is no conflict of interest in this research work.

References

  1. Menard, P.; Bott, G.J. Artificial Intelligence Misuse and Concern for Information Privacy: New Construct Validation and Future Directions. Inf. Syst. J. 2025, 35, 322–367. [Google Scholar]
  2. Ajagbe, S.A.; Awotunde, J.B.; Florez, H. Ensuring Intrusion Detection for IoT Services Through an Improved CNN. SN Comput. Sci. 2024, 5, 49. [Google Scholar] [CrossRef] [Scilit]
  3. Taiwo, G.; Vadera, S.; Alameer, A. Vision Transformers for Automated Detection of Pig Interactions in Groups. Smart Agric. Technol. 2025, 10, 100774. [Google Scholar] [CrossRef] [Scilit]
  4. Wieringa, J.; Kannan, P.K.; Ma, X.; Reutterer, T.; Risselada, H.; Skiera, B. Data Analytics in a Privacy-Concerned World. J. Bus. Res. 2021, 122, 915–925. [Google Scholar] [CrossRef] [Scilit]
  5. Fu, A.; Zhang, J.; Yang, Q. Label Inference Attacks Against Vertical Federated Learning. IEEE Trans. Inf. Forensics Secur. 2022, 17, 1162–1174. [Google Scholar] [CrossRef] [Scilit]
  6. Kairouz, P.; McMahan, H.B.; Avent, B.; Bellet, A.; Bennis, M.; Bhagoji, A.N.; Bonawitz, K.; Charles, Z.; Cormode, G.; Cummings, R.; et al. Advances and Open Problems in Federated Learning. Found. Trends Mach. Learn. 2021, 14, 1–210. [Google Scholar] [CrossRef] [Scilit]
  7. Niknam, S.; Dhillon, H.S.; Reed, J.H. Federated Learning for Wireless Communications: Motivation, Opportunities, and Challenges. IEEE Commun. Mag. 2020, 58, 46–51. [Google Scholar] [CrossRef] [Scilit]
  8. McMahan, H.B.; Moore, E.; Ramage, D.; Hampson, S.; Arcas, B.A. Communication-Efficient Learning of Deep Networks from Decentralized Data. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics; PMLR: London, UK, 2017; pp. 1273–1282. Available online: https://proceedings.mlr.press/v54/mcmahan17a.html (accessed on 5 August 2025).
  9. Adeniyi, J.K.; Ajagbe, S.A.; Adeniyi, E.A.; Mudali, P.; Adigun, M.O.; Adeniyi, T.T.; Ajibola, O. A Biometrics-Generated Private/Public Key Cryptography for a Blockchain-Based E-Voting System. Egypt. Inform. J. 2024, 25, 100447. [Google Scholar] [CrossRef] [Scilit]
  10. Chakraborty, A.; Dahal, C.; Gupta, V. Federated Retrieval-Augmented Generation: A Systematic Mapping Study. arXiv 2025, arXiv:2505.18906v1. [Google Scholar]
  11. Khan, A.; Thij, M.; Wilbik, A. Vertical Federated Learning: A Structured Literature Review. Knowl. Inf. Syst. 2025, 67, 3205–3243. [Google Scholar] [CrossRef] [Scilit]
  12. Arazzi, M.; Nicolazzo, S.; Nocera, A. A Defense Mechanism against Label Inference Attacks in Vertical Federated Learning. Neurocomputing 2025, 624, 129476. [Google Scholar] [CrossRef] [Scilit]
  13. Liu, Y.; Zou, T.; Kang, Y.; Liu, W.; He, Y.; Yi, Z.; Yang, Q. Batch Label Inference and Replacement Attacks in Black-Boxed Vertical Federated Learning. arXiv 2022, arXiv:2112.05409. [Google Scholar]
  14. Xu, J.; Zhang, Z.; Hu, R. Achieving Byzantine-Resilient Federated Learning via Layer-Adaptive Sparsified Model Aggregation. In Proceedings of the 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Tucson, AZ, USA, 26 February–6 March 2025; pp. 1508–1517. [Google Scholar]
  15. Aono, Y.; Hayashi, T.; Wang, L.; Moriai, S. Privacy-Preserving Deep Learning via Additively Homomorphic Encryption. IEEE Trans. Inf. Forensics Secur. 2017, 13, 1333–1345. [Google Scholar]
  16. Ye, M.; Shen, W.; Du, B.; Snezhko, E.; Kovalev, V.; Yuen, P.C. Vertical Federated Learning for Effectiveness, Security, Applicability: A Survey. arXiv 2024, arXiv:2405.17495. [Google Scholar]
  17. Luo, X.; Wu, Y.; Xiao, X.; Ooi, B.C. Feature Inference Attack on Model Predictions in Vertical Federated Learning. In Proceedings of the 37th International Conference on Data Engineering (ICDE), Chania, Greece, 19–22 April 2021; pp. 181–192. [Google Scholar]
  18. Wei, K.; Li, J.; Ma, C.; Ding, M.; Wei, S.; Wu, F.; Chen, G.; Ranbaduge, T. Vertical Federated Learning: Challenges, Methodologies and Experiments. arXiv 2022, arXiv:2202.04309. [Google Scholar]
  19. Al Farsi, A.; Khan, A.; Mughal, M.R.; Bait-Suwailam, M.M. Privacy and Security Challenges in Federated Learning for UAV Systems: A Systematic Review. IEEE Access 2025, 13, 86599–86615. [Google Scholar] [CrossRef] [Scilit]
  20. Hu, K.; Gong, S.; Zhang, Q.; Seng, C.; Xia, M.; Jiang, S. An Overview of Implementing Security and Privacy in Federated Learning. Artif. Intell. Rev. 2024, 57, 204. [Google Scholar] [CrossRef] [Scilit]
  21. Azeez, N.A.; Malomo, O.S.; Aaron, D.S.; Ademoye, A.A.; Okerinde, O.M.; Otolehi, U.D.; Lukman, O.O. Artificial Intelligence in Cybersecurity: A Comparative Review of Its Role across the Cyber Kill Chain. Univ. Ib. J. Sci. Log. ICT Res. 2025, 14, 153. [Google Scholar]
  22. Azeez, N.A.; Ademoye, A.A.; Malomo, O.S.; Okerinde OMAaron, D.S.; Vyver, C.V. Investigation of Augmented Datasets for Security in Internet of Medical Things (IoMT) Ecosystems. Computers 2026, 15, 369. [Google Scholar] [CrossRef] [Scilit]
  23. Finlayson, S.G.; Bowers, J.D.; Ito, J.; Zittrain, J.L.; Beam, A.L.; Kohane, I.S. Adversarial Attacks on Medical Machine Learning. Science 2019, 363, 1287–1289. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Azeez, N.A.; Adefemi, F.; Olayinka Fasina, E.P.; Venter, I.M. Evaluation of a Flexible Column-Based Access Control Security Model forMedical-Based Information. J. Comput. Sci. Its Appl. 2015, 22, 24–31. [Google Scholar]
  25. Zhan, P.; Yang, J.; Wang, H.; Zheng, C.; Wang, L. Rethinking Word-level Adversarial Attack: The Trade-off Between Efficiency, Effectiveness, and Imperceptibility. In Proceedings of the Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Torino, Italy, 20–25 May 2024; ELRA Language Resource Association: Paris, France, 2024; pp. 14037–14052. [Google Scholar]
  26. Jedrzejewski, F.V.; Thode, L.; Fischbach, J.; Gorschek, T.; Mendez, D.; Lavesson, N. Adversarial Machine Learning in Industry: A Systematic Literature Review. Comput. Secur. 2024, 145, 103988. [Google Scholar] [CrossRef] [Scilit]
  27. Azeez, N.A.; Venter, I.M. Towards ensuring scalability, interoperability and efficient access control in a multi-domain grid-based environment. Afr. Res. J. 2013, 104. [Google Scholar] [CrossRef] [Scilit]
  28. Yan, Z.; Yao, Y.; Wen, X.; Zhang, J.; Fan, K. LADSG: Label-Anonymized Distillation and Similar Gradient Substitution for Label Privacy in Vertical Federated Learning. arXiv 2025, arXiv:2506.06742. [Google Scholar]
  29. Wang, Y.; Lv, Q.; Zhang, H.; Zhao, M.; Sun, Y.; Ran, L.; Li, T. Beyond Model Splitting: Preventing Label Inference Attacks in Vertical Federated Learning with Dispersed Training. World Wide Web 2023, 26, 2691–2707. [Google Scholar] [CrossRef] [Scilit]
  30. Azeez, N.A.; Iliyas, H.D. Implementation of a 4-tier Cloud-Based Architecture for Collaborative Health Care Delivery. Niger. J. Technol. Dev. 2016, 13, 17–25. [Google Scholar] [CrossRef] [Scilit]
  31. Ding, L.; Bao, H.; Lv, Q.; Zhang, F.; Zhang, Z.; Han, J.; Ding, S. Threshold Filtering for Detecting Label Inference Attacks in Vertical Federated Learning. Electronics 2024, 13, 4376. [Google Scholar] [CrossRef] [Scilit]
  32. Bai, L.; Zhang, X.; Zhang, S.; Ye, Q.; Hu, H. ProVFL: Property Inference Attacks against Vertical Federated Learning. IEEE Trans. Inf. Forensics Secur. 2025, 20, 6529–6543. [Google Scholar] [CrossRef] [Scilit]
  33. Azeez, N.A.; Tajudeen, A. A Survey On Categorization of Threat Intelligence and Trust-Based Sharing Strategies on Cyber Attack. Vokasi Unesa Bull. Eng. Technol. Appl. Sci. 2025, 2, 128–143. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Framework of the KDkATDP defense mechanism.
Figure 1. Framework of the KDkATDP defense mechanism.
Informatics 13 00127 g001
Figure 2. Experimental results for CIFAR-10. (a) Clean accuracy vs. robust accuracy for CIFAR-10; (b) LIA success rates for CIFAR-10.
Figure 2. Experimental results for CIFAR-10. (a) Clean accuracy vs. robust accuracy for CIFAR-10; (b) LIA success rates for CIFAR-10.
Informatics 13 00127 g002
Figure 3. Evaluation of Private Leakage Index (PLI). (a) PLI for CIFAR-10; (b) PLI for CIFAR-100; (c) PLI for CINIC-10; (d) PLI for Yahoo! Answers; and (e) PLI for Criteo CTR.
Figure 3. Evaluation of Private Leakage Index (PLI). (a) PLI for CIFAR-10; (b) PLI for CIFAR-100; (c) PLI for CINIC-10; (d) PLI for Yahoo! Answers; and (e) PLI for Criteo CTR.
Informatics 13 00127 g003
Figure 4. Experimental results for CIFAR-100. (a) Clean accuracy vs. robust accuracy for CIFAR-100; (b) LIA success rates for CIFAR-100.
Figure 4. Experimental results for CIFAR-100. (a) Clean accuracy vs. robust accuracy for CIFAR-100; (b) LIA success rates for CIFAR-100.
Informatics 13 00127 g004
Figure 5. Experimental results for CINIC-10. (a) Clean accuracy vs. robust accuracy for CINIC-10; (b) LIA success rates for CINIC-10.
Figure 5. Experimental results for CINIC-10. (a) Clean accuracy vs. robust accuracy for CINIC-10; (b) LIA success rates for CINIC-10.
Informatics 13 00127 g005
Figure 6. Experimental results for Yahoo! Answers. (a) Clean accuracy vs. robust accuracy for Yahoo! Answers; (b) LIA success rates for Yahoo! Answers.
Figure 6. Experimental results for Yahoo! Answers. (a) Clean accuracy vs. robust accuracy for Yahoo! Answers; (b) LIA success rates for Yahoo! Answers.
Informatics 13 00127 g006
Figure 7. Experimental results for Criteo. (a) Clean accuracy vs. robust accuracy for Criteo; (b) LIA success rates for Criteo.
Figure 7. Experimental results for Criteo. (a) Clean accuracy vs. robust accuracy for Criteo; (b) LIA success rates for Criteo.
Informatics 13 00127 g007
Table 1. Positioning against representative VFL defenses.
Table 1. Positioning against representative VFL defenses.
MethodPrivacy DefenseRobustness DefenseGuaranteeModalities EvaluatedLIA Types
DP baselineYesNoFormal (ε)SinglePassive
KDk [12]YesNoEmpiricalVisionPassive/Direct
LADSG [28]YesNoEmpiricalTabular/TextPassive/Active/Direct
Dispersed training [29]YesNoEmpiricalVisionPassive
ProVFLNoNoNoVisionProperty inference
This workYesYesEmpiricalTabular/Text/VisionPassive/Active/Direct/Perturbed
Table 2. Overview of the datasets used.
Table 2. Overview of the datasets used.
DatasetDomainSize (Samples × Features)Task TypeEvaluation Focus
CIFAR-10Image50,000 images
32 × 32 RGB) in 10 classes
Multiclass ClassificationBaseline benchmark for adversarial robustness on simple vision tasks
CIFAR-100Image50,000 images
32 × 32 RGB) in 100 classes
Multiclass ClassificationStress-test robustness and distillation under high-granularity class
CINIC-10Image180,000 images
32 × 32 RGB in 10 classes
Multiclass ClassificationLarger-scale benchmark bridging to validate generalization
Yahoo! AnswersText50,000
5000 features across 10 classes
Textual/Natural Language ProcessingEvaluation of the model on high-dimensional sparse text representations with semantic variability.
Criteo CTRTabular80,000
13 continuous
26 categorical features
Binary ClassificationAssesses the generalization of the model on multi-modal data
Tabular Dataset: Since this is a low-dimensional dataset, there is a need to be careful with respect to the choice of models to avoid overfitting. Hence, a relatively simple single-layer feedforward network was adopted. It was implemented as a single linear layer with a ReLU activation. For structured data with very few features (6 for the active party, 7 for the passive), a deep network is unnecessary and prone to memorizing noise. This simple linear transformation provides sufficient expressive power without being over-parameterized. The top model is a shallow 3-layer MLP that receives the concatenated outputs from the bottom models. This architecture is efficient and appropriately sized for the simple binary classification task.
Table 3. Model architecture per dataset.
Table 3. Model architecture per dataset.
DatasetBottom-Model ArchitectureTop-Model Architecture
CIFAR-10Shallow VGG CNNMLP (3-Layer)
CIFAR-100Shallow VGG CNNMLP (3-Layer)
CINIC-10Shallow VGG CNNMLP (3-Layer)
Yahoo! AnswersMLP (2-Layer)MLP (2-Layer)
CriteoLinear+ReLUMLP (3-Layer)
Textual Dataset: The design leverages powerful pre-trained embeddings with a 2-layer MLP bottom model that processes pre-computed sentence embeddings generated by the a l l − m p n e t − b a s e − v 2 SentenceTransformer model. This helps to obtain a deep linguistic understanding of large language models without incurring the massive computational cost of fine-tuning the entire transformer, especially within a VFL setting. It serves as a refiner for the features by adapting the general-purpose embeddings to the specific classification task. It effectively integrates semantic information from the separate title and content embeddings before making the final topic prediction.
Table 4. Data partitioning structure across the datasets.
Table 4. Data partitioning structure across the datasets.
DatasetModeParty A
CIFAR-10Vision1 × 32 × 32
CIFAR-100Vision1 × 32 × 32
CINIC-10Vision1 × 32 × 32
Yahoo! AnswersText768
Criteo CTRTabular6
Table 5. Training objectives per model.
Table 5. Training objectives per model.
ModelPrimary GoalDefense Mechanism
OAImplement the baseline performance and subsequent vulnerability referenceNone—trained only on clean data
KD k Privacy-preservation mechanism against gradient-based label inference attacksKnowledge distillation + data anonymization
KD k +ATGoes a step further by combining privacy-preserving measures and robustness against both gradient and perturbed LIAsKD k framework + FGSM-based adversarial training
KD k +AT+DPRecovers the privacy leakage caused by the introduction of ATDifferential-privacy-style noise injection
Table 6. Implementation result for CIFAR-10.
Table 6. Implementation result for CIFAR-10.
ModelClean ACCRobust ACCRobustness GapF1-ScoreAUCPassive ASRDirect ASRActive ASRPertASRPLI
OA87.5550.4637.0987.5099.0723.0665.4035.6326.9044.34
KD k 87.5960.7226.8787.5199.0525.5856.4139.9119.0351.59
KD k +AT86.3984.641.7586.3498.9340.3834.0443.3910.0158.41
KD k +AT+DP86.3785.810.5686.3498.8618.2514.1143.3910.0181.32
Table 7. Implementation result for CIFAR-100.
Table 7. Implementation result for CIFAR-100.
ModelClean ACCRobust ACCRobustness GapF1-ScoreAUCPassive ASRDirect ASRActive ASRPertASRPLI
OA63.7439.7124.0363.5898.694.759.0720.5714.7359.34
KD k 64.6442.0322.6164.3698.692.528.1114.7011.0763.38
KD k +AT63.5257.206.3263.2798.642.557.5317.949.8487.75
KD k +AT+DP60.0959.410.6859.8298.342.526.7712.389.8496.08
Table 8. Implementation result for CINIC-10.
Table 8. Implementation result for CINIC-10.
ModelClean ACCRobust ACCRobustness GapF1-ScoreAUCPassive ASRDirect ASRActive ASRPertASRPLI
OA79.1733.8045.3779.1397.8445.2550.8268.5540.7423.37
KD k 79.5143.0936.4279.4697.9426.7240.1135.7221.0639.71
KD k +AT77.8274.363.4677.7797.6129.8626.3943.2913.3667.02
KD k +AT+DP76.6075.301.3076.5197.3216.1819.9333.5112.3182.41
Table 9. Implementation result for Yahoo! Answers.
Table 9. Implementation result for Yahoo! Answers.
ModelClean ACCRobust ACCRobustness GapF1-ScoreAUCPassive ASRDirect ASRActive ASRPerturbed ASRPLI
OA72.7665.067.7072.2495.1612.8320.8638.5211.5877.94
KD k 73.1364.468.6772.6195.229.9919.6717.8010.2879.33
KD k +AT72.9670.122.8472.3795.2318.1117.4018.2713.3278.71
KD k +AT+DP72.8469.952.8972.2695.249.7518.3426.7310.7186.68
Table 10. Implementation result for Criteo CTR.
Table 10. Implementation result for Criteo CTR.
ModelClean ACCRobust ACCRobustness GapF1-ScoreAUCPassive ASRDirect ASRActive ASRPertASRPLI
OA75.8072.263.5471.0170.8873.0847.5274.4774.4725.67
KD k 75.6772.822.8571.5270.7974.7874.1069.9136.6624.28
KD k +AT74.7274.660.0673.3169.2555.2270.1074.4773.9044.74
KD k +AT+DP75.1074.520.5873.1769.5925.5467.3063.2147.8373.88
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Azeez, N.A.; Malomo, O.S.; Okerinde, O.M.; Ademoye, A.A.; Aaron, D.S.; Vyver, C.V.D.; Ogbonna, C.E. Adversarial Training and Differential Privacy-Style Noise Injection for Privacy-Preserving Vertical Federated Learning. Informatics 2026, 13, 127. https://doi.org/10.3390/informatics13080127

AMA Style

Azeez NA, Malomo OS, Okerinde OM, Ademoye AA, Aaron DS, Vyver CVD, Ogbonna CE. Adversarial Training and Differential Privacy-Style Noise Injection for Privacy-Preserving Vertical Federated Learning. Informatics. 2026; 13(8):127. https://doi.org/10.3390/informatics13080127

Chicago/Turabian Style

Azeez, Nureni Ayofe, Oluwatobi Sunday Malomo, Omotolani Mary Okerinde, Abdullateef Akorede Ademoye, Damilola Seun Aaron, Charles Van Der Vyver, and Chijioke Erasmus Ogbonna. 2026. "Adversarial Training and Differential Privacy-Style Noise Injection for Privacy-Preserving Vertical Federated Learning" Informatics 13, no. 8: 127. https://doi.org/10.3390/informatics13080127

APA Style

Azeez, N. A., Malomo, O. S., Okerinde, O. M., Ademoye, A. A., Aaron, D. S., Vyver, C. V. D., & Ogbonna, C. E. (2026). Adversarial Training and Differential Privacy-Style Noise Injection for Privacy-Preserving Vertical Federated Learning. Informatics, 13(8), 127. https://doi.org/10.3390/informatics13080127

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop