Next Article in Journal
An Exploratory Mixed-Methods Study of Sixth-Grade Primary School Students’ Problem-Solving Strategies and Difficulties with Loops in the Educational Programming Game Rapid Router
Previous Article in Journal
Link-Time Bytecode Quickening for Java Card Virtual Machines Without Method-Component Expansion: A Formal and Analytical Study
Previous Article in Special Issue
Robust Adversarial Attack Detection in Resource-Constrained IoT Ecosystems: A Privacy-Preserving Framework Using Federated Learning
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Collaborative Federated Learning to Secure 6G-IoT with Deep Convolutional Generative Adversarial Network

1
Department of Computer Science and Information, Computer Science and Engineering College, University of Ha’il, Ha’il 55211, Saudi Arabia
2
Department of Artificial Intelligence and Data Science, Computer Science and Engineering College, University of Ha’il, Ha’il 55211, Saudi Arabia
*
Author to whom correspondence should be addressed.
Computers 2026, 15(10), 644; https://doi.org/10.3390/computers15100644
Submission received: 2 August 2026 / Revised: 16 September 2026 / Accepted: 21 September 2026 / Published: 23 September 2026

Abstract

Wireless communication developments have significantly improved communications with each generation. The upcoming sixth-generation (6G) cellular wireless standard has raised expectations, particularly for Internet of Things (IoT) applications. 6G-based IoT (6G-IoT) aims to create an intelligent, ubiquitous, and self-optimizing IoT landscape. However, 6G-IoT faces security risks, and defending against attacks is increasingly challenging because of IoT devices’ distribution and heterogeneity. Most Federated Learning (FL)-based intrusion detection systems (IDSs) aim to enhance security, offering high reliability than local training. However, traditional FL suffers from high communication latency. FL over fog infrastructure is one solution. However, fog servers may not be available at every location, or communication with IoT devices may delay the learning process. This paper proposes a distributed IDS-based collaborative FL (IDS-CFL) across different learning levels: device and fog-cloud, to reduce data transfer. Neighboring devices at the device level collaborate, leveraging their computing capabilities for faster detection. To improve accuracy and enable fast processing, we propose a Deep Convolutional Generative Adversarial Network (DCGAN) model to train data at each level. The performance is evaluated on a recent dataset, Edge-IIoTest, and compared with other distributed and centralized methods. The results show the proposed system’s effectiveness, with 96.20% accuracy and 4.5 ms detection time.

1. Introduction

Wireless communication networks have evolved substantially from the first generation (1G) to the sixth generation (6G). Each generation of wireless communication addresses the limitations of the previous one. Recently, 6G has been the latest revolution in wireless communication. 6G is the anticipated successor to 5G, providing high data rate, lower latency, enhanced energy efficiency, high mobility, high capacity, enhanced security, and privacy [1]. It provides extensive Internet of Things (IoT) connectivity and supports the development of innovative IoT applications. 6G-based IoT (6G-IoT) aims to create a more intelligent, ubiquitous, and self-optimizing IoT landscape across domains such as healthcare, smart grids, smart buildings, industrial automation, and smart city applications [2].
The expansion of 6G-IoT introduces challenges for IoT data security, which in turn affect computational complexity, storage requirements, and processing costs. The basic components of 6G-IoT networks can be highly distributed and heterogeneous, connecting different types of IoT devices and handling large volumes of data. Most of this data is real-time data and may contain sensitive information that requires high security during transmission. Therefore, protecting 6G-IoT networks while transforming data is a primary concern [3].
Existing intrusion detection systems (IDSs) can be classified into two main categories: (1) IDSs based on a centralized learning approach that relies on Machine Learning (ML) or Deep Learning (DL), and (2) IDSs based on a distributed learning approach that relies on Federated Learning (FL). Centralized ML or DL methods are widely used to detect attacks because of their simplicity and accessibility. They identify abnormal traffic from devices present in the training dataset at a single detection point [4]. However, these methods face challenges such as potential single-point failure, high resource consumption, low scalability, difficulty detecting complex attacks, and limited practically in distributed environments [5].
The distributed detection approach based on FL can train the global model across distributed data. Multiple clients train the model through the central server. This approach can provide intelligent services for participating clients in 6G networks. In addition, sharing only model parameters rather than participants’ raw data provides greater security and privacy, which are critical requirements of 6G networks [6]. However, the communication latency in traditional FL is high. Some studies have addressed this problem by using edge or fog processing capabilities for partial model aggregation. This method reduces traffic in the 6G network; however, edge-cloud or fog-cloud infrastructures may not be accessible in some locations [7]. On the other hand, 6G cloud computing can leverage the computational capabilities of intelligent neighborhood devices, enabling faster recognition. However, the literature lacks an efficient FL mechanism that fully leverages 6G infrastructure, i.e., end devices, and edge or fog/central cloud [8]. In addition, current IDSs need improved accuracy.
In recent years, a paradigm shift towards using generative models, especially Generative Adversarial Network (GAN), has improved accuracy. A GAN consists of two models: a generator that produces realistic data from random noise, and a discriminator that distinguishes between real data from generated data. The generator model tries to trick the discriminator into classifying fake data as real, while the discriminator continues to improve at distinguishing real from fake data. This adversarial process enables GANs to produce high-quality synthetic data that closely resembles real-world data, which is useful across domains such as cybersecurity and attack detection [9]. However, training slowly identifies new attack types. A Deep Convolutional Generative Adversarial Network (DCGAN) is a variant of GAN that uses Convolutional Neural Network (CNN) in the generator and discriminator. It is used for high detection accuracy and fast processing [10,11]. A challenge in the literature is that very few studies have used GANs or DCGANs in attack detection models.
This paper proposes a distributed IDS based on collaborative FL (IDS-CFL) in 6G-IoT to reduce communication latency by minimizing data transfer. To meet high computation demands, it uses fog computing. The proposed IDS-CFL provides intelligent services at different levels, including end devices and fog-cloud. Neighboring devices at the device level collaborate through Machine-to-Machine (M2M) communication, leveraging their computing capabilities for faster detection. To improve accuracy and speed up processing, we propose a Deep Convolutional Generative Adversarial Network (DCGAN) model to train data at each level. To the best of our knowledge, the proposed IDS-CFL is the first work to perform FL across end devices and fog-cloud levels to reduce communication latency.
The contributions of this paper are as follows:
  • Propose a distributed IDS in 6G-IoT to reduce communication latency based on collaborative FL (IDS-CFL) with three levels, including end devices and fog-cloud.
  • Propose a Deep Convolutional Generative Adversarial Network (DCGAN) model to improve accuracy and enable fast processing.
  • Evaluate the proposed IDS-CFL using a recent, realistic cybersecurity dataset called Edge-IIoTest with distributed and centralized methods.

2. Related Works

As this research studies attack detection in 6G-IoT, the related work categorizes recent attack detection solutions into centralized, distributed, and collaborative learning approaches.

2.1. Centralized Learning Approach

Centralized learning models are widely used as detection solutions because of their simplicity and accessibility. Saeed et al. [12] proposed attack detection in the 6G (AD6Gs) wireless network using an ML algorithm. They used Correlation Feature Selection (CFS) to implement the proposed algorithm. The experimental result achieved 99% detection accuracy. Roa et al. [13] utilized a Generative Adversarial Network (GAN) to detect anomalies in a Wireless Body Area Network (WBAN) called GAN-WBAN. They used a CNN-based architecture for the generator and discriminator to capture spatiotemporal correlations for anomaly detection. The proposed approach achieves 97% detection accuracy.
Lin and Chen [14] proposed an inpainting-based anomaly detection system to identify defects without labeled defects. Their methodology centers on an image inpainting model that detects disparities between the original and restored versions of the defective image. They implemented the proposed detection system on Self-Supervised Learning (SSL), achieving 97% accuracy. Hinojosa et al. [15] proposed an unsupervised anomaly detection framework for early fault detection within such complex industrial settings. The proposed framework is particularly advantageous because it obviates the need for labeled historical fault data, a resource often limited in real-world environments. It improves the safety and reliability of the industrial environment.

2.2. Distributed Learning Approach

Distributed learning based on the FL approach proves its effectiveness for anomaly detection. Zhang et al. [16] introduced a decentralized FL framework to detect attacks in 6G-based Unmanned Vehicles (UxVs) networks. Clients collaborate to train the model without a central server. The experimental results outperform baseline approaches such as FedAvg, achieving 77% accuracy.
Garroppo et al. [17] proposed an FL-IDS based on a CNN model to detect attacks in IoT devices and smart-building networks. The proposed method ensures sustainability, adaptability, and trustworthiness. They applied different data preprocessing methods to reduce processing load and energy consumption. They designed CNNs to enable the investigation into AI reasoning and implemented eXplainable AI (XAI) techniques. The results show that the proposed method achieves 97.55% attack detection accuracy. Korba et al. [18] presented an intrusion detection system on the Internet of Vehicle (IoV), in anticipation of the upcoming 6G technology shift. The proposed methodology integrates class-incremental learning and FL. The authors implemented the experiment using a Multi-Layer Perceptron (MLP) model. The results demonstrate the proposed system’s robustness. In addition, it has 93% accuracy and a low false positive rate.
Ma et al. [19] proposed Federated Learning and Resource-aware Embedding (FLARE) for intrusion detection in a 6G-IoT driven healthcare system. FLARE provides three technical innovations: dynamic graph modeling with optimal temporal window selection, edge-aware aggregation mechanisms that enhance local feature representation, and adaptive multi-hop embedding strategies that address computational constraints. The proposed FLARE is implemented on a Large Language Model (LLM) using CICIDS2017 and UNSW-NB15, achieving 0.935 and 0.927 f-measure and 0.947 and 0.938 AUC-ROC. Jayarajan et al. [20] proposed a hierarchical FL approach in a smart grid-based 6G network to detect DDoS attacks. The proposed work combines a cloud-based service framework with the FL setup. The experiment shows that the proposed approach is suitable for real-world environments, achieving 92% accuracy and stability.
Prathiba et al. [21] proposed a Federated Learning (FL) and edge cache-assisted cybertwin (FLCC) framework for personalized service provision in 6G-Vehicle-to-Everthing environments. Integrating cybertwin technology into 6G enables seamless connectivity between physical systems and digital platforms, providing reliable, instantaneous wireless access. The FLCC framework incorporates edge cooperation and optimization using the proposed Federated Multi-agent Deep Reinforcement Learning (FM-DRL) algorithm. Experimental results indicate that the FLCC framework outperforms baseline approaches by 17.6%. Alatawi [22] presented a novel secure adaptive FL framework with integrated explainability, called SAFEL-IoT, designed for anomaly detection in Industrial Internet of Things (IIoT) systems. SAFEL-IoT is based on a dynamic aggregation mechanism and temporal model divergence. The results achieved 63.7 s training time and 93% accuracy.

2.3. Collaborative Federated Learning

Few works use collaborative FL to detect attacks in IoT. Kianpisheh and Taleb [23] proposed a collaborative FL (CFL) approach to provide intelligent services by combining different learning layers. The next device’s computational capability is expected to enable fast detection through 6G communication. The learning mode balances detection accuracy and response time for each device. They used a Gated Recurrent Unit (GRU) model to detect Distributed Denial of Service (DDoS) attacks. The results show that the proposed approach achieves an acceptable detection accuracy of 94%. Sedjelmaci et al. [24] proposed a cooperative FL approach to detect attacks in 6G-enabled IoT systems, accounting for key 6G network issues including latency, connectivity, data rate, and energy consumption. It is implemented through multi-level FL between IoT devices and edge computing applications. The proposed solution is compared with centralized approaches and achieves 98% detection accuracy.
Luo et al. [25] proposed a distributed framework to detect attacks on the IoT access side. They developed a personalized FL-based collaborative algorithm that enables horizontal collaboration between different detection points with network parameter sharing. The proposed framework addresses differences in IoT traffic distribution at the entrance. It uses local traffic data to train personalized models that improve local detection performance while leveraging the global model. The proposed framework achieved 99.2% accuracy and f-measure. Zakaria et al. [26] proposed a novel collaborative FL framework to provide privacy among multiple agents without data sharing. They proposed a secure aggregation mechanism to secure the FL aggregation service from reverse-engineering attacks. The proposed framework demonstrates effectiveness in accuracy and f-measure, achieving 99%. Table 1 compares existing attack detection approaches in a 6G-IoT environment.

2.4. Limitations of the Current Solutions

The literature indicates that current centralized models achieve high accuracy, but these systems are tightly coupled with their data centers, limiting real-time utility in distributed 6G-IoT environments. In contrast, most FL-based distributed learning models for attack detection suffer from high communication latency and lack an efficient detection mechanism that leverages 6G infrastructure. Existing collaborative edge- or fog-based FL approaches reduce communication latency by aggregating models at the network edge instead of transmitting parameter to the cloud. In contrast, end-devices-based FL approaches ignore the aggregator server, such as the cloud and rely on the device’s computation capabilities for aggregation, which is ineffective when the end device has limited computation capability. Therefore, we propose a distributed IDS based on collaborative FL (IDS-CFL) to minimize data transfer, with different levels to overcome these limitations.

3. Methodology

Figure 1 illustrates the distributed attack model in 6G-IoT networks. First, the attackers control a set of IoT devices distributed across various locations, known as botnets, to attack target servers. To detect attacks in 6G-IoT networks with low communication latency, we propose an intrusion detection system (IDS)-based collaborative Federated Learning (FL) (IDS-CFL). In addition, to improve accuracy, a Deep Convolutional Generative Adversarial Network (DCGAN) model is used to train the data. The following sections describe the details.

3.1. Intrusion Detection System Based Collaborative Federated Learning

As shown in Figure 2, the proposed IDS-based collaborative FL (IDS-CFL) architecture comprises device, fog, and cloud levels. The device level consists of IoT devices and actuators. The fog level consists of fog servers with high computational capabilities to minimize computation delays; it is collected with Base Stations (BSs). The data for each device u with size X u is represented with X u   =   { ( x u 1 , y u 1 ) ,   …   , ( x u X u , y u X u ) } , where x is the input data and y is the label for x . Each device u trains its local model M u with its data. At the device level, devices can share data with neighbors to enable faster detection. Devices communicate via M2M to share data. To avoid bandwidth consumption and protect privacy, FL does not allow sharing raw data. It is adapted to train the global model M . To address the availability of fog/cloud infrastructure, the system optimally selects a suitable data-sharing mechanism by enabling dynamic collaboration among the learning levels the device participates in. The details are given below.

3.1.1. Cloud-Level FL

FL at the cloud level aims to train the global model M c using data from the IoT devices. First, the cloud broadcasts the initial parameter w 0 to all U devices participating in the learning process. The learning process contains the local model parameter w c for M c that aims to minimize the global loss function L c w c as
min L c w c = 1 U c ∑ u = 1 U c 1 P u X u ∑ j = 1 P u X u 1 P u X u l w d ,   x uj , y uj
where P u is the local batch for the device u participating in cloud learning. l w ,   x uj ,   y uj is the loss value for data sample j for the device u . Each device minimizes the loss function on its own data using gradient descent, with a lower bound of 0 and an upper bound of 1. The devices participating in cloud learning transmit their local parameters to the BSs at the fog level for partial aggregation. Fog servers perform partial aggregation with the FedAvg algorithm. The partial aggregation of the model at the fog level M f is calculated as
M f = 1 U f ∑ u = 1 U c ( P u X u ) · w u
where U c is the number of devices under the coverage of BSs that participate in cloud learning. P u X u is the data size of device u included in the partial aggregation process. The w u is the model weight parameter of device u . The BSs then transmit the partial aggregations to the cloud for global aggregation. The global model M c is aggregated as
M c = 1 M ∑ i = 1 M M f
After the global model update, the cloud sends the updated parameters to the BSs, which then forward them to the devices included in the cloud learning. For fog-level learning, we can use the equations above. The BSs transmit M f model parameters to the devices under their coverage.

3.1.2. Device-Level FL

At the device level, the aggregator A and neighboring devices N d can perform global model aggregation. Learning process contains the local model parameter w d for M d that aims to minimize the global loss function L c w c as
min L d w d = 1 U d ∑ u = 1 U d 1 P u X u ∑ j = 1 P u X u 1 P u X u l w d ,   x uj , y uj
As in cloud-level FL, each device minimizes the loss function on its own data using gradient descent. Afterward, the randomly selected aggregator receives the local parameters w d from its neighboring devices to perform aggregation process. The aggregator calculates the weighted average of the neighboring devices’ parameters. All aggregators receive the parameters from their neighboring devices in parallel. The aggregation process for M d model at the device level is calculated as
M d = 1 U d ∑ u = 1 U d ( P u X u ) · w u · γ u
where γ u is the weighted average coefficient. For privacy, devices do not share raw data. Each device near the aggregator is included in the device-level FL. For simplicity, we assume all neighboring devices are included in the device-level FL.

3.2. Problem Formulation

As mentioned above, the current IDSs have multiple challenges, including communication latency and learning performance. The first optimization problem considers the communication latency, which is high in FL because of the large volume of data and parameters transmissions. By minimizing data transmissions, the detection time decreases. At each FL iteration, a device’s detection time is the time needed to receive the updated parameters and train the model. The first objective function is to minimize detection time T subject to a maximum CPU frequency L constraint. It can be modeled as
z = min   ( T )
subject to
L min   ≤   L u   ≤   L max
This time is calculated based on the device’s learning level. The following represent the calculations for the devices included in the learning.
The second optimization problem considers learning performance, which must be optimized to improve detection accuracy. The second objective function maximizes detection accuracy by minimizing loss function L ( w ) at each learning level, with a massively distributed data constraint where the global data X is stored across a large number of devices. It can be modeled as
J ( w ) =   min X u   ( L ( w ) )
subject to
∑ u = 1 U X u = X

3.2.1. Detection Time in CFL

To solve the first optimization problem, we minimize the detection time at each learning level, as explained below.
  • Cloud level:
  • Local training: The local training time at device u depends on the device’s data size. The local training time is calculated as
    T loc _ u = P u X u L u k L k
    where L k is the number of CPU cycles required to train one data sample, and L u k is the device CPU frequency.
  • Model aggregation: The partial aggregation time at BS i consists of the time required to transmit the parameters from the devices that participate in the cloud level under BS’s coverage area. The partial aggregation time at BS i is calculated as
    T agr i = w u R u , i + | R i | | w g | L i k L w k
    where w g is the size of global model, it is the same size as local model w u , L w k is the number of CPU cycles required to aggregate one unit of data, and L i k is the CPU frequency of BS i . The aggregation time at the cloud server is calculated as
    T agr   = T agr i + w g R c , i + w g L c k L w k
    where R c , i is the bandwidth communication between BS i and cloud server, and L c k is the CPU cycle at the cloud server. The aggregation time is affected by the large amount and diversity of data, which makes the system more scalable.
  • Parameter transmission: The required time depends on the parameters download at device u under BS i   coverage and is based on the parameter size. The required time for the transmission between the cloud server and BS i , and between BS i   and device u , is calculated as
    T tsm ( u ) = log 2 w u R c , i + w u R u , i
In the aggregation phase, the detection time in a single iteration for device u that takes part in the cloud level is calculated as
T u = max u T loc _ u + max u log 2 T tsm ( u )
2.
Device level: At the device level, the model parameters can be shared across neighborhood devices, with one device selected as the aggregator. The detection time for the aggregator u in one iteration is calculated as
T u = max n ∈ u T loc _ n + T agr u L u k
It includes the time required for local training at device u and neighborhood devices, as well as the time to transmit the model parameters from neighborhood devices. The aggregation time is represented as
T agr _ u =   max n ∈ N u w n R n , u + N u w g L u k L w k
where R n , u is the transmission rate of the neighborhood device n to connect to device u . When the neighborhood device n participates at the device level but is not selected as an aggregator, it downloads the parameters after the aggregation process. The detection time for the neighborhood device n is calculated as
T n =   max u ∈ N u T loc _ u - + w g R n , u + T agr u K n k
3.
Device selection algorithm: Collaboration across multiple levels can improve detection accuracy but may increase service delivery delays or cause service unavailability in locations far from cloud or fog servers. Therefore, device-level collaboration can minimize time consumption and enable faster detection. Moreover, selecting a device far from the aggregator takes time and adds transmission latency. Therefore, a Gray Wolf Optimizer (GWO) optimization algorithm is used for device selection at the device level. The GWO is a metaheuristic algorithm that mimics the social hierarchy and hunting behaviors of the grey wolves to catch prey in nature. It is used to solve various problems, including global optimization problems because it has fewer parameters, simple principles, and easy implementation [27]. In GWO, five solutions represent the devices in the search space, where alpha (α) is the best solution, beta (β) is the second-best solution, delta (δ) is the third-best solution, and omega (ω) is the rest of the solutions. The wolf represents aggregator A, which is selected randomly, while prey represents neighboring devices A n . The best three solutions (α, β, δ) are used to guide the other solutions (ω) to improve the search space. The selection steps are described as follows.
  • Encircling: Aggregator A starts by forming a circle around the neighboring devices A n when hunting. It is represented mathematically as
    E = C   ×   A n   t − A ( t )
    A t + 1 = A n t − S   ×   P    
    where t is the iteration number, A n is the neighboring device position, A is the aggregator position, P is used to specify a new position of the aggregator, and S and C are coefficient vectors that are calculated as
    S = 2 a   ×   r 1 − a  
    C = 2 r 2
    where r 1 and r 2 are random vectors in [0, 1], a is a vector decreased linearly from 2 to 0 over the iterations, and is calculated as
    a = 2 − t   ×   2 max itr
  • Hunting: In this step, the best three solutions (α, β, δ) are obtained. As for the other solutions (ω), they need to update their position by moving towards the average of the three best and known positions, since they have better knowledge about the optimal location of the neighboring device. This step is represented mathematically as
    A i t + 1 = A i t − a i   ×   P i
    V i = C i   ×   A i t − A ( t )
Let V i is the positive weight associated with aggregator i   ∈   { α , β , δ } such that ∑ i V i   =   1 . The aggregators’ positions α , β , and δ are a good estimation of the average position of the optimal solution at iteration t , it is represented as
A t + 1 = ∑ i ∈ { α , β , δ } V i   ×   A i ( t + 1 )
  • Attacking: The aggregators finish the hunt by attacking the neighboring device until they stop moving. To model the attacking process, Equation (20) is used, as the parameter a balance exploration and exploitation; a decrease linearly from 2 to 0 over iterations. Consequently, parameter S takes a random value in the range [ − 2 a ,   2 a ] given by Equation (18). When S   >   1 or S   <   − 1 , aggregators take a random position; when − 1   ≤   S   ≤   1 , they are forced to move toward the neighboring device.
Algorithms 1 and 2 describe the CFL procedure and device selection.
Algorithm 1: Collaborative FL
Input: Device u   ∈   U , initial global parameter w 0 , local patch P u of device u , data size X u of device u , input data X
Output: Updated global model
//Initialization
Cloud server initialize w 0
for each iteration t   =   0 ,   1 ,   2 ,   … ,   T  do
     //Cloud-Level FL
     Select a random set U of u devices
     Broadcast w 0 to all U devices participating at the cloud level
     Receive partial aggregation M f from the fog level and perform global aggregation M c using Equation (3)
     Updated w c and w u parameters
     Send updated w u to BSs participating in the learning
     //Fog-Level FL
     Each device u sends its local parameters w u to BSs
     Fog server performs partial aggregation M f using Equation (2)
     Send M f to cloud
     Receive updated w u from cloud
     Update parameters w f and w u
     Send w u to devices U participating in the learning
     //Device-Level FL
     Select aggregator A randomly
     Aggregator A receives w d from neighboring devices N d
     Aggregator A performs aggregation M d using Equation (5)
     Update parameters w d and w u
end for
Algorithm 2: Device selection using GWO
Input: Population size of aggregator devices A
Output: Optimally aggregator position A α from Neighboring devices N d
Initialize the aggregator population randomly A n
Initialize a , S , and C
Calculate the fitness of each search agent
A α is the best solution
A β is the second-best solution
A δ is the third-best solution
while  t   <   max itr  do
     for each aggregator A n  do
           Initialize r 1 and r 2 randomly
           Update current position using Equation (23)
     end for
Update a , S , and C
Calculate fitness of each aggregator A n
Update A α , A β , and A δ
t   =   t   +   1
end while
return A α (set of devices participating in device-level FL)

3.2.2. Deep Convolutional Generative Adversarial Network

To solve the second objective function and achieve high accuracy and fast processing, we use Deep Convolutional Generative Adversarial Network (DCGAN) model to train data at each level. It enables fast processing by replacing heavy fully connected layers with efficient convolutional layers. It can learn complex patterns and feature representations from large datasets. DCGAN improves detection performance by generating synthetic attack samples to train stronger detection models. The architecture of DCGAN is shown in Figure 3, it consists of a generator GE ( z ) and a discriminator DE ( x ) .
The generator takes a random noise vector z and upsamples it by using fractionally strided convolutions to output 2D synthetic data. The discriminator distinguishes between generated and real data. The DCGAN is mathematically described as
min GE   max DE L GE ( z ) , DE ( x )
The loss function L at each level is represented as
L = E x ~ X train [ log DE X ) ] + E z ~ X z [ log ⁡ ( 1 − D GE X ) ]  
where E is the expectation, X train is the training data distribution variable, x is the input vector, X z is the noise distribution variable, and z is a noise vector.
In the generator, the data z is sampled from the random X z as the generator input. It converts the 100-dimentional noise signal into ( 4   ×   4   ×   1024 ) 3D tensor using a fully connected layer; four convolution operation are applied sequentially, each with a different number of channels. Each operation uses a ( 4   ×   4 ) convolution kernel and a stride of 2, converting the feature map array into an output map array ( 64   ×   64   ×   3 ) . Batch normalization is introduced after each transpose convolution to improve training speed, stability, and performance. After that, the ReLU function introduces nonlinearity through setting negative values to 0 and keeping positive values unchanged; this helps the network learn patterns. The synthetic data X z ′   =   GE ( X z ) is generated, the synthetic generation process can be explained as
x z   l = v l w ge l M   ×   x z l − 1 + b ge l  
where x z l is the output data z at layer lth. w ge l and b ge l are the weight and bias parameters of lth convolutional layer. v l is the ReLU activation function for all layers except the output layer, which uses Tanh activation.
Now, the synthetic data X z ′ is combined with the input data X train to generate combined data X de that used to train the discriminator. It is configured with three ( 2   ×   2 ) convolution kernel with a stride of 2.
The discriminator has four convolutional layers and takes a ( 64   ×   64   ×   3 ) input. The four convolutional layers use LeakyReLU except the output layer, which uses Sigmoid. Each operation uses a ( 4   ×   4 ) convolution kernel and a stride of 2, converting the feature map array into an output map array ( 1   ×   1 ) . After training, the discriminator generates the discriminator data Y de where Y de   =   DE ( X de ) . Finally, the data from the discriminator needs to be preprocessed, which can be defined as
Y de = y l 1 = 0 ,   y l 2 = 1                 i f   y l 1   <   y l 2 y l 1 = 1 ,   y l 2 = 0                 i f   y l 1   ≥   y l 2
where y l 1 and y l 2 are the discriminator outputs of the lth data of the discriminator. The output Y de is contain Y z and Y train , where Y z is the discriminant output of synthetic data X z ′ and Y train is the discriminant output of data X train . Now, the discriminator performance of Y z and Y train are calculated using accuracy ACU function as
ACU = TP + TN TP + TN + FP + FN
The second objective function is improving accuracy by obtaining optimal discriminator performance; we can redefine the loss function as
L GE , DE = E x ~ X train [ log ( ACU ( DE X ) ) ] + E z ~ X z [ log ⁡ ( 1 − ACU ( DE GE X ) ) ]
The optimal discriminator DE ge * is obtained as
DE ge * = arg max de L DCGAN ( GE de * ,   DE )
where GE de * is the optimal generator performance which is obtained as
GE de * = arg min ge L DCGAN ( GE ,   DE ge * )  
Algorithm 3 illustrates the overall DCGAN training process.
Algorithm 3: DCGAN algorithm
Input: Number of epochs ( n _ epochs ), batch size ( n _ batch ), initial generator weight θ ge , initial discriminator weight θ de , generator ( GE ), discriminator ( DE ), training data distribution ( X train ), noise distribution X z , input vector ( x ), noise vector ( z )
Output: Optimally trained DE ge * and GE de *
//Initialization
Initialize θ ge and θ de
                            for each epoch from 1 to n _ epochs  do
                                Shuffle training data and create minibatches
                                    for each batch from 1 to n _ batch  do
                                         //Generator training
                                                 Sample minibatch of m noise samples z 1 , z 2 , … , z m from X z
                                                   Sample minibatch of m real data samples x 1 , x 2 , … , x m from X train
                                                   Generate synthetic data { GE ( z 1 ) , GE ( z 2 ) , … , GE ( z m ) } using Equation (26)
                                                   Combine synthetic data X z ′ with X train to generate combined data X de
                                                   Update DE using Adam: ∇ θ d 1 m ∑ i = 1 m [ log DE x i + log ⁡ ( 1 − DE GE z i ) ]
                                         //Discriminator training
                                                   Evaluate discriminator output Y de using Equation (29)
                                                   Sample minibatch of m noise samples z 1 , z 2 , … , z m from X z
                                                   Generate fake data samples { GE ( z 1 ) , GE ( z 2 ) , … , GE ( z m ) }
                                                   Update GE using Adam:
                                                         ∇ θ g 1 m ∑ i = 1 m log ⁡ ( 1 − DE GE z i )
                                         end for
                                Calculate DE ge * using Equation (30)
                                Calculate GE de * using Equation (31)
                                Save DE ge * and GE de *
                            end for

4. Data Preparation

This section describes the dataset used for evaluation and the preprocessing steps used to prepare the training and testing sets.

4.1. Data Description

This paper uses a recent, comprehensive, realistic cybersecurity dataset for IoT and IIoT applications, called Edge-IIoTest [28]. It supports intrusion detection systems in both centralized and FL modes. The dataset is generated using a purpose-built IoT/IIoT testbed with a large representative set of devices, sensors, protocols, and cloud/edge configurations. Multiple devices (over 10 types) generate the IoT data. It includes five types of threats: DoS/DDoS attacks, Information gathering, Man in Middel attacks, Injection attacks and Malware attacks. The dataset contains 61 features from different sources, with 1176 features found to be highly correlated.

4.2. Data Preprocessing

IoT traffic includes normal traffic and various types of attacks. We apply the following preprocessing techniques to prepare the dataset for training.
  • Feature mapping: The features in IoT data do not consist only of numeric values. Therefore, a mapping technique is required to convert categorical values to numeric values. A common method is One-Hot Encoding (OHE), which converts each distinct value into a binary value.
  • Data normalization: Normalization is a scaling method that converts all data features onto a common scale. We removed outliers, such as null values and non-numeric entries in numerical attributes. For a fair comparison, we have normalized the dataset using Min-Max scaling. This popular method facilitates arithmetic processing by linearly mapping each feature’s range to 0–1. It can be calculated as
X scale = X − X min X max   −   X min
3.
Feature Selection: It improves predictive quality by selecting relevant features. Feature selection is the process of selecting a subset of features important for solving the detection problem and discarding unneeded features. In this paper, we use a mutual information (MI) technique for feature selection, with selection criteria based on feature dependencies: features with high mutual information are considered the best features. The input–output variables from the training set can be represented as X and Y where R = { X ,   Y } . To compute MI, follow these steps.
Step 1: Measure the dependencies between input and output variables MI   =   { X i ,   Y } as
MI ( X , Y ) = ∑ x ∈ X ∑ y ∈ Y PR ( x ,   y ) log PR ( x ,   y ) PR ( x ) × PR ( y )
where PR ( x ) and PR ( y ) are the marginal distributions, and PR ( x ,   y ) is the joint probability of X and Y .
Step 2: Remove any values less than other values and maintain the remaining in a vector V . Now, the input and output variables can be represented as R new , sorted in descending order based on the value of V .
Step 3: Define the number of selected features represented as F and select the first input–output F variable from R new as a selected features to form R F .
Step 4: Select the variable with the highest MI as a first variable in R . Now, select jth variables ( 1   ≤   j   ≤   F ) from R F and compute MI ( S ,   X j ) values. If MI ( S ,   X j )   ≤   α , R   ∪   X j , where α is the threshold value. Otherwise,   X j will have less dependency and need to be deleted.
4.
Class Balancing: To prevent bias toward the majority “Normal” class, the script identifies the minority class size and samples from it. The Synthetic Minority Over-sampling Technique (SMOTE) [29] is applied to handle the imbalanced dataset. It is based on the K-Nearest Neighbors (KNN) model, which takes samples and considers their nearest neighbors in the feature map. If ( x 1 ,   x 2 ) is a sample of a minority class and if its neighbor is selected as ( x 1 * , x 2 * ) , then the data point is synthesized as
( X 1 ,   X 2 ) =   ( x 1 ,   x 2 ) + random ( 0 ,   1 )   ×   ∆
where ∆   =   { ( x 1 * , x 1 ) ,   ( x 2 * , x 2 ) } , random ( 0 ,   1 ) is a random value between 0 and 1. Figure 4 shows the data before and after data balancing.

5. Experimental Results

This section describes the experimental setup and compares the proposed IDS-CFL with distributed and centralized methods across different experiments.

5.1. Experiment Setup

For evaluation, the experiment is conducted in a 6G-IoT simulation environment combining cloud, fog, and device levels. The NS-3 simulator was used to build a 6G network environment with 50 distributed devices collaborating in the detection process. The simulation environment was configured with network parameters that align with 6G specifications. The CFL channel model is based on uplink massive MIMO technology to achieve the high throughput and efficiency needed for 6G-IoT networks. The channel model consists of 5 BSs equipped with fog servers located between the cloud and device level, with a CPU frequency of 2.5 GHz. All fog nodes have identical computing capabilities. The transmission bandwidth is 100 MHz, and the coverage radius is 250 m. The devices’ CPUs are randomly selected between 1.5 to 2.4 GHz based on device capabilities. We use TensorFlow and Keras DL libraries to implement the DCGAN model. We adopt an 80:20 split, allocating 80% of the data for training and 20% for testing. We found that a learning rate of 0.01 is appropriate for our system. The parameter settings utilized in the implementation are detailed in Table 2.

5.2. Performance Metrics

We applied different performance metrics to evaluate the effectiveness of the proposed system, including detection time, accuracy, recall, f-measure (F1), and Area Under Curve (AUC).

5.3. Performance Analysis

The performance of the proposed IDS-CFL system is evaluated and compared with different methods, including
  • No FL: In this model, centralized learning is performed where each device sends its data to the cloud for model training.
  • Fog-level FL: In this model, the traditional FedAvg method is performed at the fog level. The devices send their data to the fog level for model training. After training, the fog servers send the local parameters to the cloud for global model aggregation. After that, the cloud sends the updated parameters to the fog level to update the model.
  • Device-level FL: In this model, at each FL iteration, the devices train the model using their data and, instead of sending local model parameters to the cloud or fog, share their parameters. The selected aggregator receives the parameters from neighboring devices and aggregates the model. After that, the aggregator sends the results back to the neighboring devices for attack detection. The aggregator device is selected randomly, and neighboring devices are selected based on the k-means clustering as in [30].

5.4. Results

The experimental performance of the proposed CFL is compared with the other methods through different experiments to evaluate its effectiveness. These experiments are explained in the following.

5.4.1. Experiment 1: Effect of Collaboration in CFL

In this experiment, we examined the involvement of devices and collaboration across different learning levels in FL using two scenarios (S1 and S2), as shown in Table 3.
  • S1: In S1, there are no devices in No FL involved in learning. In fog-level FL, a high portion of devices is involved in FL (70%), whereas device-level FL involves the lowest portion (30%). This is because 30% of devices have neighbors. CFL has the highest device involvements, at 90% in FL. BSs can cover up to 80% of devices, and devices without BS access can collaborate in FL through their neighboring devices, if available. Also, if a device has no neighbors, it can participate in FL at the fog level if fog infrastructure is available. For our CFL, 10% of devices are not covered by any BS or neighbors.
Figure 5a compares the proposed CFL with the other methods with S1 in terms of detection time. No FL achieves the lowest detection time with 3.8 ms, as it performs only local training. Device-level FL achieves 7.5 ms detection time because it takes time to transmit parameters. Training takes around 4.5 ms, and aggregation takes 3 ms. Fog-level FL has the highest detection time, around 8.3 ms, due to the overhead transmitting parameters to the BSs. The proposed CFL outperforms the device level and fog level with 4.5 ms. This is because the CFL enables collaboration between the device level and fog level, which minimizes detection time. In addition, using GWO algorithm helps to reduce detection time and increase accuracy compared with k-mean clustering selection method.
As shown in Figure 6a, our proposed CFL model achieves the highest values: 96.20% accuracy, 94.10% recall, 95.90% f-measure, and 97.05% AUC. No FL achieves the lowest performance because it does not share data. Fog-level FL outperformed the device-level FL with 94.80% accuracy, 92.99% recall, 93.52 f-measure, and 95.03% AUC. This is because fog-level FL shares parameters across 70% of devices, compared with 40% in device-level FL. The proposed CFL learns to optimize accumulated detection accuracy by exploiting fog-level FL.
2.
S2: In S2, the fog servers’ capabilities are available for 90% of devices, whereas 40% of devices have a neighboring. As shown in Figure 5b, there is no significant difference in detection time compared with S1. The time decreases by around 6.6 ms in device-level FL and 7.2 ms in fog-level FL. As shown in Figure 6b, the accuracy and the other metrices are increased in fog-level FL which scored 97.99%. In device-level FL, performance is better than in S1, with 90.06% accuracy when 40% of devices have a neighborhood. Compared with fog-level FL, our CFL reduced communication because it shares parameters across 90%. On the other hand, it outperformed No FL and device-level FL. Table 4 summarized the comparison results.

5.4.2. Experiment 2: Comparing DCGAN with Other Models

In this experiment, we apply the proposed CFL to the GAN and CNN models to validate DCGAN’s performance. The traditional GAN is performed as in [31]. The CNN architecture comprises four convolutional layers; each pair of convolutional layers is followed by one max-pooling layer. Conv 1 and Conv 2 operated with (convolution filters = 16, filter size = 3, and output shape = 16 × 196 feature matrix), max-pooling 1 with pooling size = 2, Conv 3 and Conv 4 operated with (convolution filters = 32, filter size = 3, and output shape = 16 × 128 feature matrix), and max-pooling 2 with pooling size = 2. For fully connected layers, the model consists of three layers, each followed by a dropout layer. One padding is applied before each convolution operation, and ReLU is used as the activation function in the hidden layers. Sigmoid activates the last fully connected layer.
As shown in Figure 7, CNN outperformed DCGAN and GAN with 3.5 ms, with training time of about 2 ms and aggregation time of 1.5 ms. This is because CNN solves a simple mapping problem with a single network, whereas DCGAN and GAN involve training two adversarial networks simultaneously. Figure 8 shows that the CFL with DCGAN outperformed the GAN and CNN models in accuracy, recall, f-measure, and AUC. The GAN model had the highest values after DCGAN, with scored 94.32% accuracy, 92.16% recall, 92.98% f-measure, and 94% AUC. The CNN model performed the worst among the DCGAN and GAN models. Table 5 summarizes the comparison results of different models.

5.4.3. Experiment 3: Examine the Scalability with Increasing Number of Devices

In this experiment, we test the scalability of the proposed CFL with the other methods. We increased the CFL size to 100 and 200 devices measure accuracy and detection time over 10 rounds. The experiment has been executed 10 times for each method size in each round. For devices involvements, we use the same involvements of S1 as in Table 3. Figure 9 shows average accuracy and detection time; the proposed CFL shows consistent performance stability as the number of devices increases. Compared with the other methods, our CFL outperforms them in terms of scalability. This experiment proves that the proposed CFL is scalable and stable with a large number of devices in FL. Therefore, as the number of devices in FL increases, the model learns all types of attacks more accurately. Although the CFL is constant, accuracy drops below 95% when the number of devices exceeds 200.
In summary, our proposed system can achieve performance similar to a model trained on global data by leveraging the collaborative efforts of models trained on localized data. The proposed IDS-CFL has achieved the minimum of the objective function. This is because IDS-CFL optimizes the detection accuracy and detection time as the objective function. However, the proposed CFL model with DCGAN had the highest detection time among the CNN and GAN models. It achieves the best accuracy among the other models. Thus, it achieves a tradeoff between accuracy and detection time. Although the proposed CFL achieves acceptable performance, it has limitations, such as maintaining consistent accuracy and real-time performance. The distribution and heterogeneity of IoT data across different device capabilities makes it challenging to maintain constant accuracy. Although the proposed CFL is trained on a non-IID dataset, it needs to be tested on real-time data to prove its effectiveness.

6. Conclusions

In this paper, a distributed IDS-based collaborative FL (IDS-CFL) was proposed to detect attacks in 6G-IoT networks. The collaboration is performed among three levels: device, fog, and cloud levels to reduce communication latency. A Deep Convolutional Generative Adversarial Network (DCGAN) model was proposed to improve detection accuracy and enable faster processing across different learning levels. To evaluate the performance of our system, we conducted different experiments comparing the proposed system with distributed and centralized methods. To demonstrate the effectiveness of the DCGAN model, we compared it with other models, such as a basic Convolutional Neural Network (CNN) and Generative Adversarial Network (GAN). The performance of the proposed system was evaluated on a recent Edge-IIoTest dataset, achieving the highest performance with 96.20% accuracy and 4.5 ms detection time. As future work, we aim to address the limitations. We will evaluate the proposed system in real 6G-IoT scenarios with advanced convergence strategy to maintain accuracy constant, using real-time data.

Author Contributions

Conceptualization, R.A. and B.A.; methodology, R.A.; software, A.A.; validation, R.A., B.A. and A.A.; formal analysis, R.A.; investigation, A.A.; resources, B.A.; data curation, R.A.; writing—original draft preparation, R.A.; writing—review and editing, R.A., B.A. and A.A.; visualization, B.A.; supervision, R.A.; project administration, R.A.; funding acquisition, R.A. All authors have read and agreed to the published version of the manuscript.

Funding

This research has been funded by Scientific Research Deanship at University of Ha’il -Saudi Arabia through project number BA-26 004.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Acknowledgments

This research has been funded by Scientific Research Deanship at University of Ha’il -Saudi Arabia through project number BA-26 004.

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. Tera, S.P.; Chinthaginjala, R.; Pau, G.; Kim, T.H. Towards 6G: An Overview of the Next Generation of Intelligent Network Connectivity. IEEE Access 2024, 13, 925–961. [Google Scholar] [CrossRef] [Scilit]
  2. Nguyen, D.C.; Ding, M.; Pathirana, P.N.; Seneviratne, A.; Li, J.; Niyato, D.; Dobre, O.; Poor, H.V. 6G Internet of Things: A Comprehensive Survey. IEEE Internet Things J. 2021, 9, 359–383. [Google Scholar] [CrossRef] [Scilit]
  3. Junior, E.E. Systematic Review of 6G-IoT Privacy Risks, Emerging Threats, Mitigation Strategies, and Cybersecurity. SSRN Electron. J. 2025, 19, 180–190. [Google Scholar] [CrossRef] [Scilit]
  4. Assiri, M. Artificial intelligence-based intrusion detection and secure communication model for sustainable 6G-IoT networks. Sci. Rep. 2026, 16, 12662. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Wu, Y.; Chen, J.; Lei, T.; Yu, J.; Hossain, M.S. Web 3.0 security: Backdoor attacks in federated learning-based automatic speaker verification systems in the 6G era. Future Gener. Comput. Syst. 2024, 160, 433–441. [Google Scholar] [CrossRef] [Scilit]
  6. Edegbe, G.N.; Acheme, S. A systematic review of centralized and decentralized machine learning models: Security concerns, defenses and future directions. NIPES-J. Sci. Technol. Res. 2024, 6, 161–175. [Google Scholar] [CrossRef]
  7. Chen, K.; Liu, Y. Toward Privacy-Preserving AI Standards for Federated Learning in 6G-Enabled Digital Twin Environments. IEEE Commun. Stand. Mag. 2025, 10, 265–272. [Google Scholar] [CrossRef] [Scilit]
  8. de Alwis, C.; Aouedi, O.; Xu, J.; Wang, S.; Siriwardhana, Y.; Hewa, T.; Zeydan, E.; Sandeepa, C.; Liyanage, M. Federated Learning for 6G Security: A Survey on Threats, Solutions, and Research Directions. IEEE Commun. Surv. Tutor. 2026, 28, 4883–4914. [Google Scholar] [CrossRef] [Scilit]
  9. Tomkos, I.; Klonidis, D.; Pikasis, E.; Theodoridis, S. Toward the 6G network era: Opportunities and challenges. IT Prof. 2020, 22, 34–38. [Google Scholar] [CrossRef] [Scilit]
  10. Creswell, A.; White, T.; Dumoulin, V.; Arulkumaran, K.; Sengupta, B.; Bharath, A.A. Generative adversarial networks: An overview. IEEE Signal Process. Mag. 2024, 35, 53–65. [Google Scholar] [CrossRef] [Scilit]
  11. Wu, Y.; Nie, L.; Wang, S.; Ning, Z.; Li, S. Intelligent Intrusion Detection for Internet of Things Security: A Deep Convolutional Generative Adversarial Network-enabled Approach. IEEE Internet Things J. 2021, 10, 3094–3106. [Google Scholar] [CrossRef] [Scilit]
  12. Saeed, M.M.; Saeed, R.A.; Gaid, A.S.A.; Mokhtar, R.A.; Khalifa, O.O.; Ahmed, Z.E. Attacks Detection in 6G Wireless Networks Using Machine Learning; IGI Global: Hershey, PA, USA, 2023; pp. 6–11. [Google Scholar] [CrossRef] [Scilit]
  13. Rao, V.A.; Rao, R.; Hota, C. Anomaly detection in wireless body area networks using generative adversarial networks. In Proceedings of the 2024 IEEE International Conference on Industry 4.0, Artificial Intelligence, and Communications Technology (IAICT), Bali, Indonesia, 4–6 July 2024. [Google Scholar]
  14. Lin, C.-Y.; Chen, C.-Z. Inpainting-based anomaly detection system with self-supervised learning. In Proceedings of the 2024 IEEE International Conference on Industry 4.0, Artificial Intelligence, and Communications Technology (IAICT), Bali, Indonesia, 4–6 July 2024. [Google Scholar]
  15. Hinojosa-Palafox, E.A.; Rodríguez-Elías, O.M.; Pacheco-Ramírez, J.H.; Hoyo-Montaño, J.A.; Pérez-Patricio, M.; Espejel-Blanco, D.F. A Novel Unsupervised Anomaly Detection Framework for Early Fault Detection in Complex Industrial Settings. IEEE Access 2024, 12, 181823–181845. [Google Scholar] [CrossRef] [Scilit]
  16. Zhang, J.; Luo, C.; Jiang, Y.; Min, G. Decentralized Federated Learning for Intrusion Detection in 6G-based UxV Networks. IEEE Veh. Technol. Mag. 2025, 20, 83–93. [Google Scholar] [CrossRef] [Scilit]
  17. Garroppo, R.G.; Giardina, P.G.; Landi, G.; Ruta, M. Trustworthy AI and Federated Learning for Intrusion Detection in 6G-Connected Smart Buildings. Future Internet 2025, 17, 191. [Google Scholar] [CrossRef] [Scilit]
  18. Korba, A.A.; Sebaa, S.; Mabrouki, M.; G-Doudane, Y.; Benatchba, K. A Life-long Learning Intrusion Detection System for 6G-Enabled IoV. In Proceedings of the 2024 International Wireless Communications and Mobile Computing (IWCMC); IEEE: New York, NY, USA, 2024; pp. 1773–1778. [Google Scholar] [CrossRef] [Scilit]
  19. Ma, X.; Hu, J.; Liang, S.; Wu, Y. Federated Learning and Resource-Aware Graph Neural Network for Intrusion Detection in 6G-IoT Driven Healthcare System. IEEE Internet Things J. 2025, 13, 7749–7761. [Google Scholar] [CrossRef] [Scilit]
  20. Jayarajan, J.; Mahalingam, N.; Seng, Y.K. Empowering Smart Grid Security: Towards Federated Learning in 6G-Enabled Smart Grids using Cloud. Res. Sq. 2024; preprint. [CrossRef] [Scilit] [PubMed]
  21. Prathiba, S.B.; Raja, G.; Anbalagan, S.; Gurumoorthy, S.; Kumar, N.; Guizani, M. Cybertwin-Driven Federated Learning Based Personalized Service Provision for 6G-V2X. IEEE Trans. Veh. Technol. 2022, 71, 4632–4641. [Google Scholar] [CrossRef] [Scilit]
  22. Alatawi, M.N. SAFEL-IoT: Secure Adaptive Federated Learning with Explainability for Anomaly Detection in 6G-Enabled Smart Industry 5.0. Electronics 2025, 14, 2153. [Google Scholar] [CrossRef] [Scilit]
  23. Kianpisheh, S.; Taleb, T. Collaborative Federated Learning for 6G With a Deep Reinforcement Learning Based Controlling Mechanism: A DDoS Attack Detection Scenario. IEEE Trans. Netw. Serv. Manag. 2024, 21, 4731–4749. [Google Scholar] [CrossRef] [Scilit]
  24. Sedjelmaci, H.; Kheir, N.; Boudguiga, A.; Kaaniche, N. Cooperative and smart attacks detection systems in 6G-enabled Internet of Things. In Proceedings of the ICC 2022-IEEE International Conference on Communications, Seoul, Republic of Korea, 16–20 May 2022; Available online: https://ieeexplore.ieee.org/abstract/document/9838338/ (accessed on 21 September 2023).
  25. Luo, Y.; Chen, X.; Sun, H.; Li, X.; Ge, N.; Feng, W.; Lu, J. Securing 5G/6G IoT Using Transformer and Personalized Federated Learning: An Access-Side Distributed Malicious Traffic Detection Framework. IEEE Open J. Commun. Soc. 2024, 5, 1325–1339. [Google Scholar] [CrossRef] [Scilit]
  26. El Houda, Z.A.; Naboulsi, D.; Kaddoum, G. A Privacy-Preserving Collaborative Jamming Attacks Detection Framework Using Federated Learning. IEEE Internet Things J. 2024, 11, 12153–12164. [Google Scholar] [CrossRef] [Scilit]
  27. Faris, H.; Aljarah, I.; Al-Betar, M.A.; Mirjalili, S. Grey wolf optimizer: A review of recent variants and applications. Neural Comput. Appl. 2017, 30, 413–435. [Google Scholar] [CrossRef] [Scilit]
  28. Ferrag, M.A.; Friha, O.; Hamouda, D.; Maglaras, L.; Janicke, H. Edge-IIoTset: A New Comprehensive Realistic Cyber Security Dataset of IoT and IIoT Applications for Centralized and Federated Learning. IEEE Access 2022, 10, 40281–40306. [Google Scholar] [CrossRef] [Scilit]
  29. Abunada, M.; Belhaouari, S.B.; Bensmail, H. Synthetic Minority Oversampling for Imbalanced Time Series Classification Based on Path Signature. Appl. Sci. 2026, 16, 4451. [Google Scholar] [CrossRef] [Scilit]
  30. Trindade, S.; da Fonseca, N.L.S. Multicriteria Scoring for Cluster and Client Selection in Heterogeneous Hierarchical Federated Learning. IEEE Internet Things J. 2026, 13, 16763–16779. [Google Scholar] [CrossRef] [Scilit]
  31. Kaur, R. Generative Adversarial Network (GANs) for Image Generation or Data Augmentation. Int. J. Sci. Archit. Technol. Environ. 2025, 3, 509–517. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Distributed attack model in 6G-IoT network.
Figure 1. Distributed attack model in 6G-IoT network.
Computers 15 00644 g001
Figure 2. Collaborative FL architecture.
Figure 2. Collaborative FL architecture.
Computers 15 00644 g002
Figure 3. DCGAN architecture.
Figure 3. DCGAN architecture.
Computers 15 00644 g003
Figure 4. Data distribution (a) before and (b) after data balancing.
Figure 4. Data distribution (a) before and (b) after data balancing.
Computers 15 00644 g004
Figure 5. Detection time comparison in (a) S1 and (b) S2.
Figure 5. Detection time comparison in (a) S1 and (b) S2.
Computers 15 00644 g005
Figure 6. Methods performance comparison in (a) S1 and (b) S2.
Figure 6. Methods performance comparison in (a) S1 and (b) S2.
Computers 15 00644 g006
Figure 7. Detection time of CNN, GAN, and DCGAN in CFL.
Figure 7. Detection time of CNN, GAN, and DCGAN in CFL.
Computers 15 00644 g007
Figure 8. Performance of CNN, GAN, and DCGAN in CFL.
Figure 8. Performance of CNN, GAN, and DCGAN in CFL.
Computers 15 00644 g008
Figure 9. Scalability performance.
Figure 9. Scalability performance.
Computers 15 00644 g009
Table 1. Comparison of attack detection solutions in 6G-IoT environment.
Table 1. Comparison of attack detection solutions in 6G-IoT environment.
RefWorkLearning ModelDatasetLimitations
[12]Attack detection in 6G (AD6Gs)RF and SVMCICDDoS2019It is difficult to identify new threats
[13]GAN-WBAN to detect anomaliesCNNMIMICIt is not practical with a large number of points
[14]Anomaly detection systemSSLMVTecHigh transmission overhead
[15]Anomaly detection framework in industrial environmentML2015 PHM Data ChallengeUsing unsupervised learning lacks to detect anomalies, especially those that significantly deviate from previously observed patterns
[16]Decentralized FL for intrusion detection in UxVsMLAWID-3It is not practical with complex attacks
[17]FL-IDS to detect attacks in smart buildingsCNNToN-IoTThere is no clear correlation between the impact of attacks on network and telemetry data
[18]Intrusion detection system in IoVMLP5G-NIDDIt is not effective when the number of clients increases
[19]FLARE for intrusion detection in 6G-IoT driven healthcare systemLLMCICIDS2017 and UNSW-NB15The training time is high for real-time applications such as healthcare systems
[20]A hierarchical FL approach in smart grid for 6G network to detect DDoS attacksCNNCICDDoS2019It is suitable only for a small number of clients
[21]FLCC to provide security in 6G-V2XDRLCIFAR10It is not practical with complex attacks
[22]SAFEL-IoT for anomaly detectionAESKABThe adaptive aggregation mechanisms remain limited, and formal proofs under non-convex optimization settings are necessary to establish robust guarantees
[23]Collaborative FL to detect DDoS attacksGRUCICDDoS2019Some accuracy loss might be experienced due to the issue of more locality in constructing the parameters of the model
[24]Cooperative FL to detect attacks in 6G-Enabled IoTMLUNSW-NB15It suffers from high communication latency
[25]Personalized FL-based collaborative algorithm for attack detectionTransformerN-BaIoTThe global detection ability is not accurate, and it is not effective when the clients are highly distributed
[26]Collaborative FL to enable privacy-aware distributed learningCNNWSN-DSThe response time and communication latency are high
Table 2. Parameters setting.
Table 2. Parameters setting.
ParameterValue
No. of devices50
No. of BSs5
Data rate10 Mbps
Fog CPU frequency2.5 GHz
Device CPU frequencyfrom 1.5 to 2.4
Transmission bandwidth100 MHz
Learning rate0.01
Batch size128
Epochs10
OptimizerAdam
Table 3. Devices’ involvement during learning levels.
Table 3. Devices’ involvement during learning levels.
MethodTotal Involved Device Ratio in FLDevice-Level Collaboration RatioFog-Level Collaboration Ratio
No FL000
Fog-level FLS1: 70%, S2: 90%0100%
Device-level FLS1: 30%, S2: 40%100%0
CFL90%25%75%
Table 4. Methods comparative results.
Table 4. Methods comparative results.
Method/MetricsAccuracy (%)Recall (%)F1 (%)AUC (%)DT (ms)
No FL75.0273.1174.0577.133.8
Fog-level FL94.8092.9993.5295.038.3
Device-level FL85.0982.1284.1985.037.5
Proposed CFL96.2094.1095.9097.054.5
Table 5. Comparison results of different models in CFL.
Table 5. Comparison results of different models in CFL.
Model/MetricsAccuracy (%)Recall (%)F1 (%)AUC (%)DT (ms)
CNN80.1476.9979.5681.173.5
GAN94.3292.1692.98944.1
DCGAN96.2094.1095.9097.054.5
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Almarshdi, R.; Alrashidi, B.; Alamr, A. Collaborative Federated Learning to Secure 6G-IoT with Deep Convolutional Generative Adversarial Network. Computers 2026, 15, 644. https://doi.org/10.3390/computers15100644

AMA Style

Almarshdi R, Alrashidi B, Alamr A. Collaborative Federated Learning to Secure 6G-IoT with Deep Convolutional Generative Adversarial Network. Computers. 2026; 15(10):644. https://doi.org/10.3390/computers15100644

Chicago/Turabian Style

Almarshdi, Rasha, Bedour Alrashidi, and Abrar Alamr. 2026. "Collaborative Federated Learning to Secure 6G-IoT with Deep Convolutional Generative Adversarial Network" Computers 15, no. 10: 644. https://doi.org/10.3390/computers15100644

APA Style

Almarshdi, R., Alrashidi, B., & Alamr, A. (2026). Collaborative Federated Learning to Secure 6G-IoT with Deep Convolutional Generative Adversarial Network. Computers, 15(10), 644. https://doi.org/10.3390/computers15100644

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop