Replacing the Genetic Algorithm with Multi-Objective Bacterial Foraging Optimization in XCS
Highlights
- BFOA replaces GA in XCS, creating a novel classifier optimization system that evaluates multiple fitness criteria simultaneously through a weighted sum scalarization.
- BFOA enables concurrent optimization of classifier accuracy, stability, and variance.
- BFOA-XCS achieves the top Friedman ranking among XCS variants across 19 benchmark datasets, with notable variance reduction (15.2%) over standard GA-XCS.
- In dynamic cybersecurity defense, XCS variants significantly outperform three of five deep RL baselines (DQN, Q-Learning, Policy Gradient (REINFORCE)), while PPO and SAC achieve higher overall rewards at substantially higher computational cost.
- Swarm intelligence offers a viable alternative to evolutionary operators in LCS research.
- Rule-based XCS provides competitive cybersecurity defense with interpretable policies, warmup-free competence, and 5–26× lower computational overhead compared to the top-performing neural approaches (PPO at 5.3× and SAC at 26.1× the XCS compute time per training run).
- The BFOA-XCS hybrid generalizes across diverse static and dynamic environments.
Abstract
1. Introduction
- Replacing the GA component in XCS with BFOA, extending the LCS research domain with a bio-inspired swarm intelligence algorithm. To the best of the authors’ knowledge, this is the first integration of BFOA with the LCS. The integration pattern itself is independent of the specific swarm optimizer chosen. BFOA is one instantiation, and the same template can in principle accommodate other population-based optimizers.
- Introducing a weighted multi-criteria fitness function that leverages BFOA’s population dynamics to simultaneously evaluate classifier accuracy, prediction stability, and variance reduction. These capabilities were not natively supported by the standard single-objective GA operator, which evaluates and recombines classifiers in pairs rather than as a coordinated population.
- Comprehensive empirical validation across two experimental domains: static classification on 19 machine learning datasets and dynamic cybersecurity defense in a simulated network environment with six distinct attack scenarios, demonstrating that BFOA-XCS achieves consistent ranking improvements over standard XCS, supported by medium-to-large effect sizes (Cohen’s d = 0.623–0.803) and a 15.2% variance reduction.
- Comparative evaluation against modern deep RL (DRL) baselines (Deep Q-Network (DQN), Q-Learning, Policy Gradient (REINFORCE), Proximal Policy Optimization (PPO), and Soft Actor–Critic (SAC)), demonstrating that XCS with BFOA optimization significantly outperforms three of five neural baselines (DQN, Q-Learning, and Policy Gradient) while offering interpretable decision rules, warmup-free competence, and substantially lower computational cost per training run than the top-performing neural agents (PPO at 5.3× and SAC at 26.1× the XCS compute time).
2. Related Domains
2.1. Learning Classifier Systems: The eXtended Classifier Systems Subdomain
2.2. Nature-Inspired Metaheuristic: Bacterial Foraging Optimization Algorithm
2.3. Reinforcement Learning in Cybersecurity
3. Methods and Methodology
3.1. eXtended Classifier System
3.1.1. Knowledge Representation
3.1.2. Main XCS Cycle
| Algorithm 1 Pseudocode of the XCS main reinforcement learning cycle (adapted from [58]) |
| initialize reinforcement_program rp initialize environment env initialize XCS // create population [P]; connect to evaluation function procedure RUN_XCS() 1. do { 2. σ = get sensory input from env 3. generate [M] from [P] using σ // covering may be activated 4. generate PA from [M] 5. act = select action according to PA 6. execute act in env 7. ρ = get reward from rp 8. if [A]−1 is not empty then 9. P = ρ−1 + γ·max(PA) 10. update set [A]-1 using P, possibly deleting in [P] 11. → run GA in [A]−1, possibly inserting in [P] // replaced by BFOA (Section 3.3) 12. end if 13. if rp signals end-of-problem then 14. P = ρ 15. update set [A] using P, possibly deleting in [P] 16. → run GA in [A], possibly inserting in [P] // replaced by BFOA (Section 3.3) 17. clear [A]−1 18. else 19. [A]−1 = [A] 20. ρ−1 = ρ 21. σ−1 = σ 22. end if 23. } while (termination criteria from rp are not met) end procedure |
3.2. Bacterial Foraging Optimization Algorithm
3.2.1. Standard BFOA
| Algorithm 2 Pseudocode of the standard Bacterial Foraging Optimization Algorithm (based on [18]) |
| Input: N: Number of bacteria Nc: Number of chemotactic steps Nr: Number of reproduction steps Ned: Number of elimination-dispersal steps Ns: Swim length (max swim steps per tumble) Pe: Probability of elimination LB, UB: Search space bounds F: Evaluation (fitness) function Output: Best solution found |
| 1. Initialize N bacteria at random positions within [LB, UB] 2. Evaluate health of each bacterium using F 3. for ed = 1 to Ned do // Elimination-Dispersal 4. for rs = 1 to Nr do // Reproduction 5. for cs = 1 to Nc do // Chemotaxis 6. for each bacterium i do 7. Generate random tumble direction Δ 8. Normalize: φ = Δ/||Δ|| 9. Move: xnew = xi + C·φ // C = step size 10. Evaluate F(xnew) 11. if F(xnew) < F(xi) then // Improvement found 12. xi = xnew // Accept new position 13. for s = 1 to Ns do // Swim in same direction 14. xswim = xi + C·φ 15. if F(xswim) < F(xi) then 16. xi = xswim 17. else break // Stop swimming 18. end for 19. end if 20. end for (bacterium) 21. end for (chemotaxis) 22. Sort bacteria by health (ascending) 23. Replace worst N/2 bacteria with copies of best N/2 // Reproduction 24. end for (reproduction) 25. for each bacterium i do // Elimination-Dispersal 26. if random() < Pe then 27. Relocate xi to random position within [LB, UB] 28. end if 29. end for 30. end for (elimination-dispersal) 31. return best solution found |
3.2.2. Adaptive BFOA
- Step size C: If the algorithm detects improvement (aggregate improvement > 0.01), the step size is increased to expand the search. When the population stagnates, it decays gradually. If stagnation persists beyond a threshold (five consecutive cycles without progress), the step size is reset to a balanced midpoint value, acting as a recovery mechanism.
- Swim length Ns: Increases alongside step size growth during productive phases and decreases during stagnation, bounded between minimum and maximum values.
- Elimination probability Pe: Increases during stagnation (promoting exploration through random relocation) and decreases during productive phases (preserving promising solutions).
3.2.3. Improved BFOA (IBFOA)
3.3. BFOA-XCS Integration
3.3.1. Replacing GA with BFOA in XCS
| Algorithm 3 Pseudocode of the BFOA integration into the XCS cycle (replaces GA at Algorithm 1, lines 11 and 16) |
| procedure RUN_BFOA(actionSet [A], sensory input σ) 1. if [A] is empty then return 2. N = number of classifiers in [A] 3. D = condition length + 1 // Map classifiers into a continuous-valued search space so BFOA’s // chemotactic movement can operate on the action set 4. for each classifier cli in [A] do 5. for each condition bit j do 6. if cli.condition[j] = ‘#’ then pos[i][j] = 0.5 7. else if cli.condition[j] = 1 then pos[i][j] = 0.75 8. else pos[i][j] = 0.25 9. end for 10. pos[i][D] = normalize(cli.action, targetMin, targetMax) 11. end for // Refine the action-set population through chemotaxis and // population dynamics under the multi-criteria fitness 12. Select BFOA variant (Standard or Adaptive or IBFOA) 13. Initialize optimizer with pos, bounds = [0, 1]D, fitness function F 14. optimizedpos = optimizer.optimize() // Write optimized action values back to the classifiers. Conditions // are deliberately left unchanged to preserve XCS’s binary rule // representation and avoid disrupting mature populations. 15. for each classifier cli in [A] do 16. actionraw = denormalize(optimizedpos[i][D], targetMin, targetMax) 17. if classification task then 18. cli.action = snapToNearestClass(actionraw) 19. else 20. cli.action = actionraw 21. end if 22. end for end procedure |
3.3.2. Multi-Objective Classifier Fitness Function
3.4. Cybersecurity Simulation Environment
3.4.1. Network Topology
3.4.2. State Representation
3.4.3. Action Space
- Defend Node: Removes compromise and applies defense. Rewards correctly targeting compromised (+10) or vulnerable (+3) nodes and penalizes defending nodes that are already safe (−2 for safe, −1 for already defended).
- Isolate Node: Disconnects a node from the network. Rewards isolating compromised nodes (+7), penalizes isolating clean nodes (−3).
- Patch Node: Removes all vulnerabilities from a node and applies defense. Rewards patching vulnerable nodes (+5).
- Monitor Node: Places surveillance on a node. Small reward for monitoring suspicious nodes (+1).
- Restore Node: Reconnects a previously isolated node. Rewards restoring isolated nodes (+3).
- Do Nothing: No action is taken. Receives a constant penalty (−2) to discourage passivity.
3.4.4. Attack Dynamics
3.4.5. Reward Structure
3.4.6. Scenario Profiles
4. Experimental Evaluation
4.1. Experiment 1: Benchmark Machine Learning Datasets
4.1.1. Experimental Setup
4.1.2. Results and Analysis
Overall Operator Ranking
Per-Dataset Accuracy
Effect Sizes and Pairwise Comparisons
Variance Reduction
Multi-Objective Fitness Ablation
Summary of Key Findings
4.2. Experiment 2: Dynamic Cybersecurity Defense
4.2.1. Experimental Setup
- DQN (Deep Q-Network): learning rate 0.001, discount factor γ = 0.99, ε-greedy exploration with ε = 0.1.
- Q-Learning: tabular with learning rate 0.1, γ = 0.95, ε = 0.1, and state discretization into 10 bins.
- Policy Gradient (REINFORCE): the foundational policy gradient algorithm [69], learning rate 0.001, γ = 0.95.
- PPO (Proximal Policy Optimization): learning rate 0.001, γ = 0.99, clip ε = 0.2, GAE-λ (Generalized Advantage Estimation) = 0.95, 4 epochs per update. For PPO, the clip-region check in the surrogate objective uses a non-strict comparison: the unclipped policy gradient is applied whenever surr1 ≤ surr2, ensuring that the gradient flows when the importance ratio lies inside the clip interval [1 − ε, 1 + ε], including at its boundary. This is consistent with the surrogate objective specified by Schulman et al. (2017) [70].
- SAC (Soft Actor–Critic, discrete variant): actor learning rate 0.00015, critic learning rate 0.0003, γ = 0.99, initial entropy coefficient α = 0.2 with automatic tuning, target entropy ratio 0.5, Polyak τ = 0.005, replay buffer capacity 10,000, batch size 64, warmup 128 steps.
- Random: uniformly random action selection (lower-bound baseline).
4.2.2. Results and Analysis
Overall Agent Ranking
Per-Scenario Analysis
P# Sensitivity
Classifier Population Analysis
4.2.3. Comparison with Deep Reinforcement Learning Baselines
Statistical Significance
Architectural Analysis
Computational Cost
5. Discussion
5.1. Interpretation of Results
5.2. Limitations
6. Future Work
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| BFOA | Bacterial Foraging Optimization Algorithm |
| CS-1 | Cognitive System One |
| DDoS | Distributed Denial of Service |
| DRL | Deep Reinforcement Learning |
| DQN | Deep Q-Network |
| GA | Genetic Algorithm |
| GAE | Generalized Advantage Estimation |
| IBFOA | Improved Bacterial Foraging Optimization Algorithm |
| LCS | Learning Classifier System(s) |
| NSL-KDD | Network Security Laboratory–Knowledge Discovery and Data Mining (intrusion detection benchmark dataset) |
| PPO | Proximal Policy Optimization |
| RL | Reinforcement Learning |
| SAC | Soft Actor–Critic |
| SBA | Swarm-based Algorithm(s) |
| UCI | University of California, Irvine (Machine Learning Repository) |
| UNSW-NB15 | University of New South Wales Network-Based 2015 (intrusion detection benchmark dataset) |
| XCS | eXtended Classifier System |
| ZCS | Zeroth-level Classifier System |
References
- Black, J. The Power of Knowledge: How Information and Technology Made the Modern World; Yale University Press: New Haven, CT, USA, 2014. [Google Scholar]
- Sarker, I.H. AI-Based Modeling: Techniques, Applications and Research Issues towards Automation, Intelligent and Smart Systems. SN Comput. Sci. 2022, 3, 158. [Google Scholar] [CrossRef] [PubMed]
- Siddique, N.; Adeli, H. Nature Inspired Computing: An Overview and Some Future Directions. Cogn. Comput. 2015, 7, 706–714. [Google Scholar] [CrossRef] [PubMed]
- Holland, J.H. Adaptation. In Progress in Theoretical Biology; Academic Press: New York, NY, USA, 1976; Volume 4, pp. 263–293. [Google Scholar]
- Siddique, A.; Heider, M.; Iqbal, M.; Shiraishi, H. A Survey on Learning Classifier Systems from 2022 to 2024. In Proceedings of the Genetic and Evolutionary Computation Conference Companion (GECCO ’24 Companion), Melbourne, VIC, Australia, 14–18 July 2024; Association for Computing Machinery: New York, NY, USA; pp. 1797–1806. [CrossRef]
- Holland, J.H.; Reitman, J.S. Cognitive Systems Based on Adaptive Algorithms. In Pattern-Directed Inference Systems; Waterman, D.A., Hayes-Roth, F., Eds.; Academic Press: New York, NY, USA, 1978; pp. 313–329. [Google Scholar]
- Wilson, S.W. Classifier Fitness Based on Accuracy. Evol. Comput. 1995, 3, 149–175. [Google Scholar] [CrossRef]
- Pätzel, D.; Stein, A.; Hähner, J. A Survey of Formal Theoretical Advances Regarding XCS. In Proceedings of the Genetic and Evolutionary Computation Conference Companion (GECCO ’19 Companion), Prague, Czech Republic, 13–17 July 2019; Association for Computing Machinery: New York, NY, USA; pp. 1295–1302. [CrossRef]
- Gowri, A.S.; ShanthiBala, P.; Ramdinthara, I.Z. Fog-Cloud Enabled Internet of Things Using Extended Classifier System (XCS). In Artificial Intelligence-Based Internet of Things Systems; Pal, S., De, D., Buyya, R., Eds.; Springer: Cham, Switzerland, 2022; pp. 163–189. [Google Scholar] [CrossRef]
- Roozegar, M.; Mahjoob, M.J.; Esfandyari, M.J.; Panahi, M.S. XCS-Based Reinforcement Learning Algorithm for Motion Planning of a Spherical Mobile Robot. Appl. Intell. 2016, 45, 736–746. [Google Scholar] [CrossRef]
- Storn, R.; Price, K. Differential Evolution—A Simple and Efficient Heuristic for Global Optimization over Continuous Spaces. J. Glob. Optim. 1997, 11, 341–359. [Google Scholar] [CrossRef]
- Dorigo, M.; Maniezzo, V.; Colorni, A. Ant System: Optimization by a Colony of Cooperating Agents. IEEE Trans. Syst. Man Cybern. B Cybern. 1996, 26, 29–41. [Google Scholar] [CrossRef]
- Hatamlou, A. Black Hole: A New Heuristic Optimization Approach for Data Clustering. Inf. Sci. 2013, 222, 175–184. [Google Scholar] [CrossRef]
- Rao, R.V.; Savsani, V.J.; Vakharia, D.P. Teaching–Learning-Based Optimization: A Novel Method for Constrained Mechanical Design Optimization Problems. Comput.-Aided Des. 2011, 43, 303–315. [Google Scholar] [CrossRef]
- Wang, Y.; Zhang, J.; Zhang, M.; Wang, D.; Yang, M. Enhanced Artificial Ecosystem-Based Optimization for Global Optimization and Constrained Engineering Problems. Clust. Comput. 2024, 27, 10053–10092. [Google Scholar] [CrossRef]
- Ezugwu, A.E.; Shukla, A.K.; Nath, R.; Akinyelu, A.A.; Agushaka, J.O.; Chiroma, H.; Muhuri, P.K. Metaheuristics: A comprehensive overview and classification along with bibliometric analysis. Artif. Intell. Rev. 2021, 54, 4237–4316. [Google Scholar] [CrossRef]
- Wolpert, D.H.; Macready, W.G. No Free Lunch Theorems for Optimization. IEEE Trans. Evol. Comput. 1997, 1, 67–82. [Google Scholar] [CrossRef]
- Passino, K.M. Biomimicry of Bacterial Foraging for Distributed Optimization and Control. IEEE Control Syst. Mag. 2002, 22, 52–67. [Google Scholar] [CrossRef]
- Das, S.; Biswas, A.; Dasgupta, S.; Abraham, A. Bacterial Foraging Optimization Algorithm: Theoretical Foundations, Analysis, and Applications. In Foundations of Computational Intelligence Volume 3: Global Optimization; Springer: Berlin/Heidelberg, Germany, 2009; pp. 23–55. [Google Scholar] [CrossRef]
- Mishra, S.; Bhende, C.N. Bacterial Foraging Technique-Based Optimized Active Power Filter for Load Compensation. IEEE Trans. Power Deliv. 2007, 22, 457–465. [Google Scholar] [CrossRef]
- Xin, F.; Zhang, M.; Li, J.; Luo, C. Phase Retrieval for Radar Constant–Modulus Signal Design Based on the Bacterial Foraging Optimization Algorithm. Electronics 2024, 13, 506. [Google Scholar] [CrossRef]
- Shanthi, D.L.; Chethan, N. Genetic Algorithm Based Hyper-Parameter Tuning to Improve the Performance of Machine Learning Models. SN Comput. Sci. 2023, 4, 119. [Google Scholar] [CrossRef]
- Pandey, H.M.; Chaudhary, A.; Mehrotra, D. A Comparative Review of Approaches to Prevent Premature Convergence in GA. Appl. Soft Comput. 2014, 24, 1047–1077. [Google Scholar] [CrossRef]
- Lanzi, P.L. Learning Classifier Systems from a Reinforcement Learning Perspective. Soft Comput. 2002, 6, 162–170. [Google Scholar] [CrossRef]
- Watkins, C.J.C.H.; Dayan, P. Q-Learning. Mach. Learn. 1992, 8, 279–292. [Google Scholar] [CrossRef]
- Zang, Z.; Li, D.; Wang, J.; Xia, D. Learning classifier system with average reward reinforcement learning. Knowl.-Based Syst. 2013, 40, 58–71. [Google Scholar] [CrossRef]
- Butz, M.V. Learning Classifier Systems. In Springer Handbook of Computational Intelligence; Kacprzyk, J., Pedrycz, W., Eds.; Springer: Berlin/Heidelberg, Germany, 2015; pp. 961–981. [Google Scholar] [CrossRef]
- Bull, L.; Bernadó-Mansilla, E.; Holmes, J. (Eds.) Learning Classifier Systems in Data Mining: An Introduction. In Learning Classifier Systems in Data Mining; Springer: Berlin/Heidelberg, Germany, 2008; Volume 125, pp. 1–15. [Google Scholar] [CrossRef]
- Djartov, B.; Mostaghim, S.; Papenfuß, A.; Wies, M. A Learning Classifier System Approach to Time-Critical Decision-Making in Dynamic Alternate Airport Selection. In Proceedings of the 2024 IEEE Congress on Evolutionary Computation (CEC), Yokohama, Japan, 30 June–5 July 2024; IEEE: New York, NY, USA, 2024; pp. 1–8. [Google Scholar] [CrossRef]
- Irfan, M.; Zheng, J.; Iqbal, M.; Masood, Z.; Arif, M.H. Knowledge extraction and retention based continual learning by using convolutional autoencoder-based learning classifier system. Inf. Sci. 2022, 591, 287–305. [Google Scholar] [CrossRef]
- Ferjani, R.; Rejeb, L.; Abdelkarim, C.; Ben Said, L. Evidential Supervised Classifier System: A New Learning Classifier System Dealing with Imperfect Information. Int. J. Inf. Technol. Decis. Mak. 2024, 23, 917–938. [Google Scholar] [CrossRef]
- Bull, L. A brief history of learning classifier systems: From CS-1 to XCS and its variants. Evol. Intell. 2015, 8, 55–70. [Google Scholar] [CrossRef]
- Yang, J.; Xu, H.; Jia, P. Effective search for Pittsburgh learning classifier systems via estimation of distribution algorithms. Inf. Sci. 2012, 198, 100–117. [Google Scholar] [CrossRef]
- Urbanowicz, R.J.; Granizo-Mackenzie, A.; Moore, J.H. An Analysis Pipeline with Statistical and Visualization-Guided Knowledge Discovery for Michigan-Style Learning Classifier Systems. IEEE Comput. Intell. Mag. 2012, 7, 35–45. [Google Scholar] [CrossRef]
- Urbanowicz, R.J.; Moore, J.H. Learning Classifier Systems: A Complete Introduction, Review, and Roadmap. J. Artif. Evol. Appl. 2009, 2009, 736398. [Google Scholar] [CrossRef]
- Wilson, S.W. ZCS: A zeroth level classifier system. Evol. Comput. 1994, 2, 1–18. [Google Scholar] [CrossRef]
- Stalph, P.O.; Butz, M.V. Current XCSF Capabilities and Challenges. In Learning Classifier Systems; Bacardit, J., Browne, W.N., Drugowitsch, J., Bernadó-Mansilla, E., Butz, M.V., Eds.; Springer: Berlin/Heidelberg, Germany, 2010; Volume 6471, pp. 57–69. [Google Scholar] [CrossRef]
- Smierzchała, Ł.; Kozłowski, N.; Unold, O. Anticipatory Classifier System With Episode-Based Experience Replay. IEEE Access 2023, 11, 41190–41204. [Google Scholar] [CrossRef]
- Butz, M.V.; Lanzi, P.L.; Wilson, S.W. Function approximation with XCS: Hyperellipsoidal conditions, recursive least squares, and compaction. IEEE Trans. Evol. Comput. 2008, 12, 355–376. [Google Scholar] [CrossRef]
- Yousefi, A.; Badie, K.; Ebadzadeh, M.M.; Sharifi, A. Improving the efficiency of the XCS learning classifier system using evolutionary memory. Wirel. Netw. 2024, 30, 5171–5186. [Google Scholar] [CrossRef]
- Butz, M.V.; Goldberg, D.E.; Tharakunnel, K. Analysis and improvement of fitness exploitation in XCS: Bounding models, tournament selection, and bilateral accuracy. Evol. Comput. 2003, 11, 239–277. [Google Scholar] [CrossRef] [PubMed]
- Heider, M.; Pätzel, D.; Stegherr, H.; Hähner, J. A Metaheuristic Perspective on Learning Classifier Systems. In Metaheuristics for Machine Learning: New Advances and Tools; Eddaly, M., Jarboui, B., Siarry, P., Eds.; Springer Nature: Singapore, 2023; pp. 73–98. [Google Scholar] [CrossRef]
- Shehab, M.; Sihwail, R.; Daoud, M.; Al-Mimi, H.; Abualigah, L. Nature-Inspired Metaheuristic Algorithms: A Comprehensive Review. Int. Arab J. Inf. Technol. 2024, 21, 815–831. [Google Scholar] [CrossRef]
- Darvishpoor, S.; Darvishpour, A.; Escarcega, M.; Hassanalian, M. Nature-Inspired Algorithms from Oceans to Space: A Comprehensive Review of Heuristic and Meta-Heuristic Optimization Algorithms and Their Potential Applications in Drones. Drones 2023, 7, 427. [Google Scholar] [CrossRef]
- Martinson, J.N.V.; Walk, S.T. Escherichia coli Residency in the Gut of Healthy Human Adults. EcoSal Plus 2020, 9, ESP-0003-2020. [Google Scholar] [CrossRef]
- Chen, H.; Zhu, Y.; Hu, K. Adaptive bacterial foraging optimization. Abstr. Appl. Anal. 2011, 2011, 108269. [Google Scholar] [CrossRef]
- Ahmad, F.; Zhu, D.; Sun, J. Bacterial chemotaxis: A way forward to aromatic compounds biodegradation. Environ. Sci. Eur. 2020, 32, 52. [Google Scholar] [CrossRef]
- Li, J.; Dang, J.; Bu, F.; Wang, J. Analysis and improvement of the bacterial foraging optimization algorithm. J. Comput. Sci. Eng. 2014, 8, 1–10. [Google Scholar] [CrossRef][Green Version]
- Abouhawwash, M. Innovations in Cyber Defense with Deep Reinforcement Learning: A Concise and Contemporary Review. Artif. Intell. Cybersecur. 2024, 1, 44–51. [Google Scholar] [CrossRef]
- Oh, S.H.; Jeong, M.K.; Kim, H.C.; Park, J. Applying Reinforcement Learning for Enhanced Cybersecurity against Adversarial Simulation. Sensors 2023, 23, 3000. [Google Scholar] [CrossRef]
- Naeem, M.R.; Amin, R.; Farhan, M.; Alsubaei, F.S.; Alsolami, E.; Zakaria, M.D. Cyber Security Enhancements with Reinforcement Learning: A Zero-Day Vulnerability Identification Perspective. PLoS ONE 2025, 20, e0324595. [Google Scholar] [CrossRef]
- Gueriani, A.; Kheddar, H.; Mazari, A.C. Deep Reinforcement Learning for Intrusion Detection in IoT: A Survey. In Proceedings of the 2023 2nd International Conference on Electronics, Energy and Measurement (IC2EM), Medea, Algeria, 28–29 November 2023; IEEE: Piscataway, NJ, USA, 2023; Volume 1, pp. 1–7. [Google Scholar] [CrossRef]
- Hore, S.; Shah, A.; Bastian, N.D. Deep VULMAN: A Deep Reinforcement Learning-Enabled Cyber Vulnerability Management Framework. Expert Syst. Appl. 2023, 221, 119734. [Google Scholar] [CrossRef]
- Ren, S.; Jin, J.; Niu, G.; Liu, Y. ARCS: Adaptive Reinforcement Learning Framework for Automated Cybersecurity Incident Response Strategy Optimization. Appl. Sci. 2025, 15, 951. [Google Scholar] [CrossRef]
- Hammad, A.A.; Jasim, F.T. Adaptive Cyber Defense using Advanced Deep Reinforcement Learning Algorithms: A Real-Time Comparative Analysis. J. Comput. Theor. Appl. 2025, 2, 523–535. [Google Scholar] [CrossRef]
- Alnfiai, M.M. AI-powered cyber resilience: A reinforcement learning approach for automated threat hunting in 5G networks. EURASIP J. Wirel. Commun. Netw. 2025, 2025, 68. [Google Scholar] [CrossRef]
- Bu, S.-J.; Kang, H.-B.; Cho, S.-B. Ensemble of Deep Convolutional Learning Classifier System Based on Genetic Algorithm for Database Intrusion Detection. Electronics 2022, 11, 745. [Google Scholar] [CrossRef]
- Butz, M.V.; Wilson, S.W. An Algorithmic Description of XCS. In Advances in Learning Classifier Systems; Lanzi, P.L., Stolzmann, W., Wilson, S.W., Eds.; Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2001; Volume 1996, pp. 253–272. [Google Scholar] [CrossRef]
- Yan, X.; Zhu, Y.; Zhang, H.; Chen, H.; Niu, B. An Adaptive Bacterial Foraging Optimization Algorithm with Lifecycle and Social Learning. Discret. Dyn. Nat. Soc. 2012, 2012, 409478. [Google Scholar] [CrossRef]
- Chen, M.; Ou, Y.; Qiu, X.; Wang, H. An Effective Bacterial Foraging Optimization Based on Conjugation and Novel Step-Size Strategies. In Artificial Intelligence and Security; Sun, X., Wang, J., Bertino, E., Eds.; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2020; Volume 12239, pp. 362–374. [Google Scholar] [CrossRef]
- Mantegna, R.N. Fast, accurate algorithm for numerical simulation of Lévy stable stochastic processes. Phys. Rev. E 1994, 49, 4677–4683. [Google Scholar] [CrossRef]
- Strom, B.E.; Applebaum, A.; Miller, D.P.; Nickels, K.C.; Pennington, A.G.; Thomas, C.B. MITRE ATT&CK: Design and Philosophy; Technical Report MP180360R1; revised March 2020; The MITRE Corporation: McLean, VA, USA, 2018; Available online: https://attack.mitre.org/docs/ATTACK_Design_and_Philosophy_March_2020.pdf (accessed on 21 May 2026).
- Booth, H.; Rike, D.; Witte, G.A. The National Vulnerability Database (NVD): Overview; ITL Bulletin; National Institute of Standards and Technology: Gaithersburg, MD, USA, 2013. Available online: https://csrc.nist.gov/CSRC/media/Publications/Shared/documents/itl-bulletin/itlbul2013-12.pdf (accessed on 21 May 2026).
- Frei, S.; May, M.; Fiedler, U.; Plattner, B. Large-Scale Vulnerability Analysis. In Proceedings of the 2006 SIGCOMM Workshop on Large-Scale Attack Defense (LSAD ‘06), Pisa, Italy, 11 September 2006; ACM: New York, NY, USA, 2006; pp. 131–138. [Google Scholar] [CrossRef]
- Shahzad, M.; Shafiq, M.Z.; Liu, A.X. Large Scale Characterization of Software Vulnerability Life Cycles. IEEE Trans. Dependable Secur. Comput. 2020, 17, 730–744. [Google Scholar] [CrossRef]
- Jajodia, S.; Ghosh, A.K.; Swarup, V.; Wang, C.; Wang, X.S. (Eds.) Moving Target Defense: Creating Asymmetric Uncertainty for Cyber Threats; Advances in Information Security; Springer: New York, NY, USA, 2011; Volume 54. [Google Scholar] [CrossRef]
- Demšar, J. Statistical Comparisons of Classifiers over Multiple Data Sets. J. Mach. Learn. Res. 2006, 7, 1–30. [Google Scholar]
- Cohen, J. Statistical Power Analysis for the Behavioral Sciences, 2nd ed.; Lawrence Erlbaum Associates: Hillsdale, NJ, USA, 1988. [Google Scholar]
- Williams, R.J. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Mach. Learn. 1992, 8, 229–256. [Google Scholar] [CrossRef]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal Policy Optimization Algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar] [CrossRef]
- Kim, D.H.; Abraham, A.; Cho, J.H. A Hybrid Genetic Algorithm and Bacterial Foraging Approach for Global Optimization. Inf. Sci. 2007, 177, 3918–3937. [Google Scholar] [CrossRef]
- Asgarkhani, N.; Kazemi, F.; Jankowski, R.; Formisano, A. Dynamic ensemble-learning model for seismic risk assessment of masonry infilled steel structures incorporating soil-foundation-structure interaction. Reliab. Eng. Syst. Saf. 2026, 267, 111839. [Google Scholar] [CrossRef]






| Feature | Standard BFOA | Adaptive BFOA | IBFOA |
|---|---|---|---|
| Step size | Fixed | Self-tuning (improvement-based) | Lévy flight (heavy-tailed) |
| Swim length | Fixed | Adaptive (increases/decreases) | Fixed |
| Elimination probability | Fixed | Adaptive (based on stagnation) | Fixed |
| Reproduction | Sort and replace | Sort and replace (inherited) | Omitted |
| Elitism | None | None | Global best preserved |
| Stagnation recovery | None | Parameter reset after threshold | Lévy jumps escape local optima |
| Swarming | Optional (based on gradient) | Optional (inherited) | Not used |
| Mechanism | Effect on Network State | Per-Step Dynamics | Immunity/Mitigation | Reference |
|---|---|---|---|---|
| Lateral spread | Compromised node attempts to compromise an adjacent neighbor | βvul if neighbor has known vulnerabilities; βclean otherwise (βvul > βclean) | Isolated and defended neighbors are immune | MITRE ATT&CK Lateral Movement (TA0008) [62] |
| External attacks | An external attacker compromises an undefended, vulnerable node directly | αext on each undefended, vulnerable node, independent of internal spread | Defended nodes are immune | MITRE ATT&CK Initial Access (TA0001) [62] |
| Vulnerability emergence | A new vulnerability appears on a node, simulating the discovery of a new exploit during the episode | γvul on each undefended, uncompromised node | Defended and compromised nodes do not acquire new vulnerabilities in this step | CVE/NVD vulnerability disclosure [63]; vulnerability-lifecycle literature [64,65] |
| Defense decay | An active defense lapses, returning the node to an undefended state | δdecay on each defended node | None: defense decay applies to all defended nodes | Static-defense-obsolescence framing in moving target defense literature [66] |
| Parameter | DDoS | Intrusion | Ransomware | Port Scan | Data Exfil. | Mixed |
|---|---|---|---|---|---|---|
| Spread rate (vulnerable), βvul | 0.1 | 0.06 | 0.15 | 0.04 | 0.12 | 0.08 |
| Spread rate (clean), βclean | 0.03 | 0.02 | 0.04 | 0.01 | 0.03 | 0.02 |
| External attack rate, αext | 0.04 | 0.01 | 0.01 | 0.06 | 0.02 | 0.03 |
| Defense attack rate, δdecay | 0.02 | 0.05 | 0.03 | 0.04 | 0.03 | 0.03 |
| New vulnerability rate, γvul | 0.04 | 0.08 | 0.02 | 0.12 | 0.04 | 0.06 |
| Initial compromised | 2 | 1 | 1 | 0 | 1 | 1 |
| Initial vulnerable (%) | 40 | 50 | 20 | 60 | 30 | 35 |
| Terminal threshold (%) | 60 | 60 | 55 | 60 | 60 | 60 |
| No. | Dataset | Source | Category | Instances | Target | Classes |
|---|---|---|---|---|---|---|
| 1 | Breast Cancer | UCI | Classification | 569 | diagnosis | 2 |
| 2 | German Credit | UCI | Classification | 1000 | class | 2 |
| 3 | Heart Disease | UCI | Classification | 297 | class | 2 |
| 4 | Ionosphere | UCI | Classification | 351 | class | 2 |
| 5 | Iris | UCI | Classification | 150 | class | 3 |
| 6 | Neurofibromatosis | Other | Classification | 296 | class | 2 |
| 7 | Wine | UCI | Classification | 178 | class | 3 |
| 8 | Auto MPG | UCI | Regression | 392 | mpg | 41 bins |
| 9 | Boston Housing | UCI | Regression | 506 | MEDV | 41 bins |
| 10 | Concrete Strength | UCI | Regression | 1030 | Strength | 41 bins |
| 11 | Paddy Yield | Other | Regression | 2789 | yield | 41 bins |
| 12 | Wine Quality | UCI | Regression | 1143 | quality | 6 1 |
| 13 | NSL-KDD (Binary) 2 | Cyber | Cybersecurity | 25,192 | label | 2 |
| 14 | NSL-KDD (Multi) 2 | Cyber | Cybersecurity | 25,194 | attack_cat | 5 |
| 15 | UNSW-NB15 (Binary) 2 | Cyber | Cybersecurity | 16,466 | label | 2 |
| 16 | UNSW-NB15 (Multi) 2 | Cyber | Cybersecurity | 16,466 | attack_cat | 10 |
| 17 | KDD Test+ | Cyber | Cybersecurity | 22,544 | label | 2 |
| 18 | Ransomware (RG) | Cyber | Cybersecurity | 2675 | RG | 2 |
| 19 | Ransomware (Family) | Cyber | Cybersecurity | 2675 | family | 41 bins 1 |
| Rank | Operator | Avg. Rank |
|---|---|---|
| 1 | IBFOA | 2.395 |
| 2 | NONE | 2.553 |
| 3 | BFOA | 3.000 |
| 4 | GA | 3.368 |
| 5 | Adaptive BFOA | 3.684 |
| Dataset | NONE | GA | BFOA | Adaptive BFOA | IBFOA |
|---|---|---|---|---|---|
| Breast Cancer | 0.762 | 0.758 | 0.720 | 0.717 | 0.772 |
| German Credit | 0.688 | 0.704 | 0.687 | 0.687 | 0.697 |
| Heart Disease | 0.766 | 0.764 | 0.729 | 0.741 | 0.751 |
| Ionosphere | 0.759 | 0.777 | 0.741 | 0.754 | 0.739 |
| Iris | 0.937 | 0.920 | 0.883 | 0.837 | 0.883 |
| Neurofibromatosis | 0.559 | 0.544 | 0.566 | 0.548 | 0.556 |
| Wine | 0.800 | 0.718 | 0.777 | 0.747 | 0.741 |
| Auto MPG | 0.381 | 0.379 | 0.381 | 0.381 | 0.381 |
| Boston Housing | 0.518 | 0.517 | 0.518 | 0.519 | 0.518 |
| Concrete Strength | 0.563 | 0.563 | 0.563 | 0.563 | 0.563 |
| Paddy Yield | 0.328 | 0.328 | 0.328 | 0.328 | 0.328 |
| Wine Quality | 0.512 | 0.433 | 0.524 | 0.537 | 0.539 |
| NSL-KDD (Binary) | 0.534 | 0.544 | 0.543 | 0.532 | 0.534 |
| NSL-KDD (Multi) | 0.365 | 0.368 | 0.394 | 0.372 | 0.365 |
| UNSW-NB15 (Binary) | 0.449 | 0.452 | 0.462 | 0.462 | 0.449 |
| UNSW-NB15 (Multi) | 0.288 | 0.249 | 0.211 | 0.198 | 0.211 |
| KDD Test+ | 0.681 | 0.569 | 0.655 | 0.618 | 0.654 |
| Ransomware (RG) | 0.885 | 0.785 | 0.863 | 0.836 | 0.875 |
| Ransomware (Family) | 0.772 | 0.772 | 0.771 | 0.771 | 0.771 |
| Comparison | Wilcoxon p | Cohen’s d | Magnitude | Direction |
|---|---|---|---|---|
| GA vs. NONE | 0.220 | 0.803 | Large | NONE > GA |
| GA vs. BFOA | 0.457 | 0.693 | Medium | BFOA > GA |
| GA vs. Adaptive BFOA | 1.000 | 0.623 | Medium | Adaptive BFOA > GA |
| GA vs. IBFOA | 0.433 | 0.737 | Medium | IBFOA > GA |
| Operator | Variance Reduction | Interpretation |
|---|---|---|
| NONE | 4.3% | Slight reduction |
| BFOA | 4.7% | Slight reduction |
| Adaptive BFOA | 1.8% | Slight reduction |
| IBFOA | 15.2% | Notable reduction |
| Metric | MO = True | MO = False |
|---|---|---|
| Friedman χ2 | 8.895 | 6.200 |
| Friedman p-value | 0.064 | 0.185 |
| Best Friedman rank | IBFOA (2.395) | NONE (2.447) |
| GA Friedman rank | 4th (3.368) | 4th (3.158) |
| Operator | Variance reduction (MO = true) | Variance reduction (MO = false) |
| NONE | 4.3% ↓ | −1.7% (increase) |
| BFOA | 4.7% ↓ | −5.2% (increase) |
| Adaptive BFOA | 1.8% ↓ | −9.3% (increase) |
| IBFOA | 15.2% ↓ | 6.7% ↓ |
| Rank | Agent | Mean Reward | Std | Friedman Rank | Avg Time |
|---|---|---|---|---|---|
| 1 | PPO | +307.57 | 283.53 | 1.167 | 2 m 8 s |
| 2 | SAC | +110.65 | 140.39 | 1.833 | 10 m 26 s |
| 3 | Adaptive BFOA-XCS_P0.50 | +6.60 | 43.60 | 4.833 | 24 s |
| 4 | GA-XCS_P0.50 | +6.42 | 44.20 | 6.500 | 24 s |
| 5 | IBFOA-XCS_P0.50 | +6.41 | 44.66 | 6.500 | 24 s |
| 6 | NONE-XCS_P0.50 | +6.05 | 44.26 | 7.000 | 24 s |
| 7 | BFOA-XCS_P0.50 | +5.82 | 43.45 | 7.000 | 24 s |
| 8 | IBFOA-XCS_P0.60 | +3.78 | 41.06 | 8.167 | 16 s |
| 9 | GA-XCS_P0.60 | +3.70 | 40.74 | 7.500 | 16 s |
| 10 | NONE-XCS_P0.60 | +3.57 | 40.11 | 7.667 | 16 s |
| 11 | BFOA-XCS_P0.60 | +3.09 | 40.50 | 9.333 | 16 s |
| 12 | Adaptive BFOA-XCS_P0.60 | +2.52 | 40.12 | 10.500 | 16 s |
| 23 | DQN | −28.34 | 43.87 | 23.000 | 38 s |
| 24 | Q-Learning | −47.98 | 21.77 | 24.000 | <1 s |
| 25 | Random | −80.12 | 20.29 | 25.167 | <1 s |
| 26 | Policy Gradient | −86.12 | 11.22 | 25.833 | 4 s |
| Scenario | SAC | Adaptive BFOA-XCS | BFOA-XCS | NONE-XCS | DQN | PPO | Random |
|---|---|---|---|---|---|---|---|
| DDoS | +25.41 | −26.64 | −27.30 | −28.50 | −41.54 | +87.45 | −68.75 |
| Intrusion | +197.88 | +80.20 | +80.85 | +81.85 | −0.45 | +528.36 | −54.20 |
| Ransomware | +265.83 | +37.09 | +32.86 | +34.97 | +1.37 | +563.30 | −119.80 |
| Port Scan | −6.11 | −53.04 | −53.40 | −54.17 | −60.96 | −29.73 | −83.51 |
| Data Exfil. | +106.88 | +12.12 | +11.30 | +11.46 | −24.99 | +284.72 | −82.70 |
| Mixed | +74.03 | −10.13 | −9.42 | −9.30 | −43.50 | +411.33 | −71.77 |
| Operator | Scenario | Final Population Size | Mean Condition Specificity | Mean Match-Set Size |
|---|---|---|---|---|
| NONE | DDoS | 500 | 0.4994 ± 0.0010 | 38.06 ± 1.41 |
| NONE | Intrusion | 500 | 0.5003 ± 0.0014 | 34.33 ± 0.89 |
| NONE | Ransomware | 500 | 0.5005 ± 0.0013 | 30.29 ± 1.29 |
| NONE | Port Scan | 500 | 0.5001 ± 0.0007 | 40.35 ± 1.37 |
| NONE | Data Exfiltration | 500 | 0.5000 ± 0.0015 | 34.59 ± 1.21 |
| NONE | Mixed | 500 | 0.4996 ± 0.0009 | 36.25 ± 0.82 |
| GA | DDoS | 500 | 0.4999 ± 0.0010 | 37.76 ± 1.92 |
| GA | Intrusion | 500 | 0.4998 ± 0.0009 | 33.96 ± 0.69 |
| GA | Ransomware | 500 | 0.5000 ± 0.0010 | 29.67 ± 0.69 |
| GA | Port Scan | 500 | 0.5001 ± 0.0014 | 40.70 ± 1.03 |
| GA | Data Exfiltration | 500 | 0.5001 ± 0.0009 | 34.25 ± 0.73 |
| GA | Mixed | 500 | 0.5001 ± 0.0011 | 36.43 ± 1.20 |
| BFOA | DDoS | 500 | 0.4997 ± 0.0011 | 38.61 ± 1.45 |
| BFOA | Intrusion | 500 | 0.5000 ± 0.0010 | 33.64 ± 0.92 |
| BFOA | Ransomware | 500 | 0.5003 ± 0.0009 | 29.66 ± 1.27 |
| BFOA | Port Scan | 500 | 0.5004 ± 0.0009 | 39.89 ± 1.33 |
| BFOA | Data Exfiltration | 500 | 0.5002 ± 0.0011 | 34.55 ± 1.04 |
| BFOA | Mixed | 500 | 0.4993 ± 0.0011 | 36.68 ± 1.01 |
| Adaptive BFOA | DDoS | 500 | 0.5003 ± 0.0008 | 37.70 ± 1.58 |
| Adaptive BFOA | Intrusion | 500 | 0.4994 ± 0.0008 | 33.74 ± 1.00 |
| Adaptive BFOA | Ransomware | 500 | 0.4996 ± 0.0012 | 29.54 ± 1.20 |
| Adaptive BFOA | Port Scan | 500 | 0.5000 ± 0.0013 | 39.91 ± 1.52 |
| Adaptive BFOA | Data Exfiltration | 500 | 0.4997 ± 0.0011 | 34.03 ± 0.87 |
| Adaptive BFOA | Mixed | 500 | 0.5010 ± 0.0012 | 36.25 ± 0.66 |
| IBFOA | DDoS | 500 | 0.5002 ± 0.0008 | 37.53 ± 1.78 |
| IBFOA | Intrusion | 500 | 0.4996 ± 0.0008 | 33.68 ± 0.85 |
| IBFOA | Ransomware | 500 | 0.4998 ± 0.0010 | 30.16 ± 0.63 |
| IBFOA | Port Scan | 500 | 0.4997 ± 0.0009 | 40.16 ± 1.00 |
| IBFOA | Data Exfiltration | 500 | 0.4999 ± 0.0010 | 34.44 ± 1.01 |
| IBFOA | Mixed | 500 | 0.5001 ± 0.0011 | 36.64 ± 0.93 |
| Comparison | Wilcoxon p | Cohen’s d | Magnitude | Winner |
|---|---|---|---|---|
| SAC vs. BFOA-XCS | <0.001 | 1.000 | Large | SAC |
| DQN vs. BFOA-XCS | <0.001 | 0.776 | Medium | BFOA-XCS |
| Q-Learning vs. BFOA-XCS | <0.001 | 1.552 | Large | BFOA-XCS |
| PPO vs. BFOA-XCS | <0.001 | 1.179 | Large | PPO |
| Policy Gradient vs. BFOA-XCS | <0.001 | 2.873 | Large | BFOA-XCS |
| Random vs. BFOA-XCS | <0.001 | 2.513 | Large | BFOA-XCS |
| NONE-XCS vs. BFOA-XCS | 0.868 | 0.005 | Negligible | — |
| NONE-XCS vs. Adaptive BFOA-XCS | 0.284 | 0.012 | Negligible | — |
| Agent | Avg Time/Run | Overall Mean (Cumulative Reward) | Cost Ratio vs. XCS |
|---|---|---|---|
| SAC | 10 m 26 s | +110.65 | 26.1× |
| PPO | 2 m 8 s | +307.57 | 5.3× |
| DQN | 38 s | −28.34 | 1.6× |
| XCS variants (P# = 0.50 avg) | 24 s | +6.26 | 1.0× (baseline) |
| Policy Gradient | 4 s | −86.12 | 0.2× |
| Q-Learning | <1 s | −47.98 | — |
| Random | <1 s | −80.12 | — |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Novak, D.; Fister, I., Jr.; Dugonik, J. Replacing the Genetic Algorithm with Multi-Objective Bacterial Foraging Optimization in XCS. Mathematics 2026, 14, 1947. https://doi.org/10.3390/math14111947
Novak D, Fister I Jr., Dugonik J. Replacing the Genetic Algorithm with Multi-Objective Bacterial Foraging Optimization in XCS. Mathematics. 2026; 14(11):1947. https://doi.org/10.3390/math14111947
Chicago/Turabian StyleNovak, Damijan, Iztok Fister, Jr., and Jani Dugonik. 2026. "Replacing the Genetic Algorithm with Multi-Objective Bacterial Foraging Optimization in XCS" Mathematics 14, no. 11: 1947. https://doi.org/10.3390/math14111947
APA StyleNovak, D., Fister, I., Jr., & Dugonik, J. (2026). Replacing the Genetic Algorithm with Multi-Objective Bacterial Foraging Optimization in XCS. Mathematics, 14(11), 1947. https://doi.org/10.3390/math14111947

