Generalized Zero-Shot Learning for Evolving Network Device Identification
Abstract
1. Introduction
- (1)
- A lightweight Transformer architecture, named NetFormer, is proposed for traffic feature extraction. By removing complex components in conventional Transformers and introducing a stream mapping mechanism, the model achieves efficient discriminative feature learning while significantly reducing computational cost, and effectively captures temporal dependencies and statistical patterns in traffic data.
- (2)
- A Weighted Conditional Variational Autoencoder (W-CVAE) generation module is proposed, integrating multi-scale Maximum Mean Discrepancy (MMD) and attribute-feature contrastive learning. The multi-scale MMD constraint alleviates mode collapse, while the contrastive loss reduces the semantic gap between attributes and traffic features, enabling the generation of high-quality pseudo-samples with strong semantic fidelity for unseen classes.
- (3)
- An adaptive bias-calibrated hybrid prototype classification strategy is introduced. By dynamically adjusting the decision boundaries between seen and unseen classes, the proposed method effectively mitigates the seen-class bias in GZSL and achieves a more balanced allocation of prediction confidence.
- (4)
- Extensive experiments are conducted on real-world IoT traffic datasets. The results demonstrate that HALO significantly outperforms state-of-the-art baselines, across multiple evaluation metrics achieving a ZSL accuracy of 97.6% and a GZSL harmonic mean of 90.9%—representing a 16.7% and 8.1% absolute improvement over the state-of-the-art baseline (i.e., ZEST [10]), respectively, validating the effectiveness and superiority of the proposed framework.
2. Related Work
2.1. IoT Device Identification
2.2. Zero-Shot Learning
3. The HALO Overview
4. The Design Details
4.1. Feature Extraction with NetFormer
| Algorithm 1 NetFormer: training and feature extraction |
|
4.2. Conditional Generation Module with W-CVAE
| Algorithm 2 W-CVAE: training and generation |
|
4.2.1. Architecture of Attribute-Conditioned VAE
4.2.2. Composite Objective Function
- Reconstruction Loss. We use mean squared error (MSE) to measure the difference between the generated feature and the real feature x, ensuring that the generated samples preserve the distributional characteristics of the original data:
- Maximum Mean Discrepancy (MMD) Loss. To overcome the posterior collapse problem commonly associated with traditional KL divergence, we introduce MMD with multi-scale Gaussian kernels to constrain the distribution of the latent variable z to approximate a standard normal distribution. This ensures that during the testing phase, sampling z from the standard normal distribution and conditioning on unseen class attributes can decode valid pseudo-samples:
- Attribute-Feature Contrastive Loss. To bridge the heterogeneity gap between the feature space and the semantic attribute space, we adopt the InfoNCE contrastive loss. By maximizing the mutual information between the projected feature vector and its corresponding projected attribute vector, the model learns a consistent attribute-feature representation, significantly improving generalization to unseen classes:where denotes cosine similarity and is a temperature parameter.
4.2.3. Pseudo-Sample Generation
4.3. Hybrid Prototype Classification with Adaptive Bias Calibration
- Hybrid Prototype Construction
- Adaptive Bias Calibration
5. Evaluation
5.1. Experimental Setup
- 1.
- ZEST [10]: A pioneering and complete ZSL framework specifically designed for IoT device identification. It employs a Conditional Variational Autoencoder (CVAE) to map semantic attribute vectors into the feature space for pseudo-sample synthesis. We treat it as the state of the art.
- 2.
- Diff-K: In this variant, we replace the generative part of HALO with a conditional diffusion model [33] currently used in the ZSL field, which generates features via a denoising network conditioned on attribute vectors.
- 3.
- WGAN-GP: In this variant, we replace the generative part of our method with a Wasserstein GAN with gradient penalty [28] currently used in the ZSL field, where the generator produces features conditioned on noise and attributes.
- 4.
- ProNet [34]: A classical few-shot method that learns a metric space through episodic training on seen classes. Each class is represented by its prototype (the mean embedding of support samples), and classification follows the nearest prototype in Euclidean distance. During testing, for each k-shot scenario (), k unseen support samples are randomly drawn to compute unseen prototypes, while seen prototypes are obtained from all seen training embeddings.
5.2. Experimental Results
5.2.1. Comparison with Baselines
5.2.2. Comparison of Feature Extractors
5.2.3. Ablation Study
- (a)
- NetFormer only: Directly test the NetFormer trained on a closed set on an open set containing unseen classes.
- (b)
- Bi-LSTM replacement: Replace NetFormer with Bi-LSTM as the feature extractor.
- (c)
- Remove MMD and contrastive losses: Omit both losses from the generation module.
- (d)
- Remove bias calibration: Disable the bias calibration mechanism in the hybrid prototype classifier.
- (e)
- Replace hybrid prototype with MLP/SVM: Substitute the hybrid prototype classifier with a standard MLP or SVM.
5.2.4. Sensitivity Analysis
5.2.5. Visualization with t-SNE
6. Discussion
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Antonakakis, M.; April, T.; Bailey, M.D.; Bernhard, M.; Bursztein, E.; Cochran, J.; Durumeric, Z.; Halderman, J.A.; Invernizzi, L.; Kallitsis, M.; et al. Understanding the Mirai Botnet. In Proceedings of the 26th USENIX Security Symposium, USENIX Security 2017, Vancouver, BC, Canada, 16–18 August 2017; Kirda, E., Ristenpart, T., Eds.; USENIX Association: Berkeley, CA, USA, 2017; pp. 1093–1110. [Google Scholar]
- Herwig, S.; Harvey, K.; Hughey, G.; Roberts, R.; Levin, D. Measurement and Analysis of Hajime, a Peer-to-peer IoT Botnet. In Proceedings of the 26th Annual Network and Distributed System Security Symposium, NDSS 2019, San Diego, CA, USA, 24–27 February 2019; The Internet Society: Reston, VA, USA, 2019. [Google Scholar]
- Chowdhury, R.R.; Che-Idris, A.; Abas, P.E. A Deep Learning Approach for Classifying Network Connected IoT Devices Using Communication Traffic Characteristics. J. Netw. Syst. Manag. 2023, 31, 26. [Google Scholar] [CrossRef]
- Sivanathan, A.; Gharakheili, H.H.; Loi, F.; Radford, A.; Wijenayake, C.; Vishwanath, A.; Sivaraman, V. Classifying IoT Devices in Smart Environments Using Network Traffic Characteristics. IEEE Trans. Mob. Comput. 2019, 18, 1745–1759. [Google Scholar] [CrossRef]
- Martín, M.L.; Carro, B.; Sánchez-Esguevillas, A.; Lloret, J. Network Traffic Classifier with Convolutional and Recurrent Neural Networks for Internet of Things. IEEE Access 2017, 5, 18042–18050. [Google Scholar] [CrossRef]
- Luo, Y.; Chen, X.; Ge, N.; Feng, W.; Lu, J. Transformer-Based Device-Type Identification in Heterogeneous IoT Traffic. IEEE Internet Things J. 2023, 10, 5050–5062. [Google Scholar] [CrossRef]
- Hendrycks, D.; Gimpel, K. A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks. In Proceedings of the 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, 24–26 April 2017. Conference Track Proceedings. [Google Scholar]
- Ortiz, J.; Crawford, C.H.; Le, F. DeviceMien: Network device behavior modeling for identifying unknown IoT devices. In Proceedings of the International Conference on Internet of Things Design and Implementation, IoTDI 2019, Montreal, QC, Canada, 15–18 April 2019; Landsiedel, O., Nahrstedt, K., Eds.; ACM: New York, NY, USA, 2019; pp. 106–117. [Google Scholar] [CrossRef]
- Xian, Y.; Lampert, C.H.; Schiele, B.; Akata, Z. Zero-Shot Learning—A Comprehensive Evaluation of the Good, the Bad and the Ugly. IEEE Trans. Pattern Anal. Mach. Intell. 2019, 41, 2251–2265. [Google Scholar] [CrossRef] [PubMed]
- Wu, B.; Gysel, P.; Divakaran, D.M.; Gurusamy, M. ZEST: Attention-based Zero-Shot Learning for Unseen IoT Device Classification. In Proceedings of the NOMS 2024 IEEE Network Operations and Management Symposium, Seoul, Republic of Korea, 6–10 May 2024; IEEE: New York, NY, USA, 2024; pp. 1–9. [Google Scholar] [CrossRef]
- Kim, T.; Oh, D.H.; Lee, D.; Kim, S.; Choi, S.; Lee, S.; Kim, T. A Framework for ZSL-based Asset Identification using Multi-modal Data. In Proceedings of the 3rd International Conference on Foundation and Large Language Models, FLLM 2025, Vienna, Austria, 25–28 November 2025; IEEE: New York, NY, USA, 2025; pp. 600–605. [Google Scholar] [CrossRef]
- Miettinen, M.; Marchal, S.; Hafeez, I.; Asokan, N.; Sadeghi, A.; Tarkoma, S. IoT SENTINEL: Automated Device-Type Identification for Security Enforcement in IoT. In Proceedings of the 37th IEEE International Conference on Distributed Computing Systems, ICDCS 2017, Atlanta, GA, USA, 5–8 June 2017; Lee, K., Liu, L., Eds.; IEEE Computer Society: New York, NY, USA, 2017; pp. 2177–2184. [Google Scholar] [CrossRef]
- Liu, C.; He, L.; Xiong, G.; Cao, Z.; Li, Z. FS-Net: A Flow Sequence Network For Encrypted Traffic Classification. In Proceedings of the 2019 IEEE Conference on Computer Communications, INFOCOM 2019, Paris, France, 29 April–2 May 2019; IEEE: New York, NY, USA, 2019; pp. 1171–1179. [Google Scholar] [CrossRef]
- Lucas, J.; Tucker, G.; Grosse, R.B.; Norouzi, M. Understanding Posterior Collapse in Generative Latent Variable Models. In Proceedings of the Deep Generative Models for Highly Structured Data, ICLR 2019 Workshop, New Orleans, LA, USA, 6 May 2019. [Google Scholar]
- Chao, W.; Changpinyo, S.; Gong, B.; Sha, F. An Empirical Study and Analysis of Generalized Zero-Shot Learning for Object Recognition in the Wild. In Proceedings of the Computer Vision-ECCV 2016-14th European Conference, Amsterdam, The Netherlands, 11–14 October 2016; Leibe, B., Matas, J., Sebe, N., Welling, M., Eds.; Proceedings, Part II; Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2016; pp. 52–68. [Google Scholar] [CrossRef]
- Wynne, G.; Duncan, A.B. A Kernel Two-Sample Test for Functional Data. J. Mach. Learn. Res. 2022, 23, 1–51. [Google Scholar]
- van den Oord, A.; Li, Y.; Vinyals, O. Representation Learning with Contrastive Predictive Coding. arXiv 2018, arXiv:1807.03748. [Google Scholar]
- Meidan, Y.; Bohadana, M.; Shabtai, A.; Guarnizo, J.D.; Ochoa, M.; Tippenhauer, N.O.; Elovici, Y. ProfilIoT: A machine learning approach for IoT device identification based on network traffic analysis. In Proceedings of the Symposium on Applied Computing, SAC 2017, Marrakech, Morocco, 3–7 April 2017; Seffah, A., Penzenstadler, B., Alves, C., Peng, X., Eds.; ACM: New York, NY, USA, 2017; pp. 506–509. [Google Scholar] [CrossRef]
- Aneja, S.; Aneja, N.; Islam, M.S. IoT Device Fingerprint using Deep Learning. In Proceedings of the IEEE International Conference on Internet of Things and Intelligence System, IOTAIS 2018, Bali, Indonesia, 1–3 November 2018; IEEE: New York, NY, USA, 2018; pp. 174–179. [Google Scholar] [CrossRef]
- Sharma, H.; Kumar, P.; Sharma, K. Identification of Device Type Using Transformers in Heterogeneous Internet of Things Traffic; Springer: Berlin/Heidelberg, Germany, 2023; pp. 471–481. [Google Scholar] [CrossRef]
- Wang, X.; Wang, Y.; Lai, Y.; Hao, Z.; Liu, A.X. Reliable Open-Set Network Traffic Classification. IEEE Trans. Inf. Forensics Secur. 2025, 20, 2313–2328. [Google Scholar] [CrossRef]
- Yang, Q.; He, W.; Chen, M.; Du, H.; Shao, S.; Wu, F.; Liu, S.; Ji, Y.; Ren, K. End-to-End Open-Set Semi-Supervised Learning for Fine-Grained Encrypted Traffic Classification. IEEE Trans. Inf. Forensics Secur. 2026, 21, 1347–1362. [Google Scholar] [CrossRef]
- Frome, A.; Corrado, G.S.; Shlens, J.; Bengio, S.; Dean, J.; Ranzato, M.; Mikolov, T. DeViSE: A Deep Visual-Semantic Embedding Model. In Proceedings of the Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems 2013, Lake Tahoe, NV, USA, 5–8 December 2013; Burges, C.J.C., Bottou, L., Ghahramani, Z., Weinberger, K.Q., Eds.; ACM: New York, NY, USA, 2013; pp. 2121–2129. [Google Scholar]
- Akata, Z.; Reed, S.E.; Walter, D.; Lee, H.; Schiele, B. Evaluation of output embeddings for fine-grained image classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA, 7–12 June 2015; IEEE Computer Society: Washington, DC, USA, 2015; pp. 2927–2936. [Google Scholar] [CrossRef]
- Radovanovic, M.; Nanopoulos, A.; Ivanovic, M. Hubs in Space: Popular Nearest Neighbors in High-Dimensional Data. J. Mach. Learn. Res. 2010, 11, 2487–2531. [Google Scholar]
- Fu, Y.; Hospedales, T.M.; Xiang, T.; Gong, S. Transductive Multi-View Zero-Shot Learning. IEEE Trans. Pattern Anal. Mach. Intell. 2015, 37, 2332–2345. [Google Scholar] [CrossRef] [PubMed]
- Mirza, M.; Osindero, S. Conditional Generative Adversarial Nets. arXiv 2014, arXiv:1411.1784. [Google Scholar] [CrossRef]
- Gulrajani, I.; Ahmed, F.; Arjovsky, M.; Dumoulin, V.; Courville, A.C. Improved Training of Wasserstein GANs. In Proceedings of the Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, Long Beach, CA, USA, 4–9 December 2017; Guyon, I., von Luxburg, U., Bengio, S., Wallach, H.M., Fergus, R., Vishwanathan, S.V.N., Garnett, R., Eds.; ACM: New York, NY, USA, 2017; pp. 5767–5777. [Google Scholar]
- Mu, Z.; Shi, X.; Dogan, S. GMA-SAWGAN-GP: A Novel Data Generative Framework to Enhance IDS Detection Performance. arXiv 2026, arXiv:2603.28838. [Google Scholar]
- Kingma, D.P.; Welling, M. Auto-Encoding Variational Bayes. In Proceedings of the International Conference on Learning Representations (ICLR), Banff, AB, Canada, 14–16 April 2014. [Google Scholar]
- Schönfeld, E.; Ebrahimi, S.; Sinha, S.; Darrell, T.; Akata, Z. Generalized Zero- and Few-Shot Learning via Aligned Variational Autoencoders. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, 16–20 June 2019; IEEE: New York, NY, USA, 2019; pp. 8247–8255. [Google Scholar] [CrossRef]
- Chen, S.; Wang, W.; Xia, B.; Peng, Q.; You, X.; Zheng, F.; Shao, L. FREE: Feature Refinement for Generalized Zero-Shot Learning. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, 10–17 October 2021; IEEE: New York, NY, USA, 2021; pp. 122–131. [Google Scholar] [CrossRef]
- Ye, Z.; Gowda, S.N.; Chen, S.; Huang, X.; Xu, H.; Khan, F.S.; Jin, Y.; Huang, K.; Jin, X. ZeroDiff: Solidified Visual-semantic Correlation in Zero-Shot Learning. In Proceedings of the Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, 24–28 April 2025. [Google Scholar]
- Snell, J.; Swersky, K.; Zemel, R. Prototypical Networks for Few-shot Learning. In Proceedings of the Advances in Neural Information Processing Systems 30 (NeurIPS 2017), Long Beach, CA, USA, 4–9 December 2017; Curran Associates, Inc.: Red Hook, NY, USA, 2017. [Google Scholar]
- Dai, J.; Xu, X.; Gao, H.; Wang, X.; Xiao, F. SHAPE: A Simultaneous Header and Payload Encoding Model for Encrypted Traffic Classification. IEEE Trans. Netw. Serv. Manag. 2023, 20, 1993–2012. [Google Scholar] [CrossRef]
- Zhou, G.; Guo, X.; Liu, Z.; Li, T.; Li, Q.; Xu, K. TrafficFormer: An Efficient Pre-trained Model for Traffic Data. In Proceedings of the 2025 IEEE Symposium on Security and Privacy (SP), San Francisco, CA, USA, 12–15 May 2025; pp. 1844–1860. [Google Scholar] [CrossRef]
- Peng, L.; Xie, X.; Huang, S.; Wang, Z.; Cui, Y. Ptu: Pre-Trained Model for Network Traffic Understanding. In Proceedings of the 2024 IEEE 32nd International Conference on Network Protocols (ICNP), Charleroi, Belgium, 28–31 October 2024; pp. 1–12. [Google Scholar] [CrossRef]





| Method | ZSL-Acc | GZSL-S | GZSL-U | GZSL-H |
|---|---|---|---|---|
| HALO | 0.9757 | 0.9866 | 0.8430 | 0.9091 |
| ZEST | 0.8089 | 0.9972 | 0.7082 | 0.8282 |
| DIFF-K | 0.7891 | 0.9624 | 0.7887 | 0.8669 |
| WGAN-GP | 0.8151 | 0.9616 | 0.7996 | 0.8732 |
| Shots | ZSL-Acc | GZSL-S | GZSL-U | GZSL-H |
|---|---|---|---|---|
| 1-shot | 0.9756 | 0.9924 | 0.7064 | 0.8253 |
| 5-shot | 0.9767 | 0.9835 | 0.8231 | 0.8962 |
| 10-shot | 0.9828 | 0.9890 | 0.8259 | 0.9001 |
| Feature Extractor | AC (%) | F1 (%) | Inference Time (ms) | Params (M) |
|---|---|---|---|---|
| Bi-LSTM | 98.05 | 97.99 | 0.0150 | 0.16 |
| CNN-based | 96.36 | 96.26 | 0.0112 | 0.08 |
| NetFormer (ours) | 98.91 | 98.87 | 0.0130 | 0.12 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Wang, Z.; Jin, M.; Tang, Z.; Chen, D.; Wei, X.; You, L. Generalized Zero-Shot Learning for Evolving Network Device Identification. Electronics 2026, 15, 2320. https://doi.org/10.3390/electronics15112320
Wang Z, Jin M, Tang Z, Chen D, Wei X, You L. Generalized Zero-Shot Learning for Evolving Network Device Identification. Electronics. 2026; 15(11):2320. https://doi.org/10.3390/electronics15112320
Chicago/Turabian StyleWang, Zhihua, Minghui Jin, Zhenyu Tang, Duo Chen, Xingshen Wei, and Lizhao You. 2026. "Generalized Zero-Shot Learning for Evolving Network Device Identification" Electronics 15, no. 11: 2320. https://doi.org/10.3390/electronics15112320
APA StyleWang, Z., Jin, M., Tang, Z., Chen, D., Wei, X., & You, L. (2026). Generalized Zero-Shot Learning for Evolving Network Device Identification. Electronics, 15(11), 2320. https://doi.org/10.3390/electronics15112320

