Hybrid Time–Position Embedding for Provenance-Based Intrusion Detection
Abstract
1. Introduction
- We propose a novel provenance representation method using a hybrid time–position embedding mechanism. This method generates provenance graphs that better express system behavior by embedding procedural order and time intervals from system logs. It can be broadly applied to the embedding process of raw security data and contributes to improving dataset quality, particularly during preprocessing and dataset construction.
- We demonstrate that models trained on provenance graphs generated using our log embedding technique can more effectively capture the semantic information in system logs. Through experiments on the DARPA E3 benchmark dataset, we show that the GNN model trained on our proposed representation converges more quickly than baseline methods, indicating the effectiveness of our embedding in expressing node attributes in the provenance graph.
2. Related Work
2.1. Provenance- and Causality-Based Anomaly Detection Approaches
2.2. Semantic Representation for IDS
2.3. Temporal and Positional Encoding for Provenance Graphs
3. Method
3.1. Provenance Graph Parsing
3.2. Semantic Information and Hybrid Time–Position Embedding
| Algorithm 1 Hybrid Time–Position Embedding |
|
3.3. Graph Representation Learning
3.4. Attack Detection Phase
4. Experiment and Evaluation
- RQ1. Does Hybrid Time–Position Embedding improve node classification performance over traditional embedding methods?
- RQ2. Does Hybrid Time–Position Embedding achieve better convergence during iterative refinement than positional-only embedding?
- RQ3. Is the proposed method effective in detecting specific APT phases such as Initial Access, Persistence, and Exfiltration across the attack lifecycle?
4.1. Dataset
4.2. Baselines
- ThreaTrace: ThreaTrace identifies abnormal behavior by modeling the distributional patterns of system calls (syscall distributions). It characterizes process behavior through multi-model frameworks to maximize detection performance.
- FLASH: FLASH leverages Word2Vec-based semantic embeddings combined with positional encoding to construct meaningful node representations. It integrates a GNN with a lightweight classifier and applies iterative refinement to boost detection accuracy. FLASH also demonstrated superior scalability through its efficient architecture.
4.3. Implementation Details
4.4. Evaluation Metrics
- Precision, Recall, and F1-score: Precision measures the accuracy of malicious alerts, while Recall (Detection Rate) quantify the ability to capture all actual attack nodes. The F1-score provides the harmonic mean of these two, representing the overall balanced performance.
- False-Alarm Rate (FAR): Also known as the false positive rate, FAR is a crucial metric in the IDS domain. It represents the probability that benign system activities are incorrectly flagged as malicious (FP/(FP + TN)).
- PR-AUC: This represents the area under the Precision–Recall curve. In our ensemble-based detection, the PR-AUC is calculated by varying the consensus threshold (the number of models flagging an anomaly), offering a comprehensive view of the model’s performance beyond a single operational point.
4.5. RQ1: Classification Performance by Embedding Strategy
4.6. RQ2: Convergence Properties in Iterative Refinement
4.7. RQ3: Detection Performance by Attack Stage
4.8. Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| IDS | Intrusion Detection System |
| APT | Advanced Persistent Threat |
| ML | Machine Learning |
| DAG | Directed Acyclic Graph |
| UUID | Universally Unique Identifier |
| GNN | Graph Neural Network |
| TC | Transparent Computing |
| E3 | Engagement 3 (DARPA Transparent Computing Dataset) |
| ReLU | Rectified Linear Unit |
| OS | Operating System |
| LCS | Longest Common Subsequence |
| IP | Internet Protocol |
| CPU | Central Processing Unit |
| vCPU | Virtual Central Processing Unit |
| RAM | Random Access Memory |
Appendix A. Mapping Rules for Attack Phase Identification
| Attack Phase | Associated Processes (Exec) | System Call Actions |
|---|---|---|
| Initial Access | nginx, smtpd, imapd, inetd, sshd, proxymap, anvil, master, pickup, qmgr, ipop3d | EVENT_ACCEPT, EVENT_BIND, EVENT_RECVFROM, EVENT_RECVMSG |
| Execution | bash, sh, csh, python2.7, php-fpm, atrun, cron, expr, jot, minions, main, test, resizewin, fortune | EVENT_EXECUTE, EVENT_FORK, EVENT_MODIFY_PROCESS, EVENT_SIGNAL, EVENT_EXIT |
| Persistence | cron, atrun, sshd, inetd, screen, mkdir, cp, mount, nohup, tmux-1002, rc | EVENT_CREATE_OBJECT, EVENT_WRITE, EVENT_LINK, EVENT_MODIFY_FILE_ATTRIBUTES, EVENT_CLOSE |
| Privilege Escalation | sudo, su, pkg, doas | EVENT_CHANGE_PRINCIPAL |
| Defense Evasion | rm, unlink, mv, newsyslog, sleep, syslogd, chgrp, lockf, dd, chmod, chown, cleanup, history, ipfw, pfctl | EVENT_UNLINK, EVENT_RENAME, EVENT_TRUNCATE |
| Discovery | find, netstat, ifconfig, ps, lsof, df, ls, tail, head, cat, top, vmstat, sysctl, dmesg, uptime, hostname, route, env, kenv, less, mount, wc, cmp, date, stat, whoami, id, groups, uname, tty, limits, pwait, kldstat | EVENT_OPEN, EVENT_READ, EVENT_LSEEK, EVENT_FLOWS_TO, EVENT_FCNTL, EVENT_MMAP |
| Collection | cat, grep, egrep, awk, sed, sort, bzip2, xz, bzcat, dd, cp, tee, tar, zip, gzip, uniq, tr, cut, nawk, mktemp, md5 | EVENT_READ, EVENT_OPEN |
| Exfiltration | wget, mail, sendmail, mailwrapper, alpine, msgs, sshd, curl, ftp, nc, scp, ssh | EVENT_SENDTO, EVENT_SENDMSG, EVENT_CONNECT |
References
- Kushner, D. The real story of stuxnet. IEEE Spectr. 2013, 50, 48–53. [Google Scholar] [CrossRef] [Scilit]
- Haggard, S.; Lindsay, J.R. North Korea and the Sony Hack: Exporting Instability Through Cyberspace; JSTOR: New York, NY, USA, 2015. [Google Scholar]
- Krasznay, C. Case Study: The Notpetya Campaign. In Információés Kiberbiztonság; Ludovika Egyetemi Kiadó: Budapest, Hungary, 2020; pp. 485–499. [Google Scholar]
- Alkhadra, R.; Abuzaid, J.; AlShammari, M.; Mohammad, N. Solar winds hack: In-depth analysis and countermeasures. In Proceedings of the 2021 12th International Conference on Computing Communication and Networking Technologies (ICCCNT); IEEE: Piscataway, NJ, USA, 2021; pp. 1–7. [Google Scholar]
- Alshamrani, A.; Myneni, S.; Chowdhary, A.; Huang, D. A survey on advanced persistent threats: Techniques, solutions, challenges, and research opportunities. IEEE Commun. Surv. Tutorials 2019, 21, 1851–1877. [Google Scholar] [CrossRef] [Scilit]
- Talib, M.A.; Nasir, Q.; Nassif, A.B.; Mokhamed, T.; Ahmed, N.; Mahfood, B. APT beaconing detection: A systematic review. Comput. Secur. 2022, 122, 102875. [Google Scholar] [CrossRef] [Scilit]
- Li, Z.; Cheng, X.; Sun, L.; Zhang, J.; Chen, B. A hierarchical approach for advanced persistent threat detection with attention-based graph neural networks. Secur. Commun. Netw. 2021, 2021, 9961342. [Google Scholar] [CrossRef] [Scilit]
- Ahmad, R.; Alsmadi, I.; Alhamdani, W.; Tawalbeh, L. Zero-day attack detection: A systematic literature review. Artif. Intell. Rev. 2023, 56, 10733–10811. [Google Scholar] [CrossRef] [Scilit]
- Ali, S.; Rehman, S.U.; Imran, A.; Adeem, G.; Iqbal, Z.; Kim, K.I. Comparative evaluation of ai-based techniques for zero-day attacks detection. Electronics 2022, 11, 3934. [Google Scholar] [CrossRef] [Scilit]
- Santhosh Kumar, S.; Selvi, M.; Kannan, A. A comprehensive survey on machine learning-based intrusion detection systems for secure communication in internet of things. Comput. Intell. Neurosci. 2023, 2023, 8981988. [Google Scholar] [CrossRef] [Scilit]
- Thakkar, A.; Lohiya, R. A review on machine learning and deep learning perspectives of IDS for IoT: Recent updates, security issues, and challenges. Arch. Comput. Methods Eng. 2021, 28, 3211–3243. [Google Scholar] [CrossRef] [Scilit]
- Bilot, T.; Jiang, B.; Li, Z.; El Madhoun, N.; Al Agha, K.; Zouaoui, A.; Pasquier, T. Sometimes Simpler is Better: A Comprehensive Analysis of State-of-the-Art Provenance-Based Intrusion Detection Systems. In Proceedings of the 34th USENIX Security Symposium (USENIX Security 25); USENIX: Berkeley, CA, USA, 2025; pp. 7193–7212. [Google Scholar]
- Zipperle, M.; Gottwalt, F.; Chang, E.; Dillon, T. Provenance-based intrusion detection systems: A survey. ACM Comput. Surv. 2022, 55, 1–36. [Google Scholar] [CrossRef] [Scilit]
- Alsaheel, A.; Nan, Y.; Ma, S.; Yu, L.; Walkup, G.; Celik, Z.B.; Zhang, X.; Xu, D. {ATLAS}: A sequence-based learning approach for attack investigation. In Proceedings of the 30th USENIX Security Symposium (USENIX Security 21); USENIX: Berkeley, CA, USA, 2021; pp. 3005–3022. [Google Scholar]
- Wu, L.; Xie, Y.; Wu, Y.; Liang, J.; Li, X. Provenance Based Intrusion Detection via Measuring Provenance Sequence Similarity. In Proceedings of the 2022 International Conference on Blockchain Technology and Information Security (ICBCTIS); IEEE: Piscataway, NJ, USA, 2022; pp. 198–201. [Google Scholar]
- Liu, M.; Xue, Z.; Xu, X.; Zhong, C.; Chen, J. Host-based intrusion detection system with system calls: Review and future trends. ACM Comput. Surv. (CSUR) 2018, 51, 1–36. [Google Scholar] [CrossRef] [Scilit]
- Goyal, A.; Han, X.; Wang, G.; Bates, A. Sometimes, you aren’t what you do: Mimicry attacks against provenance graph host intrusion detection systems. In Proceedings of the 30th Network and Distributed System Security Symposium, San Diego, CA, USA, 27 February–3 March 2023. [Google Scholar]
- Inam, M.A.; Chen, Y.; Goyal, A.; Liu, J.; Mink, J.; Michael, N.; Gaur, S.; Bates, A.; Hassan, W.U. Sok: History is a vast early warning system: Auditing the provenance of system intrusions. In Proceedings of the 2023 IEEE Symposium on Security and Privacy (SP); IEEE: Piscataway, NJ, USA, 2023; pp. 2620–2638. [Google Scholar]
- Scarselli, F.; Gori, M.; Tsoi, A.C.; Hagenbuchner, M.; Monfardini, G. The graph neural network model. IEEE Trans. Neural Netw. 2008, 20, 61–80. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, Q.; Hassan, W.U.; Li, D.; Jee, K.; Yu, X.; Zou, K.; Rhee, J.; Chen, Z.; Cheng, W.; Gunter, C.A.; et al. You Are What You Do: Hunting Stealthy Malware via Data Provenance Analysis. In Proceedings of the NDSS, San Diego, CA, USA, 23–26 February 2020. [Google Scholar]
- Le, Q.; Mikolov, T. Distributed representations of sentences and documents. In Proceedings of the International Conference on Machine Learning. PMLR; ML Research Press: Cambridge, MA, USA, 2014; pp. 1188–1196. [Google Scholar]
- Irshad, H.; Ciocarlie, G.; Gehani, A.; Yegneswaran, V.; Lee, K.H.; Patel, J.; Jha, S.; Kwon, Y.; Xu, D.; Zhang, X. Trace: Enterprise-wide provenance tracking for real-time apt detection. IEEE Trans. Inf. Forensics Secur. 2021, 16, 4363–4376. [Google Scholar] [CrossRef] [Scilit]
- Wang, S.; Wang, Z.; Zhou, T.; Sun, H.; Yin, X.; Han, D.; Zhang, H.; Shi, X.; Yang, J. Threatrace: Detecting and tracing host-based threats in node level through provenance graph learning. IEEE Trans. Inf. Forensics Secur. 2022, 17, 3972–3987. [Google Scholar] [CrossRef] [Scilit]
- Hamilton, W.; Ying, Z.; Leskovec, J. Inductive representation learning on large graphs. In Advances in Neural Information Processing Systems; NIPS: Grenada, Spain, 2017; Volume 30. [Google Scholar]
- Zeng, J.; Zhang, C.; Liang, Z. Palantír: Optimizing attack provenance with hardware-enhanced system observability. In Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security; ACM: New York, NY, USA, 2022; pp. 3135–3149. [Google Scholar]
- Wu, Y.; Xie, Y.; Liao, X.; Zhou, P.; Feng, D.; Wu, L.; Li, X.; Wildani, A.; Long, D. Paradise: Real-time, generalized, and distributed provenance-based intrusion detection. IEEE Trans. Dependable Secur. Comput. 2022, 20, 1624–1640. [Google Scholar] [CrossRef] [Scilit]
- Xu, Z.; Fang, P.; Liu, C.; Xiao, X.; Wen, Y.; Meng, D. Depcomm: Graph summarization on system audit logs for attack investigation. In Proceedings of the 2022 IEEE Symposium on Security and Privacy (SP); IEEE: Piscataway, NJ, USA, 2022; pp. 540–557. [Google Scholar]
- Wu, L.; Xie, Y.; Li, J.; Feng, D.; Liang, J.; Wu, Y. Angus: Efficient active learning strategies for provenance based intrusion detection. Cybersecurity 2025, 8, 6. [Google Scholar] [CrossRef] [Scilit]
- Shen, Y.; Stringhini, G. {ATTACK2VEC}: Leveraging temporal word embeddings to understand the evolution of cyberattacks. In Proceedings of the 28th USENIX Security Symposium (USENIX Security 19); USENIX: Berkeley, CA, USA, 2019; pp. 905–921. [Google Scholar]
- Zeng, J.; Chua, Z.L.; Chen, Y.; Ji, K.; Liang, Z.; Mao, J. WATSON: Abstracting Behaviors from Audit Logs via Aggregation of Contextual Semantics. In Proceedings of the NDSS, Online, 21–25 February 2021. [Google Scholar]
- Van Ede, T.; Aghakhani, H.; Spahn, N.; Bortolameotti, R.; Cova, M.; Continella, A.; Van Steen, M.; Peter, A.; Kruegel, C.; Vigna, G. Deepcase: Semi-supervised contextual analysis of security events. In Proceedings of the 2022 IEEE Symposium on Security and Privacy (SP); IEEE: Piscataway, NJ, USA, 2022; pp. 522–539. [Google Scholar]
- Rehman, M.U.; Ahmadi, H.; Hassan, W.U. Flash: A comprehensive approach to intrusion detection via provenance graph representation learning. In Proceedings of the 2024 IEEE Symposium on Security and Privacy (SP); IEEE: Piscataway, NJ, USA, 2024; pp. 3552–3570. [Google Scholar]
- Jia, Z.; Xiong, Y.; Nan, Y.; Zhang, Y.; Zhao, J.; Wen, M. {MAGIC}: Detecting advanced persistent threats via masked graph representation learning. In Proceedings of the 33rd USENIX Security Symposium (USENIX Security 24); USENIX: Berkeley, CA, USA, 2024; pp. 5197–5214. [Google Scholar]
- Han, X.; Pasquier, T.; Bates, A.; Mickens, J.; Seltzer, M. Unicorn: Runtime provenance-based detector for advanced persistent threats. arXiv 2020, arXiv:2001.01525. [Google Scholar] [CrossRef] [Scilit]
- Yang, F.; Xu, J.; Xiong, C.; Li, Z.; Zhang, K. {PROGRAPHER}: An anomaly detection system based on provenance graph embedding. In Proceedings of the 32nd USENIX Security Symposium (USENIX Security 23); USENIX: Berkeley, CA, USA, 2023; pp. 4355–4372. [Google Scholar]
- Meng, L.; Xi, R.; Li, Z.; Zhu, H. PG-AID: An Anomaly-based Intrusion Detection Method Using Provenance Graph. In Proceedings of the 2024 27th International Conference on Computer Supported Cooperative Work in Design (CSCWD); IEEE: Piscataway, NJ, USA, 2024; pp. 2522–2527. [Google Scholar]
- Jiang, B.; Bilot, T.; El Madhoun, N.; Al Agha, K.; Zouaoui, A.; Iqbal, S.; Han, X.; Pasquier, T. ORTHRUS: Achieving High Quality of Attribution in Provenance-based Intrusion Detection Systems. In Proceedings of the Security Symposium (USENIX Sec’25); USENIX: Berkeley, CA, USA, 2025. [Google Scholar]
- Mikolov, T.; Chen, K.; Corrado, G.; Dean, J. Efficient estimation of word representations in vector space. arXiv 2013, arXiv:1301.3781. [Google Scholar] [CrossRef] [Scilit]
- Liu, F.; Wen, Y.; Zhang, D.; Jiang, X.; Xing, X.; Meng, D. Log2vec: A heterogeneous graph embedding based approach for detecting cyber threats within enterprise. In Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security; ACM: New York, NY, USA, 2019; pp. 1777–1794. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. In Advances in Neural Information Processing Systems; NIPS: Grenada, Spain, 2017; Volume 30. [Google Scholar]




| Dataset | Events | Nodes | Execs | Paths | Malicious Nodes |
|---|---|---|---|---|---|
| Cadets | 2,059,154 | 362,645 | 85 | 96,019 | 12,858 |
| Trace | 2,472,290 | 1,271,939 | 1 | 52,989 | 68,173 |
| Method | Metric | Cadets | Trace |
|---|---|---|---|
| ThreaTrace [23] | Precision | 0.90 | 0.72 |
| Recall | 0.99 | 0.99 | |
| F1-score | 0.95 | 0.83 | |
| Flash [32] | Precision | 0.952 | 0.951 |
| Recall | 0.999 | 0.986 | |
| F1-score | 0.962 | 0.969 | |
| FAR | 0.002964 | 0.003124 | |
| PR-AUC | 0.9102 | 0.9838 | |
| Proposed | Precision | 0.967 | 0.953 |
| Recall | 0.999 | 0.988 | |
| F1-score | 0.983 | 0.969 | |
| FAR | 0.00132 | 0.003077 | |
| PR-AUC | 0.9257 | 0.9869 |
| Attack Stage | Cadets | Trace | ||||
|---|---|---|---|---|---|---|
| TP | FP | FN | TP | FP | FN | |
| Discovery | 8 | 8 | 70 | 17 | 500 | 0 |
| Collection | 0 | 2 | 0 | 0 | 0 | 0 |
| Execution | 12,831 | 40 | 0 | 8 | 0 | 0 |
| Privilege Escalation | 0 | 1 | 0 | 0 | 0 | 0 |
| Defense Evasion | 0 | 8 | 0 | 1 | 1026 | 0 |
| Initial Access | 10 | 217 | 0 | 0 | 33 | 0 |
| Persistence | 2 | 155 | 0 | 12 | 1529 | 0 |
| Exfiltration | 0 | 3 | 1 | 67,338 | 387 | 0 |
| Unknown Patterns | 0 | 11 | 6 | 0 | 0 | 790 |
| Total Impact | 12,851 | 445 | 77 | 67,376 | 3475 | 790 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Gong, S.; Cho, J.; Choi, K.K. Hybrid Time–Position Embedding for Provenance-Based Intrusion Detection. Electronics 2026, 15, 1004. https://doi.org/10.3390/electronics15051004
Gong S, Cho J, Choi KK. Hybrid Time–Position Embedding for Provenance-Based Intrusion Detection. Electronics. 2026; 15(5):1004. https://doi.org/10.3390/electronics15051004
Chicago/Turabian StyleGong, Seonghyeon, Jake Cho, and Kyuwon Ken Choi. 2026. "Hybrid Time–Position Embedding for Provenance-Based Intrusion Detection" Electronics 15, no. 5: 1004. https://doi.org/10.3390/electronics15051004
APA StyleGong, S., Cho, J., & Choi, K. K. (2026). Hybrid Time–Position Embedding for Provenance-Based Intrusion Detection. Electronics, 15(5), 1004. https://doi.org/10.3390/electronics15051004

