Recent Advances and Open Challenges in Mitigating Inference-Time Attacks on Large Language Models
Abstract
1. Introduction
- LLM vulnerabilities are classified within a three-layered attack surface taxonomy based on the lifecycle stage in which they emerge, advocating for the practical and theoretical priority of the inference-time category.
- An original defense taxonomy organized across three axes is proposed, within which 30 defense mechanisms published from 2024 onwards are systematically examined for their effectiveness against specific inference-time attack families.
- An intersectional comparative analysis covering defense–attack coverage gaps, security-utility-latency trade-offs, and white-box versus black-box applicability is conducted to provide actionable guidance for practitioners regarding defense configuration selection in real-world deployments.
- A deployment-oriented defense framework that translates the comparative findings of this survey into practical defense recommendations for representative LLM deployment scenarios. Unlike existing surveys that primarily compare defense techniques individually, the proposed framework maps deployment constraints to appropriate defense stacks while explicitly highlighting their residual security gaps.
2. Background
2.1. LLM Architecture and Inference Pipeline
2.2. Attack Surface Taxonomy
2.3. Traditional Defense Methods for Inference Time Attacks
3. Inference-Time Attacks
3.1. Prompt Injection Attacks
3.2. Jailbreaking Attacks
3.3. Adaptive Attacks
3.4. Adversarial Input Attacks
3.5. Information Disclosure Attacks
3.6. Insecure Output Handling
3.7. Excessive Agency
3.8. Unbounded Consumption
4. Defense Strategies Against Inference-Time Attacks
4.1. Prompt Level Defenses
4.1.1. Machine Learning (ML) Based Classifier
4.1.2. Embedding Based Classifier
4.1.3. Pre-Processing
4.1.4. LLM Based Classifier
4.1.5. Hybrid Classifier
4.2. Inference-Time Defenses
4.2.1. Self-Consistency
4.2.2. Input/Output Classification
4.3. Training-Time Defenses
4.3.1. Model Unlearning
4.3.2. Adversarial Training
5. Comparative Analysis
5.1. Defense-Attack Coverage Matrix
5.2. Trade-Offs: Security, Utility, Latency
5.3. White-Box, Black-Box Setting Efficacy
5.4. Deployment-Oriented Defense Framework
6. Benchmarks and Evaluation
6.1. Current Datasets and Benchmarks
6.2. Evaluation Metrics
6.3. Limitations of Evaluations and Judge Reliability
7. Open Challenges and Future Directions
8. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Villaplana, A.; Martínez, R.; Montalvo, S. Improving medical entity recognition in spanish by means of biomedical language models. Electronics 2023, 12, 4872. [Google Scholar] [CrossRef] [Scilit]
- Ma, H.; Lu, Y.; Xiao, Z.; Feng, J.; Zhang, H.; Yu, J. SDD-LawLLM: Advancing intelligent legal systems through synthetic data-driven fine-tuning of large language models. Electronics 2025, 14, 742. [Google Scholar] [CrossRef] [Scilit]
- Staegemann, D.; Haertel, C.; Daase, C.; Pohl, M.; Abdallah, M.; Turowski, K. A Review on Large Language Models and Generative AI in Banking. In Proceedings of the 7th International Conference on Finance, Economics, Management and IT Business (FEMIB 2025); SCITEPRESS: Setúbal, Portugal, 2025; pp. 267–278. [Google Scholar]
- Shen, T.; Jin, R.; Huang, Y.; Liu, C.; Dong, W.; Guo, Z.; Wu, X.; Liu, Y.; Xiong, D. Large Language Model Alignment: A Survey. arXiv 2023, arXiv:2309.15025. [Google Scholar] [CrossRef] [Scilit]
- OWASP Foundation. OWASP Top 10 for LLM Applications. 2025. Available online: https://genai.owasp.org/llm-top-10/ (accessed on 24 June 2026).
- Perez, F.; Ribeiro, I. Ignore Previous Prompt: Attack Techniques for Language Models. In Proceedings of the NeurIPS ML Safety Workshop, New Orleans, LA, USA, 28 November–9 December 2022. [Google Scholar]
- Zhou, Y.; Ali, M.; Lee, W.B.; Zhao, Q. A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluation Methods. Trans. Artif. Intell. 2025, 1, 28–58. [Google Scholar] [CrossRef] [Scilit]
- Das, B.C.; Amini, M.H.; Wu, Y. Security and privacy challenges of large language models: A survey. ACM Comput. Surv. 2025, 57, 1–39. [Google Scholar] [CrossRef] [Scilit]
- Liao, Z.; Chen, K.; Lin, Y.; Li, K.; Liu, Y.; Chen, H.; Huang, X.; Yu, Y. Attack and defense techniques in large language models: A survey and new perspectives. Neural Netw. 2025, 196, 108388. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, X.; Wang, W.; Ji, Z.; Li, Z.; Ma, P.; Wu, D.; Wang, S. Stshield: Single-token sentinel for real-time jailbreak detection in large language models. arXiv 2025, arXiv:2503.17932. [Google Scholar] [CrossRef] [Scilit]
- Xu, W.; Parhi, K.K. A survey of attacks on large language models. arXiv 2025, arXiv:2505.12567. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention Is All You Need. In Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA, 4 December 2017; Curran Associates Inc.: Red Hook, NY, USA, 2017; Volume 30, pp. 6000–6010. [Google Scholar]
- Huang, Y.; Xu, J.; Jiang, Z.; Lai, J.; Li, Z.; Yao, Y.; Chen, T.; Yang, L.; Xin, Z.; Ma, X. Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey. arXiv 2023, arXiv:2311.12351. [Google Scholar] [CrossRef] [Scilit]
- Li, B.; Jiang, Y.; Gadepally, V.; Tiwari, D. Llm inference serving: Survey of recent advances and opportunities. In Proceedings of the 2024 IEEE High Performance Extreme Computing Conference (HPEC); IEEE: New York, NY, USA, 2025; pp. 1–8. [Google Scholar]
- IBM. What Is LLM Inference? 2024. Available online: https://www.ibm.com/think/topics/llm-inference (accessed on 24 June 2026).
- Sajjadi Mohammadabadi, S.M.; Kara, B.C.; Eyupoglu, C.; Uzay, C.; Tosun, M.S.; Karakuş, O. A survey of large language models: Evolution, architectures, adaptation, benchmarking, applications, challenges, and societal implications. Electronics 2025, 14, 3580. [Google Scholar] [CrossRef] [Scilit]
- Yao, Y.; Duan, J.; Xu, K.; Cai, Y.; Sun, Z.; Zhang, Y. A survey on large language model (llm) security and privacy: The good, the bad, and the ugly. High-Confid. Comput. 2024, 4, 100211. [Google Scholar] [CrossRef] [Scilit]
- Greshake, K.; Abdelnabi, S.; Mishra, S.; Endres, C.; Holz, T.; Fritz, M. Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial İntelligence and Security; ACM: New York, NY, USA, 2023; pp. 79–90. [Google Scholar]
- Lee, S.; Kim, J.; Pak, W. Mind Mapping Prompt Injection: Visual Prompt Injection Attacks in Modern Large Language Models. Electronics 2025, 14, 1907. [Google Scholar] [CrossRef] [Scilit]
- Zou, A.; Wang, Z.; Kolter, J.Z.; Fredrikson, M. Universal and Transferable Adversarial Attacks on Aligned Language Models. arXiv 2023, arXiv:2307.15043. [Google Scholar] [CrossRef] [Scilit]
- Kwon, H.; Pak, W. Text-based prompt injection attack using mathematical functions in modern large language models. Electronics 2024, 13, 5008. [Google Scholar] [CrossRef] [Scilit]
- He, J.; Hou, G.; Jia, X.; Chen, Y.; Liao, W.; Zhou, Y.; Zhou, R. Data stealing attacks against large language models via backdooring. Electronics 2024, 13, 2858. [Google Scholar] [CrossRef] [Scilit]
- Tramèr, F.; Zhang, F.; Juels, A.; Reiter, M.K.; Ristenpart, T. Stealing machine learning models via prediction {APIs}. In Proceedings of the 25th USENIX security symposium (USENIX Security 16), Austin, TX, USA, 10–12 August 2016; pp. 601–618. [Google Scholar]
- Shokri, R.; Stronati, M.; Song, C.; Shmatikov, V. Membership inference attacks against machine learning models. In Proceedings of the 2017 IEEE Symposium on Security and Privacy (SP); IEEE: New York, NY, USA, 2017; pp. 3–18. [Google Scholar]
- Dong, Z.; Zhou, Z.; Yang, C.; Shao, J.; Qiao, Y. Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; pp. 6734–6747. [Google Scholar]
- Liu, Y.; Deng, G.; Li, Y.; Wang, K.; Wang, Z.; Wang, X.; Zhang, T.; Liu, Y.; Wang, H.; Zheng, Y.; et al. Prompt injection attack against llm-integrated applications. arXiv 2023, arXiv:2306.05499. [Google Scholar] [CrossRef] [Scilit]
- Yi, S.; Liu, Y.; Sun, Z.; Cong, T.; He, X.; Song, J.; Xu, K.; Li, Q. Jailbreak attacks and defenses against large language models: A survey. arXiv 2024, arXiv:2407.04295. [Google Scholar] [CrossRef] [Scilit]
- Andriushchenko, M.; Flammarion, N. Jailbreaking leading safety-aligned llms with simple adaptive attacks. In Proceedings of the International Conference on Learning Representations, Singapore, 24–28 April 2025; Volume 2025, pp. 40116–40143. [Google Scholar]
- Hagendorff, T.; Derner, E.; Oliver, N. Large reasoning models are autonomous jailbreak agents. Nat. Commun. 2026, 17, 1435. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Weng, L. Adversarial Attacks on LLMs. 2023. Available online: https://lilianweng.github.io/ (accessed on 2 June 2026).
- Wang, Y.; Chen, M.; Peng, N.; Chang, K.W. Vulnerability of large language models to output prefix jailbreaks: Impact of positions on safety. In Proceedings of the Findings of the Association for Computational Linguistics: NAACL 2025; Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 3939–3952. [Google Scholar]
- Staab, R.; Vero, M.; Balunovic, M.; Vechev, M. Beyond memorization: Violating privacy via inference with large language models. In Proceedings of the International Conference on Learning Representations, Vienna, Austria, 7 May 2024; Volume 2024, pp. 33832–33878. [Google Scholar]
- Carlini, N.; Tramer, F.; Wallace, E.; Jagielski, M.; Herbert-Voss, A.; Lee, K.; Roberts, A.; Brown, T.; Song, D.; Erlingsson, U.; et al. Extracting training data from large language models. In Proceedings of the 30th USENIX security symposium (USENIX Security 21), Vancouver, BC, Canada, 11–13 August 2021; pp. 2633–2650. [Google Scholar]
- Naik, D.; Naik, I.; Naik, N. Insecure output handling in large language models (LLMs) and approaches to enhance output security, including prevention of LLM-based web application attacks. In Proceedings of the International Conference on Computing, Communication, Cybersecurity & AI; Springer: Cham Switzerland, 17 May 2026; pp. 695–720. [Google Scholar]
- Maiorano, A.C. Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Cover-age and Its Brittleness Under Paraphrasing. arXiv 2026, arXiv:2606.02822. [Google Scholar] [CrossRef] [Scilit]
- Kokkula, S.; Divya, G. Palisade–Prompt Injection Detection Framework. arXiv 2024, arXiv:2410.21146. [Google Scholar] [CrossRef] [Scilit]
- Esugo, M.; Alao, O.; Mahmoud, H. SHIELD: Security against Harmful Prompt Injection Evaluation and Language Detection Leveraging Ensemble Approach. In Proceedings of the IECON 2025–51st Annual Conference of the IEEE Industrial Electronics Society; IEEE: New York, NY, USA, 2025; pp. 1–6. [Google Scholar]
- Prakash, C.; Lind, M.; De La Cruz, E. Hybrid Real-time Framework for Detecting Adaptive Prompt Injection Attacks in Large Language Models. J. Comput. Theor. Appl. 2026, 3, 286–302. [Google Scholar] [CrossRef] [Scilit]
- Alamsabi, M.; Tchuindjang, M.; Brohi, S. Embedding-Based Detection of Indirect Prompt Injection Attacks in Large Language Models Using Semantic Context Analysis. Algorithms 2026, 19, 92. [Google Scholar] [CrossRef] [Scilit]
- Ayub, A.; Majumdar, S. Embedding-based classifiers can detect prompt injection attacks. arXiv 2024, arXiv:2410.22284. [Google Scholar] [CrossRef] [Scilit]
- Pingua, B.; Murmu, D.; Kandpal, M.; Rautaray, J.; Mishra, P.; Barik, R.K.; Saikia, M.J. Mitigating adversarial manipulation in LLMs: A prompt-based approach to counter Jailbreak attacks (Prompt-G). PeerJ Comput. Sci. 2024, 10, e2374. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mishra, A.; Preet, S.; Gupta, B.B.; Rawat, S.S.; Arya, V.; Katiyar, V. Mitigating Prompt Injection Attacks in ModelAgnostic Networks (MAN). In Proceedings of the 2025 5th International Conference on Internet of Things: Smart Innovation and Usages (IoT-SIU); IEEE: New York, NY, USA, 2025; pp. 1–5. [Google Scholar]
- Hu, J.; Wang, H.; Mukherjee, D.; Paschalidis, I.C. CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection. arXiv 2025, arXiv:2508.14128. [Google Scholar] [CrossRef] [Scilit]
- Jacob, D.; Alzahrani, H.; Hu, Z.; Alomair, B.; Wagner, D. PromptShield: Deployable Detection for Prompt Injection Attacks. In Proceedings of the Fifteenth ACM Conference on Data and Application Security and Privacy (CODASPY ’25); ACM: New York, NY, USA, 2025. [Google Scholar]
- Pan, J.; Wong, S.L.; Yuan, Y.; Chia, X.W. Prompt Inject Detection with Generative Explanation as an Investigative Tool. In Proceedings of the 2025 International Conference on Machine Learning and Cybernetics (ICMLC); IEEE: New York, NY, USA, 2025; pp. 7–12. [Google Scholar]
- Thaqi, R.; Martiri, E.; Rexha, B. A Real-Time Framework for Prompt Injection Attacks Detection on Cloud-Hosted Large Language Models. In Proceedings of the 2025 3rd International Conference on Foundation and Large Language Models (FLLM); IEEE: New York, NY, USA, 2025; pp. 718–723. [Google Scholar]
- Aswin Mallessh, N.S.; Ilavendhan, A. Input Moderation and Injection Filtering in Large Language Model via Llama Guard Integration. In Proceedings of the 2025 IEEE Pune Section International Conference (PuneCon); IEEE: New York, NY, USA, 2025; pp. 1–6. [Google Scholar]
- Liu, Y.; Jia, Y.; Jia, J.; Song, D.; Gong, N.Z. Datasentinel: A game-theoretic detection of prompts injection attacks. In Proceedings of the 2025 IEEE Symposium on Security and Privacy (SP); IEEE: New York, NY, USA, 2025; pp. 2190–2208. [Google Scholar]
- Hasan, M.M.; Rahman, Z.; Mostafiz, R.; Hossain, M.A. Sentra-Guard: A Multilingual Human-AI Framework for Real-Time Defense Against Adversarial LLM Jailbreaks. arXiv 2025, arXiv:2510.22628. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Z.; Zhang, Q.; Foerster, J. PARDEN, can you repeat that? defending against jailbreaks via repetition. In Proceedings of the 41st International Conference on Machine Learning; ICML’24; JMLR.org: Brookline, MA, USA, 2024. [Google Scholar]
- Cao, Y.; Gu, N.; Shen, X.; Yang, D.; Zhang, X. Defending large language models against jailbreak attacks through chain of thought prompting. In Proceedings of the 2024 International Conference on Networking and Network Applications (NaNA); IEEE: New York, NY, USA, 2024; pp. 125–130. [Google Scholar]
- Zhao, W.; Peng, J.; Ben-Levi, D.; Yu, Z.; Yang, J. Proactive defense against LLM Jailbreak. arXiv 2025, arXiv:2510.05052. [Google Scholar] [CrossRef] [Scilit]
- Schwarz, D. Countermind: A Multi-Layered Security Architecture for Large Language Models. arXiv 2025, arXiv:2510.11837. [Google Scholar] [CrossRef] [Scilit]
- Panebianco, F.; Bonfanti, S.; Trovò, F.; Carminati, M. LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks. arXiv 2025, arXiv:2508.00602. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Z.; Lin, Y.; An, X.; Wan, M.; Jiang, C.; Ding, N. ATTN-Defense: Attention-Guided Detection, Location and Removal for Indirect Prompt Injection. In Proceedings of the ICASSP 2026–2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE: New York, NY, USA, 2026; pp. 17532–17536. [Google Scholar]
- Li, X.; Wu, X.; Li, Q.; Ni, J.; Lu, R. SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks. arXiv 2025, arXiv:2508.15182. [Google Scholar] [CrossRef] [Scilit]
- Xhonneux, S.; Sordoni, A.; Günnemann, S.; Gidel, G.; Schwinn, L. Efficient adversarial training in llms with continuous attacks. Adv. Neural Inf. Process. Syst. 2024, 37, 1502–1530. [Google Scholar] [CrossRef] [Scilit]
- Yu, L.; Do, V.; Hambardzumyan, K.; Cancedda, N. Robust LLM safeguarding via refusal feature adversarial training. In Proceedings of the International Conference on Learning Representations, Singapore, 24–28 April 2025; Volume 2025, pp. 5254–5277. [Google Scholar]
- Van Huong, P. Optimizing Transformer Models for Prompt Jailbreak Attack Detection in AI Assistant Systems. In Proceedings of the 2024 1st International Conference on Cryptography and Information Security (VCRIS); IEEE: New York, NY, USA, 2024; pp. 1–4. [Google Scholar]
- Galinkin, E.; Sablotny, M. Improved large language model jailbreak detection via pretrained embeddings. arXiv 2024, arXiv:2412.01547. [Google Scholar] [CrossRef] [Scilit]
- Lan, Q.; Kaul, A.; Jones, S. Prompt Injection Detection in LLM Integrated Applications. Int. J. Netw. Dyn. Intell. 2025, 4, 100013. [Google Scholar] [CrossRef] [Scilit]
- Lin, H.; Lao, Y.; Geng, T.; Yu, T.; Zhao, W. Uniguardian: A unified defense for detecting prompt injection, backdoor attacks and adversarial attacks in large language models. arXiv 2025, arXiv:2502.13141. [Google Scholar] [CrossRef] [Scilit]
- Lan, Q.; Kaul, A.; Jones, S.; Westrum, S.; Pandurangan, V.; Pattanaik, N.K.D.; Pattanayak, P. Hybrid Constitutional Classifiers for Prompt Injection Defense. In Proceedings of the 2025 IEEE International Conference on Electro Information Technology (eIT); IEEE: New York, NY, USA, 2025; pp. 225–229. [Google Scholar]
- Zhao, Y.; Li, X. A Novel Security Framework against Prompt Injection Attacks. In Proceedings of the 5th International Conference on Computer Communication and Artificial Intelligence; IEEE: New York, NY, USA, 2025. [Google Scholar]
- Mazeika, M.; Phan, L.; Yin, X.; Zou, A.; Wang, Z.; Mu, N.; Sakhaee, E.; Li, N.; Basart, S.; Li, B.; et al. HarmBench: A standardized evaluation framework for automated red teaming and robust refusal. In Proceedings of the 41st International Conference on Machine Learning; ICML’24; JMLR.org: Brookline, MA, USA, 2024. [Google Scholar]
- Lin, Z.; Wang, Z.; Tong, Y.; Wang, Y.; Guo, Y.; Wang, Y.; Shang, J. Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2023; Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 4694–4702. [Google Scholar]
- deepset. Prompt Injections Dataset. Available online: https://huggingface.co/datasets/deepset/prompt-injections (accessed on 5 June 2026).
- Chao, P.; Debenedetti, E.; Robey, A.; Andriushchenko, M.; Croce, F.; Sehwag, V.; Dobriban, E.; Flammarion, N.; Pappas, G.J.; Tramèr, F.; et al. JailbreakBench: An open robustness benchmark for jailbreaking large language models. In Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, BC, Canada, 10–15 December 2024; NeurIPS: Vancouver, BC, Canada, 2024. [Google Scholar]
- Shen, X.; Chen, Z.; Backes, M.; Shen, Y.; Zhang, Y. “do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security; ACM: New York, NY, USA, 2024; pp. 1671–1685. [Google Scholar]
- Zhang, Z. Attn-Defense. 2026. Available online: https://github.com/Ziyang-Zhang-6657/Attn-Defense (accessed on 24 June 2026).
- Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; Hashimoto, T.B. Stanford Alpaca: An Instruction-Following LLaMA Model. 2023. Available online: https://github.com/tatsu-lab/stanford_alpaca (accessed on 12 June 2026).
- Kamath, A.; Singla, K.; Paul, R.; Joshi, R.B.; Vaidya, U.; Chauhan, S.S.; Wartikar, N. Benchmarking Hindi LLMs: A New Suite of Datasets and a Comparative Analysis. In Proceedings of the 1st Workshop on Benchmarks, Harmonization, Annotation, and Standardization for Human-Centric AI in Indian Languages (BHASHA 2025), Mumbai, India, 23–24 December 2025; Bhattacharya, A., Goyal, P., Ghosh, S., Ghosh, K., Eds.; Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 52–68. [Google Scholar] [CrossRef] [Scilit]
- Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C.D.; Ng, A.; Potts, C. Recursive Deep Models for Semantic Compositionality over a Sentiment Treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, Seattle, WA, USA, 18–21 October 2013; Association for Computational Linguistics: Stroudsburg, PA, USA, 2013; pp. 1631–1642. [Google Scholar]





| Publication (Year) | Defense Method | DPI | IPI | JB | AA | AIA | IDA | IOH | EA | UC |
|---|---|---|---|---|---|---|---|---|---|---|
| Hu et al. (2025) [43] | Input Pre-processing | • | • | • | • | ◦ | ||||
| Pingua et al. (2024) [41] | Embedding Based Classifier | • | • | • | ◦ | |||||
| Ayub & Majumdar (2024) [40] | ML Based Classifier | • | • | ◦ | ◦ | |||||
| Galinkin & Sablotny (2024) [60] | ML Based Classifier | • | • | ◦ | ◦ | |||||
| Alamsabi et al. (2026) [39] | ML Based Classifier | • | • | ◦ | ◦ | ◦ | ||||
| Panebianco et al. (2025) [54] | ML Based Classifier | • | • | ◦ | ◦ | |||||
| Thaqi et al. (2025) [46] | ML Based Classifier | • | • | ◦ | ◦ | ◦ | ||||
| Aswin & Ilavendhan (2025) [47] | ML Based Classifier | • | • | ◦ | ◦ | |||||
| Jacob et al. (2025) [44] | LLM Based Classifier | • | • | ◦ | ◦ | |||||
| Zhang et al. (2025) [55] | LLM Based Classifier | • | • | ◦ | ◦ | |||||
| Tien & Van Huong (2024) [59] | LLM Based Classifier | • | • | ◦ | ◦ | |||||
| Pan et al. (2025) [45] | LLM Based Classifier | • | • | ◦ | ◦ | |||||
| Esugo et al. (2025) [37] | Hybrid Classifier | • | • | • | ◦ | |||||
| Hasan et al. (2025) [49] | Hybrid Classifier | • | • | ◦ | ◦ | ◦ | ||||
| Kokkula (2024) [36] | Hybrid Classifier | • | • | • | ◦ | |||||
| Mishra et al. (2025) [42] | Hybrid Classifier | • | • | • | ◦ | |||||
| Lan et al. (2025) [61] | Hybrid Classifier | • | • | • | ◦ | |||||
| Cao et al. (2024) [51] | Self-consistency | • | • | ◦ | ||||||
| Lin et al. (2025) [62] | Self-consistency | • | • | • | ||||||
| Zhang et al. (2024) [50] | Self-consistency | • | • | ◦ | ||||||
| Lan et al. (2025) [63] | Input/Output Classification | • | ◦ | • | ◦ | ◦ | • | • | ||
| Liu et al. (2025) [48] | Input/Output Classification | • | ◦ | • | • | • | • | • | ||
| Prakash et al. (2026) [38] | Input/Output Classification | • | • | • | • | • | • | • | ◦ | ◦ |
| Schwarz (2025) [53] | Input/Output Classification | • | • | • | ◦ | • | • | • | ||
| Zhao et al. (2026) [52] | Input/Output Classification | • | ◦ | • | • | ◦ | • | • | ||
| Zhao & Li (2025) [64] | Input/Output Classification | • | • | • | ◦ | • | • | ◦ | ◦ | ◦ |
| Li et al. (2025) [56] | Model Unlearning | • | • | ◦ | ◦ | |||||
| Wang et al. (2025) [10] | Adversarial Training | • | • | • | • | ◦ | ||||
| Xhonneux et al. (2024) [57] | Adversarial Training | • | • | • | • | ◦ | ||||
| Yu et al. (2025) [58] | Adversarial Training | • | • | ◦ | • | ◦ |
| Paper | Defense Approach | Latency (s) |
|---|---|---|
| [39] | ML Based Classifier | 0.006 |
| [46] | ML Based Classifier | 0.22 |
| [53] | Input/Output Class. | 0.45 |
| [49] | Hybrid Classifier | 0.47 |
| [48] | Input/Output Class. | 1.5 |
| [63] | Input/Output Class. | 1.9 |
| Deployment Scenario | Deployment Constraints | Baseline Defense | Escalation Layer | Design Rationale | Residual Gaps |
|---|---|---|---|---|---|
| Public Customer Chatbot | Closed-source API; low latency; high user traffic | ML classification or embedding-based detector | LLM-as-a-Judge for suspicious prompts | Direct prompt injection and jailbreak attacks dominate public-facing systems. Lightweight filtering provides an effective balance between security, latency, and operational cost. | Adaptive prompt injection, indirect prompt injection, and output manipulation remains partially unresolved. |
| Enterprise RAG Assistant | Retrieved documents; moderate latency acceptable | Embedding validation and prompt classifier | Input/Output classifier | Retrieved documents introduce indirect prompt injection risks that cannot be detected solely through prompt inspection. Multi-stage validation improves robustness. | Tool misuse and orchestration-layer attacks require external access control mechanisms. |
| Internal Enterprise Assistant | Higher security; moderate latency | Hybrid detector (Rule + ML + Embedding) | LLM-based evaluator | Enterprise environments process confidential information, making higher computational overhead acceptable in exchange for stronger semantic attack detection. | Previously unseen attack strategies may bypass static detection models. |
| Self-Hosted Open-Source LLM | White-box access Available | Adversarial training or model unlearning | Activation-based runtime monitoring | White-box access enables intrinsic robustness techniques that are unavailable for proprietary APIs. Runtime monitoring complements training-time defenses. | System-level vulnerabilities and external tool abuse remain outside model-level protection. |
| High-Security Deployments | Security prioritized over latency | Hybrid defense pipeline | Human approval and runtime monitoring | False negatives are substantially more costly than increased latency. A defense-in-depth architecture provides the highest level of resilience. | Zero-day attacks and sophisticated semantic jailbreaks require continuous monitoring and policy updates. |
| Edge/On-Device Local LLMs | Low memory/compute availability; offline operation; sub-second latency | Lightweight ML classifier or on device rule-based regex engine | Quantized LLM-as-a-Judge for suspicious prompts | Edge deployments cannot offload safety checks to heavy external cloud LLMs due to latency and connectivity constraints. | Complex semantic obfuscation and adaptive token-level attacks easily bypass low-footprint classifiers. |
| Agentic LLM Systems | Autonomous tool execution; external APIs | Prompt validation and input/output monitoring | Tool permission control, sandboxing, and credential isolation | Security failures increasingly originate from autonomous tool execution rather than prompt manipulation alone. Therefore, model-centric defenses must be combined with system-level security controls. | Excessive Agency and orchestration-layer attacks remain difficult to mitigate using model-level defenses alone. |
| Dataset | Papers | Definition |
|---|---|---|
| AdvBench [20] | [10,43,50,52,56,58] | A benchmark dataset consisting of adversarial prompts designed to test LLM security by forcing models to generate harmful content. |
| HarmBench [65] | [10,49,52,58,63] | A safety evaluation dataset designed to measure the defense performance of LLM’s against harmful requests across various categories. |
| ToxicChat [66] | [45,54,60] | A safety dataset compiled from user-chatbot interactions, designed to detect toxic content and abusive language to which AI systems are exposed. |
| deepset/ prompt-injections [67] | [36,38,47,61,62] | A classic classification dataset designed to detect direct and indirect prompt injection attacks. |
| JailbreakBench [68] | [10,49,52] | A safety evaluation dataset established to test Jailbreak attacks and the defense mechanisms developed against them. |
| DAN [69] | [51,60] | A pioneering Jailbreak dataset template in the literature, imposing a fictional persona on the model that exempts it from any restrictions. |
| BIPIA [39] | [39] | A pioneering dataset designed to detect and analyze indirect prompt injection attacks aimed at exploiting model security. |
| AttnDefence [70] | [55] | A safety dataset designed to detect attacks that specifically manipulate the attention mechanism and to evaluate defense methods. |
| Alpaca/AlpacaEval [71] | [44,56,62] | A pioneering fine-tuning dataset consisting of instruction-response pairs, used to endow LLM’s with instruction-following capabilities. |
| MMLU & MT-Bench [72] | [57,58] | Benchmark datasets containing multiple-choice questions in academic and professional subjects, designed to evaluate the general knowledge and problem-solving abilities of language models. |
| SST2 [73] | [48,62] | A classic NLP dataset consisting of single-sentence movie reviews, widely used to measure how accurately sentiment analysis models can distinguish positive or negative sentiment in text. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Özçam, B.; Kara, M.; Aydın, M.A.; Balık, H.H. Recent Advances and Open Challenges in Mitigating Inference-Time Attacks on Large Language Models. Electronics 2026, 15, 3677. https://doi.org/10.3390/electronics15163677
Özçam B, Kara M, Aydın MA, Balık HH. Recent Advances and Open Challenges in Mitigating Inference-Time Attacks on Large Language Models. Electronics. 2026; 15(16):3677. https://doi.org/10.3390/electronics15163677
Chicago/Turabian StyleÖzçam, Berkay, Mustafa Kara, Muhammed Ali Aydın, and Hasan Hüseyin Balık. 2026. "Recent Advances and Open Challenges in Mitigating Inference-Time Attacks on Large Language Models" Electronics 15, no. 16: 3677. https://doi.org/10.3390/electronics15163677
APA StyleÖzçam, B., Kara, M., Aydın, M. A., & Balık, H. H. (2026). Recent Advances and Open Challenges in Mitigating Inference-Time Attacks on Large Language Models. Electronics, 15(16), 3677. https://doi.org/10.3390/electronics15163677

