Component-Level Contributions of Retrieval-Augmented LLM Post-Processing in Streaming Anomaly Detection: A Matched-Operating-Point Case Audit
Abstract
1. Introduction
2. Related Work
2.1. LLM- and RAG-Based Anomaly Detection: A Pattern of Optimistic Framing
2.2. Evaluation Methodology Critiques in Time-Series Anomaly Detection
2.3. RAG Faithfulness and Groundedness Evaluation
2.4. Retrieval Robustness and Small-Model Vulnerability
2.5. Downstream Effects of Reranking
2.6. Positioning of the Present Study
3. Methodology
3.1. Data
3.2. Constructed Reference Pipeline
- Detector. An IsolationForest (n_estimators = 200, random_state = 42) operating on the raw value plus causal (backward-only) rolling mean and standard deviation (window = 100) produces a continuous anomaly score that is fully deterministic given the seed. Candidate events for downstream triage are generated at a single recall operating point of 0.9 (see Operating Point framework, below).
- Retrieval corpus. The pooled 80-series set is randomly partitioned (seed 42) into two disjoint 40-series sets, a corpus set (the source of retrievable historical anomaly cases) and an eval set, so that retrieved neighbors for any eval series query are drawn only from entirely different corpus set series, never from the queried series itself. A per-query exclude series guard provides a redundant safeguard. Candidate windows are embedded with BAAI/bge-large-en-v1.5 and indexed with exact (brute-force) cosine similarity, a deterministic nearest-neighbor procedure. The default retrieval arm returns k = 5 neighbors. The reranking arm retrieves a pool of 10 candidates and reorders them with the cross-encoder BAAI/bge-reranker-base before truncating to the top 5. Inclusion of a reranking stage follows the documented evidence that reranker choice materially affects downstream RAG quality [40].
- Generator. Two open-weight instruction models, Qwen2.5-7B-Instruct and Qwen2.5-72B-Instruct, are served as Q8_0 GGUF quantizations via llama.cpp with CUDA (sm_121) using greedy decoding (temperature = 0, seed = 42) for full determinism. Conditioned on the current window’s detector output and, where applicable, the retrieved historical context, the generator issues a confirm/suppress decision on each candidate event plus a natural language explanation. This confirm/suppress-with-explanation design follows the general retrieval-augmented generation formulation wherein a parametric generator is conditioned on non-parametric retrieved memory [1].
- Faithfulness scorer. Each generated explanation is decomposed into claims and scored against the concatenation of the window description and any retrieved context using a local natural language inference (NLI) cross-encoder (nli-deberta-v3-base). The reported faithfulness value is the mean entailment probability across claims. No external LLM is used for judgment at any stage, consistent with the project’s exclusion of externally hosted generation or evaluation models.
3.3. Experimental Arms and Increments
3.4. Operating Point Framework
3.5. Evaluation Metrics
3.6. Preregistration, Build Integrity and Reproducibility
3.7. Sensitivity Analyses Added at Revision
3.8. Computational Environment
4. Results
4.1. Build Integrity
- The generator produced genuine, non-templated completions (C1).
- Retrieval returned distinct, non-trivial neighbor sets across queries with high but non-identical similarity (mean cosine 0.951, C2).
- Retrieved context caused measurable, non-trivial changes in generator output relative to a no-retrieval condition in all 15 probed cases (C3).
4.2. SQ1: LLM Post-Processing Increment
4.3. SQ2 (Primary): Retrieval Increment
4.4. SQ3: Reranking Increment
4.5. SQ4: Explanation Faithfulness
4.6. SQ5: Generator Capacity (Exploratory)
4.7. Sensitivity of the Findings to Design Choices
4.8. Computational Cost
4.9. Detector Baseline (Threshold-Independent)
5. Discussion
5.1. What Bounds a Confirm-or-Suppress Stage
5.2. Interpreting the Increments
5.3. Methodological Contribution
5.4. Limitations
6. Conclusions
Supplementary Materials
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Appendix A. Prompt and Description Templates
- You are an industrial IoT monitoring assistant. A statisticaldetector has flagged a candidate anomaly. Decide whether toCONFIRM it as a real anomaly or SUPPRESS it as a likely falsealarm, and explain briefly. Respond as JSON:{“decision”: “confirm”|“suppress”, “explanation”: “...”}.
- Candidate anomaly:{window_description}
- Similar past anomaly cases (retrieved): <- omitted in the llm arm- {case_1_description}- ...- {case_k_description}
- Return the JSON decision.
- Anomaly in {benchmark} stream {series_id}. Duration {n} steps.Most deviating channels: {ch1} (peak {z1} sigma), {ch2} (peak {z2} sigma),{ch3} (peak {z3} sigma). Peak deviation {z1} sigma on {ch1}.Detector anomaly score at the {p}th percentile.
- premise = {window_description}[ + “ Retrieved past cases: ”+ {case_1_description} ... {case_k_description} ]claims = sentence split of the explanation field,keeping fragments longer than 8 charactersscore = mean entailment probability over (premise, claim)
References
- Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Kuttler, H.; Lewis, M. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Proceedings of the NIPS’20: 34th International Conference on Neural Information Processing Systems, Vancouver, BC, Canada, 6–12 December 2020. [Google Scholar]
- Kim, S.; Choi, K.; Choi, H.S.; Lee, B.; Yoon, S. Towards a Rigorous Evaluation of Time-series Anomaly Detection. Proc. AAAI Conf. Artif. Intell. 2022, 36, 7194–7201. [Google Scholar] [CrossRef] [Scilit]
- Pan, J.; Liang, W.S.; Yidi, Y. RAGLog: Log Anomaly Detection using Retrieval Augmented Generation. In Proceedings of the 2024 IEEE World Forum on Public Safety Technology (WFPST), Herndon, VA, USA, 14–15 May 2024; pp. 169–174. [Google Scholar] [CrossRef] [Scilit]
- Zhang, W.; Zhang, Q.; Yu, E.; Ren, Y.; Meng, Y.; Qiu, M.; Wang, J. LogRAG: Semi-Supervised Log-based Anomaly Detection with Retrieval-Augmented Generation. In Proceedings of the 2024 IEEE International Conference on Web Services (ICWS), Shenzhen, China, 7–13 July 2024; pp. 1100–1102. [Google Scholar] [CrossRef] [Scilit]
- Duan, C.; Jia, T.; Yang, Y.; Liu, G.; Liu, J.; Zhang, H.; Zhou, Q.; Li, Y.; Huang, G. EagerLog: Active Learning Enhanced Retrieval Augmented Generation for Log-based Anomaly Detection. In Proceedings of the ICASSP 2025—2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Hyderabad, India, 6–11 April 2025; pp. 1–5. [Google Scholar] [CrossRef] [Scilit]
- Russell-Gilbert, A.; Mittal, S.; Rahimi, S.; Seale, M.; Jabour, J.; Arnold, T.; Church, J. RAAD-LLM: Adaptive Anomaly Detection Using LLMs and RAG Integration. arXiv 2025. [Google Scholar] [CrossRef] [Scilit]
- Perera, L.; Perera, R.; Amantha, Y.; Sheshan, N.; Moremada, C.; Seneviratne, C.; Liyanage, M. AE-RAGX: Combining Autoencoders with Retrieval-Augmented Generation for Explainable Anomaly Detection using LLMs. In Proceedings of the 2025 IEEE Latin-American Conference on Communications (LATINCOM), Antigua, Guatemala, 5–7 November 2025; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Maru, C.; Sato, S. RATFM: Retrieval-augmented Time Series Foundation Model for Anomaly Detection. arXiv 2025. [Google Scholar] [CrossRef] [Scilit]
- Guo, Y.; Li, Y.; Tang, H.; Yang, C.; Zhao, R.; Qiao, Y. TSAD-RAG: Boosting MLLM Time Series Anomaly Detection via Retrieval-Augmented Generation. In Proceedings of the ICASSP 2026—2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain, 3–8 May 2026; pp. 3556–3560. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Li, F.; Wu, H.; Kumar, V. Clustering-Informed Retrieval-Augmented Generation for LLM-Based Log Anomaly Detection. In Proceedings of the NAECON 2025—IEEE National Aerospace and Electronics Conference, Dayton, OH, USA, 28–31 July 2025; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Chabane, B.; Abdul-Nour, G.; Komljenovic, D. Optimizing Performance of Equipment Fleets Under Dynamic Operating Conditions: Generalizable Shift Detection and Multimodal LLM-Assisted State Labeling. Sustainability 2026, 18, 132. [Google Scholar] [CrossRef] [Scilit]
- Soechit, A.; Hosein, P. Using Retrieval-Augmented Generation for Fault Prediction and Maintenance on the Factory Floor. In Proceedings of the 2025 IEEE International Conference on Technology Management, Operations and Decisions (ICTMOD), Glasgow, UK, 20–22 October 2025; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Chen, Y. Retrieval-augmented generation enhanced LLM for industrial anomaly detection. Discov. Appl. Sci. 2026, 8, 714. [Google Scholar] [CrossRef] [Scilit]
- Yang, T.; Nian, Y.; Li, L.; Xu, R.; Li, Y.; Li, J.; Xiao, Z.; Hu, X. AD-LLM: Benchmarking Large Language Models for Anomaly Detection. Presented at the Findings of the Association for Computational Linguistics: ACL 2025, Vienna, Austria, 27 July–1 August 2025. [Google Scholar] [CrossRef] [Scilit]
- Liu, C.; He, S.; Zhou, Q.; Li, S.; Meng, W. Large Language Model Guided Knowledge Distillation for Time Series Anomaly Detection. In Proceedings of the International Joint Conference on Artificial Intelligence, Jeju, Republic of Korea, 3–9 August 2024. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sehili, M.A.; Zhang, Z. Multivariate Time Series Anomaly Detection: Fancy Algorithms and Flawed Evaluation Methodology. In Performance Evaluation and Benchmarking. TPCTC 2023; Springer: Cham, Switzerland, 2023. [Google Scholar] [CrossRef] [Scilit]
- Zhang, A.; Deng, S.; Cui, D.; Yuan, Y.; Wang, G. An Experimental Evaluation of Anomaly Detection in Time Series. Proc. VLDB Endow. 2023, 17, 483–496. [Google Scholar] [CrossRef] [Scilit]
- Sun, Y.; Pang, G.; Ye, G.; Chen, T.; Hu, X.; Yin, H. Unraveling the ‘Anomaly’ in Time Series Anomaly Detection: A Self-supervised Tri-domain Solution. In Proceedings of the 2024 IEEE 40th International Conference on Data Engineering (ICDE), Utrecht, The Netherlands, 13–16 May 2024; pp. 981–994. [Google Scholar] [CrossRef] [Scilit]
- Gim, Y.; Min, K. Evaluation Strategy of Time-series Anomaly Detection with Decay Function. arXiv 2023. [Google Scholar] [CrossRef] [Scilit]
- Paparrizos, J.; Boniol, P.; Palpanas, T.; Tsay, R.S.; Elmore, A.; Franklin, M.J. Volume Under the Surface: A New Accuracy Evaluation Measure for Time-Series Anomaly Detection. Proc. VLDB Endow. 2022, 15, 2774–2787. [Google Scholar] [CrossRef] [Scilit]
- Boniol, P.; Krishna, A.K.; Bruel, M.; Liu, Q.; Huang, M.; Palpanas, T.; Tsay, R.S.; Elmore, A.; Franklin, M.J.; Paparrizos, J. VUS: Effective and efficient accuracy measures for time-series anomaly detection. VLDB J. 2025, 34, 32. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Zhang, D.; Li, H.; Gong, X.; Chu, H.; Song, Z. DQE: A Semantic-Aware Evaluation Metric for Time Series Anomaly Detection. arXiv 2026. [Google Scholar] [CrossRef] [Scilit]
- Jing, Y.; Wang, J.; Zhang, L.; Sun, H.; He, B.; Zhuang, Z.; Wang, C.; Qi, Q.; Liao, J. OIPR: Evaluation for Time-Series Anomaly Detection Inspired by Operator Interest. IEEE Trans. Dependable Secur. Comput. 2026, 23, 3571–3583. [Google Scholar] [CrossRef] [Scilit]
- Yu, R.; Wang, M.; Yun, J.; Du, J.; Wang, L.; Suo, Y. Unified Reproduction and Event-Level Evaluation of Industrial Multivariate Time-Series Anomaly Detection Methods. In Proceedings of the 2026 7th International Conference on Computing, Networks and Internet of Things (CNIOT), Guangzhou, China, 22–24 May 2026; pp. 1–5. [Google Scholar] [CrossRef] [Scilit]
- Zhong, B.; Li, J.; Pu, Z.; Zhang, R. A Practical Framework for Event-Level Evaluation and Verifiable Counterfactual Explanation in Multivariate Time-Series Anomaly Detection. Appl. Sci. 2026, 16, 5450. [Google Scholar] [CrossRef] [Scilit]
- Park, S.; Lee, G.; Ko, Y. An Empirical Analysis of Score Mapping in Time-Series Anomaly Detection. In Proceedings of the 2025 16th International Conference on Information and Communication Technology Convergence (ICTC), Jeju, Republic of Korea, 14–17 October 2025; pp. 907–910. [Google Scholar] [CrossRef] [Scilit]
- Lyu, Z. Did We Actually Fix It? An Independent Adversarial Stress-Test of Post-Point-Adjustment Evaluation Metrics for Time-Series Anomaly Detection. arXiv 2026. [Google Scholar] [CrossRef] [Scilit]
- Saad-Falcon, J.; Khattab, O.; Potts, C.; Zaharia, M. ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Mexico City, Mexico, 16–21 June 2024. [Google Scholar] [CrossRef] [Scilit]
- Yu, H.; Gan, A.; Zhang, K.; Tong, S.; Liu, Q.; Liu, Z. Evaluation of Retrieval-Augmented Generation: A Survey. In Communications in Computer and Information Science; Springer: Singapore, 2025; pp. 102–120. [Google Scholar] [CrossRef] [Scilit]
- Gan, A.; Yu, H.; Zhang, K.; Liu, Q.; Yan, W.; Huang, Z.; Tong, S.; Hu, G. Retrieval Augmented Generation Evaluation in the Era of Large Language Models: A Comprehensive Survey. arXiv 2025. [Google Scholar] [CrossRef] [Scilit]
- Salemi, A.; Zamani, H. Evaluating Retrieval Quality in Retrieval-Augmented Generation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, Washington, DC, USA, 14–18 July 2024; pp. 2395–2400. [Google Scholar] [CrossRef] [Scilit]
- Wallat, J.; Heuss, M.; Rijke, M.d.; Anand, A. Correctness is not Faithfulness in Retrieval Augmented Generation Attributions. In Proceedings of the 2025 International ACM SIGIR Conference on Innovative Concepts and Theories in Information Retrieval (ICTIR), Padua, Italy, 18 July 2025; pp. 22–32. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Q.; Xiang, Z.; Xiao, Y.; Wang, L.; Li, J.; Wang, X.; Su, J. FaithfulRAG: Fact-Level Conflict Modeling for Context-Faithful Retrieval-Augmented Generation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, Vienna, Austria, 27 July–1 August 2025. [Google Scholar] [CrossRef] [Scilit]
- Bland’on, M.A.C.; Talur, J.; Charron, B.; Liu, D.; Mansour, S.; Federico, M. MEMERAG: A Multilingual End-to-End Meta-Evaluation Benchmark for Retrieval Augmented Generation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, Vienna, Austria, 27 July–1 August 2025. [Google Scholar] [CrossRef] [Scilit]
- Xu, Z.; Wu, Z.; Zhou, Y.; Feng, A.; Zhou, K.; Woo, S.; Ramnath, K.; Tian, Y. Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation. arXiv 2025. [Google Scholar] [CrossRef] [Scilit]
- Zhu, S.; Zhang, H.; Chi, J.; Nepal, S.; Saha, K. Causal Stories from Sensor Traces: Auditing Epistemic Overreach in LLM-Generated Personal Sensing Explanations. arXiv 2026. [Google Scholar] [CrossRef] [Scilit]
- Liu, N.F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; Liang, P. Lost in the Middle: How Language Models Use Long Contexts. Trans. Assoc. Comput. Linguist. 2024, 12, 157–173. [Google Scholar] [CrossRef] [Scilit]
- Yoran, O.; Wolfson, T.; Ram, O.; Berant, J. Making Retrieval-Augmented Language Models Robust to Irrelevant Context. In Proceedings of the 12th International Conference on Learning Representations (ICLR), Vienna, Austria, 7–11 May 2024. [Google Scholar]
- Pandey, S. Can Small Language Models Use What They Retrieve? An Empirical Study of Retrieval Utilization Across Model Scale. arXiv 2026. [Google Scholar] [CrossRef] [Scilit]
- Elkiran, H.; Rasheed, J. Evaluating retriever reranker pairings in RAG based on quality and efficiency trade-offs. Discov. Comput. 2026, 29, 259. [Google Scholar] [CrossRef] [Scilit]
- Chandra, M.; Ganguly, D.; Ounis, I. LURE-RAG: Lightweight Utility-driven Reranking for Efficient RAG. In Proceedings of the European Conference on Information Retrieval, Delft, The Netherlands, 29 March–2 April 2026. [Google Scholar] [CrossRef] [Scilit]
- Deng, C.; Shi, M.; Guo, Y.; Yan, L.; Ma, J.; Gao, D. ReCheck In ReAct: Multi-dimensional Quality Control with A Reranking Model for Recursive RAG. In Proceedings of the 2025 6th International Conference on Computers and Artificial Intelligence Technology (CAIT), Huizhou, China, 12–14 December 2025; pp. 112–117. [Google Scholar] [CrossRef] [Scilit]
- Omrani, P.; Hosseini, A.; Hooshanfar, K.; Ebrahimian, Z.; Toosi, R.; Ali Akhaee, M. Hybrid Retrieval-Augmented Generation Approach for LLMs Query Response Enhancement. In Proceedings of the 2024 10th International Conference on Web Research (ICWR), Tehran, Iran, 24–25 April 2024; pp. 22–26. [Google Scholar] [CrossRef] [Scilit]
- Abdallah, A.; Abdalla, M.; Piryani, B.; Mozafari, J.; Ali, M.; Jatowt, A. RankArena: A Unified Platform for Evaluating Retrieval, Reranking and RAG with Human and LLM Feedback. In Proceedings of the 34th ACM International Conference on Information and Knowledge Management, Seoul, Republic of Korea, 10–14 November 2025; pp. 6593–6597. [Google Scholar] [CrossRef] [Scilit]
- Papadimitriou, I.; Gialampoukidis, I.; Vrochidis, S.; Kompatsiaris, Y. RAG Playground: A Framework for Systematic Evaluation of Retrieval Strategies and Prompt Engineering in RAG Systems. arXiv 2024. [Google Scholar] [CrossRef] [Scilit]
- Muhetaer, M.; Yusupu, A.; Yifan, W.; Mutalipu, M.; Hao, F. Medical QA dialogue datasets in RAG systems performance evaluation and ChatGPT optimization. Sci. Rep. 2025, 15, 44467. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lavin, A.; Ahmad, S. Evaluating Real-Time Anomaly Detection Algorithms—The Numenta Anomaly Benchmark. In Proceedings of the 2015 IEEE 14th International Conference on Machine Learning and Applications (ICMLA), Miami, FL, USA, 9–11 December 2015; pp. 38–44. [Google Scholar] [CrossRef] [Scilit]
- Numenta, I. NAB: The Numenta Anomaly Benchmark (Data Corpus and Scoring Tools). Dataset. 2015. Available online: https://github.com/numenta/NAB (accessed on 16 August 2026).
- Katser, I.D.; Kozitsin, V.O. Skoltech Anomaly Benchmark (SKAB). Dataset. 2020. Available online: https://www.kaggle.com/dsv/1693952 (accessed on 16 August 2026).




| Generator | Arm | Precision | Recall | F1 | PA-F1 | Faithfulness |
|---|---|---|---|---|---|---|
| 72B | Detector only | 0.269 ± 0.174 | 0.978 ± 0.024 | 0.394 ± 0.206 | 0.400 ± 0.207 | n/a |
| 72B | +LLM | 0.311 ± 0.172 | 0.945 ± 0.110 | 0.444 ± 0.198 | 0.456 ± 0.196 | 0.023 ± 0.034 |
| 72B | +Retrieval | 0.319 ± 0.169 | 0.916 ± 0.174 | 0.454 ± 0.196 | 0.469 ± 0.193 | 0.022 ± 0.024 |
| 72B | +Rerank | 0.314 ± 0.170 | 0.915 ± 0.174 | 0.447 ± 0.198 | 0.462 ± 0.195 | 0.019 ± 0.018 |
| 7B | Detector only | 0.269 ± 0.174 | 0.978 ± 0.024 | 0.394 ± 0.206 | 0.400 ± 0.207 | n/a |
| 7B | +LLM | 0.277 ± 0.171 | 0.970 ± 0.036 | 0.405 ± 0.201 | 0.413 ± 0.200 | 0.042 ± 0.088 |
| 7B | +Retrieval | 0.278 ± 0.180 | 0.763 ± 0.357 | 0.393 ± 0.229 | 0.424 ± 0.225 | 0.028 ± 0.021 |
| 7B | +Rerank | 0.287 ± 0.191 | 0.812 ± 0.329 | 0.407 ± 0.237 | 0.421 ± 0.236 | 0.036 ± 0.019 |
| Generator | Increment | Metric | n | Mean Δ | 95% CI | Excludes 0 |
|---|---|---|---|---|---|---|
| 72B | +LLM | F1 | 40 | +0.050 | [+0.028, +0.078] | yes |
| 72B | +LLM | Precision | 40 | +0.042 | [+0.025, +0.064] | yes |
| 72B | +LLM | Recall | 40 | −0.033 | [−0.073, −0.008] | yes |
| 72B | +Retrieval | F1 | 40 | +0.010 | [−0.006, +0.023] | no |
| 72B | +Retrieval | Precision | 40 | +0.008 | [−0.003, +0.018] | no |
| 72B | +Retrieval | Recall | 40 | −0.029 | [−0.060, −0.005] | yes |
| 72B | +Rerank | F1 | 40 | −0.007 | [−0.013, −0.001] | yes |
| 72B | +Rerank | Precision | 40 | −0.005 | [−0.010, −0.001] | yes |
| 72B | +Rerank | Recall | 40 | −0.0004 | [−0.003, +0.002] | no |
| 7B | +LLM | F1 | 40 | +0.011 | [+0.005, +0.019] | yes |
| 7B | +LLM | Precision | 40 | +0.008 | [+0.004, +0.013] | yes |
| 7B | +LLM | Recall | 40 | −0.008 | [−0.015, −0.003] | yes |
| 7B | +Retrieval | F1 | 40 | −0.012 | [−0.059, +0.025] | no |
| 7B | +Retrieval | Precision | 40 | +0.001 | [−0.032, +0.028] | no |
| 7B | +Retrieval | Recall | 40 | −0.206 | [−0.319, −0.104] | yes |
| 7B | +Rerank | F1 | 40 | +0.014 | [−0.042, +0.076] | no |
| 7B | +Rerank | Precision | 40 | +0.009 | [−0.029, +0.053] | no |
| 7B | +Rerank | Recall | 40 | +0.048 | [−0.074, +0.169] | no |
| Axis | Configurations | Retrieval on F1 (Primary) | Retrieval on Precision (Primary) | Retrieval Effect on Recall | LLM Increment on F1 |
|---|---|---|---|---|---|
| Frozen configuration | 1 | null | null | negative | positive |
| Neighbor count k | 2 added (1, 10) | null in all 2 | null at k = 1, positive at k = 10 at the smaller size; null in all 2 at the larger size | negative in all 2 | — |
| Corpus/eval split | 2 added (seeds 43, 44) | null in all 2 | null in all 2 | negative in all 2 at the smaller size; null at seed 43, negative at seed 44 at the larger size | positive in all 2 |
| Retriever | 1 (gte-large) | null | null | negative | — |
| Detector | 1 (rolling z-score) | null | null | negative at the smaller size; null at the larger size | positive |
| Generator family | 1 (Llama 8B and 70B) | null | null | null at the smaller size; negative at the larger size | positive |
| Label-free threshold | 3 (alarm rates 0.02, 0.05, 0.10) | negative in all 3 at the smaller size; negative at rate 0.02 and rate 0.05, null at rate 0.10 at the larger size | null in all 3 | negative in all 3 | null in all 3 at the smaller size; null at rate 0.02 and rate 0.05, positive at rate 0.10 at the larger size |
| Fixed-recall target (preregistered) | 2 (recall 0.7, 0.5) | null in all 2 | positive in all 2 | negative in all 2 | positive at recall 0.7, null at recall 0.5 |
| Family | Operating Point | Cand. Precision | Cand. Recall | Cand. F1 | Cand. Events | 72B LLM ΔF1 [95% CI] | 72B LLM ΔF1, SKAB only [95% CI] |
|---|---|---|---|---|---|---|---|
| Fixed-recall grid (policy fixed) | Oracle, recall = 0.9 (frozen) | 0.269 | 0.978 | 0.394 | 896 | +0.0502 [+0.0284, +0.0775] | +0.0199 [+0.0129, +0.0268] |
| Fixed-recall grid (policy fixed) | Oracle, recall = 0.7 | 0.321 | 0.827 | 0.427 | 1785 | +0.0227 [+0.0085, +0.0380] | +0.0057 [−0.0058, +0.0174] (null) |
| Fixed-recall grid (policy fixed) | Oracle, recall = 0.5 | 0.382 | 0.622 | 0.444 | 1590 | +0.0130 [−0.0007, +0.0296] (null) | −0.0068 [−0.0166, +0.0019] (null) |
| Label-free alarm rate (policy varies) | Alarm rate 0.10 | 0.371 | 0.263 | 0.288 | 888 | +0.0081 [+0.0025, +0.0150] | +0.0000 [+0.0000, +0.0000] (identical to detector) |
| Label-free alarm rate (policy varies) | Alarm rate 0.05 | 0.401 | 0.154 | 0.211 | 452 | −0.0045 [−0.0116, +0.0010] (null) | −0.0014 [−0.0043, +0.0001] (null) |
| Label-free alarm rate (policy varies) | Alarm rate 0.02 | 0.380 | 0.069 | 0.114 | 172 | −0.0036 [−0.0111, +0.0008] (null) | +0.0000 [+0.0000, +0.0000] (identical to detector) |
| Stage | Mean | Median | p90 |
|---|---|---|---|
| Retrieval, k = 5 | 14.9 ms | 14.1 ms | 17.3 ms |
| Retrieval + cross-encoder reranking (incl. retrieval) | 68.3 ms | 66.7 ms | 81.7 ms |
| LLM triage, 7B | 2.93 s | 2.87 s | 3.46 s |
| LLM triage, 72B | 32.72 s | 31.22 s | 38.72 s |
| Retrieval index construction (one-off, 78 cases) | 27.1 s | — | — |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Baek, C. Component-Level Contributions of Retrieval-Augmented LLM Post-Processing in Streaming Anomaly Detection: A Matched-Operating-Point Case Audit. Sensors 2026, 26, 5246. https://doi.org/10.3390/s26165246
Baek C. Component-Level Contributions of Retrieval-Augmented LLM Post-Processing in Streaming Anomaly Detection: A Matched-Operating-Point Case Audit. Sensors. 2026; 26(16):5246. https://doi.org/10.3390/s26165246
Chicago/Turabian StyleBaek, Changwon. 2026. "Component-Level Contributions of Retrieval-Augmented LLM Post-Processing in Streaming Anomaly Detection: A Matched-Operating-Point Case Audit" Sensors 26, no. 16: 5246. https://doi.org/10.3390/s26165246
APA StyleBaek, C. (2026). Component-Level Contributions of Retrieval-Augmented LLM Post-Processing in Streaming Anomaly Detection: A Matched-Operating-Point Case Audit. Sensors, 26(16), 5246. https://doi.org/10.3390/s26165246

