Anomalies in AI Outputs Beyond Input Data Quality: The Significance of Reasoning
Abstract
1. Introduction
2. Materials and Methods
2.1. Systematic Review Protocol
2.2. Theoretical Framework
3. Results
3.1. Data Sources and Datasets
- simpleqa_hallucination_rates.csv (7 models, 4 variables): hallucination rates from the SimpleQA benchmark [27], extracted from the official system cards of OpenAI, Google, and Anthropic (2024–2026). Used to quantify the magnitude of the hallucination problem in state-of-the-art models (Section 3.2).
- knowledge_vs_reasoning.csv (7 models, 4 variables): accuracy decomposed into knowledge vs. reasoning according to the protocol by Thapa et al. [13], based on 11 biomedical benchmarks. Used to demonstrate the independence of reasoning anomalies (Section 3.2).
- irrelevant_context_impact.csv (5 conditions, 3 variables): effect of irrelevant context based on GSM-IC [11] and GSM8K [28]. Used to demonstrate that correct data does not guarantee correct reasoning (Section 3.2).
- mitigation_effectiveness.csv (6 techniques, 5 variables): Comparative effectiveness of mitigation techniques, consolidating results from Self-Consistency [29], CoVe [30], Tree of Thoughts [31], Process Reward Models [32], SATLM/VERGE [33,34], and multi-agent debate [35]. Used to evaluate existing interventions (Section 3.4).
- cot_faithfulness.csv (7 families, 4 variables): Chain-of-Thought faithfulness by model family according to Young et al. [20]. Used to analyze the reliability of reasoning traces (Section 3.5).
- benchmark_saturation.csv (10 benchmarks, 5 variables): saturation analysis, consolidating data from public leaderboards and original papers. Used to evaluate the validity of current benchmarks (Section 3.6).
3.2. Taxonomy of Reasoning Anomalies
- There is no unified framework that integrates the detection of factual hallucinations, reasoning failures, and CoT inaccuracies.
- The connection between functional awareness mechanisms and the detection of reasoning anomalies lacks concrete operational translation.
- The distinction between knowledge errors and reasoning errors requires standardized evaluation frameworks.
3.3. Detection Methods for Reasoning Anomalies
3.3.1. Self-Consistency Verification
3.3.2. Semantic Entropy
3.3.3. Internal State Probing
3.3.4. Formal Verification
3.4. Mitigation Techniques
3.5. Chain-of-Thought Faithfulness Analysis
3.6. Benchmark Saturation and New Metrics
3.7. Proposed Framework Architecture
4. Discussion
- Computational overhead, although Semantic Entropy Probes [46] reduce it.
- Formal verification is limited to formalizable domains.
- The framework does not address phenomenal consciousness [26]. In particular, Layer 4 formal verification introduces additional latency in inference time that has not been quantified in this work. The activation of SAT/SMT solvers and multi-agent debate increases computational cost by a non-negligible amount. Future implementations must measure this overhead in milliseconds per query and establish selective activation thresholds to ensure the framework’s viability in commercial deployments with latency constraints.
Limitations and Scope of Validity
5. Conclusions
Supplementary Materials
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Zhao, W.X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z.; et al. A Survey of Large Language Models. arXiv 2026, arXiv:2303.18223. [Google Scholar] [CrossRef]
- Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F.L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. GPT-4 Technical Report. arXiv 2024, arXiv:2303.08774. [Google Scholar] [CrossRef]
- Huang, L.; Yu, W.; Ma, W.; Zhong, W.; Feng, Z.; Wang, H.; Chen, Q.; Peng, W.; Feng, X.; Qin, B.; et al. A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions. ACM Trans. Inf. Syst. 2025, 43, 42. [Google Scholar] [CrossRef]
- Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y.J.; Madotto, A.; Fung, P. Survey of Hallucination in Natural Language Generation. ACM Comput. Surv. 2023, 55, 248. [Google Scholar] [CrossRef]
- Gao, Y.; Xiong, Y.; Gao, X.; Jia, K.; Pan, J.; Bi, Y.; Dai, Y.; Sun, J.; Wang, M.; Wang, H. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv 2024, arXiv:2312.10997. [Google Scholar] [CrossRef]
- Liu, N.F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; Liang, P. Lost in the Middle: How Language Models Use Long Contexts. Trans. Assoc. Comput. Linguist. 2024, 12, 157–173. [Google Scholar] [CrossRef]
- Niu, C.; Wu, Y.; Zhu, J.; Xu, S.; Shum, K.; Zhong, R.; Song, J.; Zhang, T. RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Bangkok, Thailand, 2024; pp. 10862–10878. [Google Scholar]
- Zou, W.; Geng, R.; Wang, B.; Jia, J. PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models. In Proceedings of the 34th USENIX Security Symposium; USENIX Association: Seattle, WA, USA, 2025; pp. 3827–3844. [Google Scholar]
- Asai, A.; Wu, Z.; Wang, Y.; Sil, A.; Hajishirzi, H. Self-RAG: Learning to Retrieve, Generate, and Critique Through Self-Reflection. In Proceedings of the 12th International Conference on Learning Representations (ICLR 2024), Vienna, Austria, 7–11 May 2024. [Google Scholar]
- Yan, S.-Q.; Gu, J.-C.; Zhu, Y.; Ling, Z.-H. Corrective Retrieval Augmented Generation. arXiv 2024, arXiv:2401.15884. [Google Scholar] [CrossRef]
- Shi, F.; Chen, X.; Misra, K.; Scales, N.; Dohan, D.; Chi, E.; Schärli, N.; Zhou, D. Large Language Models Can Be Easily Distracted by Irrelevant Context. In Proceedings of the 40th International Conference on Machine Learning; JMLR: Honolulu, HI, USA, 2023; pp. 31210–31227. [Google Scholar]
- Yang, M.; Huang, E.; Zhang, L.; Surdeanu, M.; Wang, W.Y.; Pan, L. How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing; Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 13329–13347. [Google Scholar]
- Thapa, R.; Wu, Q.; Wu, K.; Zhang, H.; Zhang, A.; Wu, E.; Ye, H.; Bedi, S.; Aresh, N.; Boen, J.; et al. Disentangling Reasoning and Knowledge in Medical Large Language Models. arXiv 2025, arXiv:2505.11462. [Google Scholar] [CrossRef]
- Jin, M.; Luo, W.; Cheng, S.; Wang, X.; Hua, W.; Tang, R.; Wang, W.Y.; Zhang, Y. Disentangling Memory and Reasoning Ability in Large Language Models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 1681–1701. [Google Scholar]
- Song, P.; Han, P.; Goodman, N. Large Language Model Reasoning Failures. arXiv 2026, arXiv:2602.06176. [Google Scholar] [CrossRef]
- Tyen, G.; Mansoor, H.; Carbune, V.; Chen, P.; Mak, T. LLMs Cannot Find Reasoning Errors, But Can Correct Them Given the Error Location. In Proceedings of the Findings of the Association for Computational Linguistics ACL 2024; Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; pp. 13894–13908. [Google Scholar]
- Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.; Le, Q.; Zhou, D. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Proceedings of the 36th International Conference on Neural Information Processing System; Curran Associates Inc.: New Orleans, LA, USA, 2022; pp. 24824–24837. [Google Scholar]
- Lanham, T.; Chen, A.; Radhakrishnan, A.; Steiner, B.; Denison, C.; Hernandez, D.; Li, D.; Durmus, E.; Hubinger, E.; Kernion, J.; et al. Measuring Faithfulness in Chain-of-Thought Reasoning. arXiv 2023, arXiv:2307.13702. [Google Scholar] [CrossRef]
- Turpin, M.; Michael, J.; Perez, E.; Bowman, S.R. Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting. In Proceedings of the 37th International Conference on Neural Information Processing Systems; Curran Associates Inc.: New Orleans, LA, USA, 2023; pp. 74952–74965. [Google Scholar]
- Young, R.J. Lie to Me: How Faithful Is Chain-of-Thought Reasoning in Reasoning Models? arXiv 2026, arXiv:2603.22582. [Google Scholar] [CrossRef]
- Paul, D.; West, R.; Bosselut, A.; Faltings, B. Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning. In EMNLP 2024 Conference on Empirical Methods in Natural Language Processing, Findings of EMNLP; Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; pp. 15012–15032. [Google Scholar] [CrossRef]
- Thota, Y.R.; Rafatirad, S.; Houman, H.; Nikoubin, T. When Models Ignore Definitions: Measuring Semantic Override Hallucinations in LLM Reasoning. arXiv 2026, arXiv:2602.17520. [Google Scholar] [CrossRef]
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 Statement: An Updated Guideline for Reporting Systematic Reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef]
- Block, N. On a Confusion about a Function of Consciousness. Behav. Brain Sci. 1995, 18, 227–247. [Google Scholar] [CrossRef]
- Baars, B.J.; Geld, N.; Kozma, R. Global Workspace Theory (GWT) and Prefrontal Cortex: Recent Developments. Front. Psychol. 2021, 12, 749868. [Google Scholar] [CrossRef]
- Arévalo-Royo, J.; Latorre-Biel, J.-I.; Flor-Montalvo, F.-J. Cognitive Systems and Artificial Consciousness: What It Is Like to Be a Bat Is Not the Point. Metrics 2025, 2, 11. [Google Scholar] [CrossRef]
- Wei, J.; Bosma, M.; Zhao, V.Y.; Guu, K.; Yu, A.W.; Lester, B.; Du, N.; Dai, A.M.; Le, Q.V. Finetuned Language Models Are Zero-Shot Learners. arXiv 2022, arXiv:2109.01652. [Google Scholar] [CrossRef]
- Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. Training Verifiers to Solve Math Word Problems. arXiv 2021, arXiv:2110.14168. [Google Scholar] [CrossRef]
- Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.; Chi, E.; Narang, S.; Chowdhery, A.; Zhou, D. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv 2023, arXiv:2203.11171. [Google Scholar] [CrossRef]
- Dhuliawala, S.; Komeili, M.; Xu, J.; Raileanu, R.; Li, X.; Celikyilmaz, A.; Weston, J. Chain-of-Verification Reduces Hallucination in Large Language Models. In Proceedings of the Findings of the Association for Computational Linguistics ACL 2024; Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; pp. 3563–3578. [Google Scholar]
- Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T.L.; Cao, Y.; Narasimhan, K. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. In Proceedings of the 37th International Conference on Neural Information Processing Systems; Curran Associates Inc.: New Orleans, LA, USA, 2023; pp. 11809–11822. [Google Scholar]
- Lightman, H.; Kosaraju, V.; Burda, Y.; Edwards, H.; Baker, B.; Lee, T.; Leike, J.; Schulman, J.; Sutskever, I.; Cobbe, K. Let’s Verify Step by Step. arXiv 2023, arXiv:2305.20050. [Google Scholar] [CrossRef]
- Ye, X.; Chen, Q.; Dillig, I.; Durrett, G. SATLM: Satisfiability-Aided Language Models Using Declarative Prompting. In Proceedings of the 37th International Conference on Neural Information Processing Systems; Curran Associates Inc.: New Orleans, LA, USA, 2023; pp. 45548–45580. [Google Scholar]
- Singh, V.; Cassel, D.; Weir, N.; Feng, N.; Bayless, S. VERGE: Formal Refinement and Guidance Engine for Verifiable LLM Reasoning. arXiv 2026, arXiv:2601.20055. [Google Scholar] [CrossRef]
- Du, Y.; Li, S.; Torralba, A.; Tenenbaum, J.B.; Mordatch, I. Improving Factuality and Reasoning in Language Models through Multiagent Debate. In Proceedings of the 41st International Conference on Machine Learning; JMLR: Vienna, Austria, 2024; pp. 11733–11763. [Google Scholar]
- Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; Steinhardt, J. Measuring Mathematical Problem Solving with the MATH Dataset. arXiv 2021, arXiv:2103.03874. [Google Scholar] [CrossRef]
- Lin, S.; Hilton, J.; Evans, O. TruthfulQA: Measuring How Models Mimic Human Falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2022; pp. 3214–3252. [Google Scholar]
- Kazemi, M.; Fatemi, B.; Bansal, H.; Palowitch, J.; Anastasiou, C.; Mehta, S.V.; Jain, L.K.; Aglietti, V.; Jindal, D.; Chen, P.; et al. BIG-Bench Extra Hard. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 26473–26501. [Google Scholar]
- Farquhar, S.; Kossen, J.; Kuhn, L.; Gal, Y. Detecting Hallucinations in Large Language Models Using Semantic Entropy. Nature 2024, 630, 625–630. [Google Scholar] [CrossRef]
- Manakul, P.; Liusie, A.; Gales, M. SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing; Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 9004–9017. [Google Scholar]
- Chua, J.; Rees, E.; Batra, H.; Bowman, S.R.; Michael, J.; Perez, E.; Turpin, M. Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought. arXiv 2025, arXiv:2403.05518. [Google Scholar] [CrossRef]
- Mündler, N.; He, J.; Jenko, S.; Vechev, M. Self-Contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation. arXiv 2024, arXiv:2305.15852. [Google Scholar] [CrossRef]
- Li, J.; Cheng, X.; Zhao, W.X.; Nie, J.Y.; Wen, J.R. HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models. In Proceedings of the EMNLP 2023 Conference on Empirical Methods in Natural Language Processin; Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 6449–6464. [Google Scholar] [CrossRef]
- Min, S.; Krishna, K.; Lyu, X.; Lewis, M.; Yih, W.T.; Koh, P.W.; Iyyer, M.; Zettlemoyer, L.; Hajishirzi, H. FActScore: Fine-Grained Atomic Evaluation of Factual Precision in Long Form Text Generation. In Proceedings of the EMNLP 2023 Conference on Empirical Methods in Natural Language Processing; Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 12076–12100. [Google Scholar] [CrossRef]
- Taubenfeld, A.; Sheffer, T.; Ofek, E.; Feder, A.; Goldstein, A.; Gekhman, Z.; Yona, G. Confidence Improves Self-Consistency in LLMs. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2025; Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 20090–20111. [Google Scholar]
- Kossen, J.; Han, J.; Razzak, M.; Schut, L.; Malik, S.; Gal, Y. Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs. arXiv 2024, arXiv:2406.15927. [Google Scholar] [CrossRef]
- Azaria, A.; Mitchell, T. The Internal State of an LLM Knows When It’s Lying. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2023; Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 967–976. [Google Scholar]
- Damirchi, H.; De la Jara, I.M.; Abbasnejad, E.; Shamsi, A.; Zhang, Z.; Shi, J. Truth as a Trajectory: What Internal Representations Reveal About Large Language Model Reasoning. arXiv 2026, arXiv:2603.01326. [Google Scholar] [CrossRef]
- Sriramanan, G.; Bharti, S.; Sadasivan, V.S.; Saha, S.; Kattakinda, P.; Feizi, S. LLM-Check: Investigating Detection of Hallucinations in Large Language Models. Adv. Neural Inf. Process. Syst. 2024, 37, 34188–34216. [Google Scholar] [CrossRef]
- Chuang, Y.S.; Qiu, L.; Hsieh, C.Y.; Krishna, R.; Kim, Y.; Glass, J. Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps. In Proceedings of the EMNLP 2024 Conference on Empirical Methods in Natural Language Processing; Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; pp. 1419–1436. [Google Scholar] [CrossRef]
- Cao, C.; Yang, J.; Li, H.; Pan, K.; Zhao, Z.; Chen, Z.; Tian, Y.; Wu, L.; He, C.; Han, S.; et al. Pushing the Boundaries of Natural Reasoning: Interleaved Bonus from Formal-Logic Verification. arXiv 2026, arXiv:2601.22642. [Google Scholar] [CrossRef]
- Khalifa, M.; Agarwal, R.; Logeswaran, L.; Kim, J.; Peng, H.; Lee, M.; Lee, H.; Wang, L. Process Reward Models That Think. arXiv 2025, arXiv:2504.16828. [Google Scholar] [CrossRef]
- Huang, J.; Chen, X.; Mishra, S.; Zheng, H.S.; Yu, A.W.; Song, X.; Zhou, D. Large Language Models Cannot Self-Correct Reasoning Yet. arXiv 2024, arXiv:2310.01798. [Google Scholar] [CrossRef]
- Kamoi, R.; Zhang, Y.; Zhang, N.; Han, J.; Zhang, R. When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs. Trans. Assoc. Comput. Linguist. 2024, 12, 1417–1440. [Google Scholar] [CrossRef]
- Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; Cao, Y. React: Synergizing Reasoning and Acting in Language Models. In Proceedings of the 11th International Conference on Learning Representations (ICLR 2023), Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
- Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; Yao, S. Reflexion: Language Agents with Verbal Reinforcement Learning. Adv. Neural Inf. Process. Syst. 2023, 36, 8634–8652. [Google Scholar]
- Gao, L.; Madaan, A.; Zhou, S.; Alon, U.; Liu, P.; Yang, Y.; Callan, J.; Neubig, G. PAL: Program-Aided Language Models. Proc. Mach. Learn. Res. 2022, 202, 10764–10799. [Google Scholar]
- Phan, L.; Gatti, A.; Han, Z.; Li, N.; Hu, J.; Zhang, H.; Zhang, C.B.C.; Shaaban, M.; Ling, J.; Shi, S.; et al. Humanity’s Last Exam. arXiv 2026, arXiv:2501.14249. [Google Scholar] [CrossRef]
- Colelough, B.C.; Regli, W. Neuro-Symbolic AI in 2024: A Systematic Review. arXiv 2025, arXiv:2501.05435. [Google Scholar] [CrossRef]
- Fabiano, F.; Ganapini, M.B.; Loreggia, A.; Mattei, N.; Murugesan, K.; Pallagani, V.; Rossi, F.; Srivastava, B.; Venable, K.B. Thinking Fast and Slow in Human and Machine Intelligence. Commun. ACM 2025, 68, 72–79. [Google Scholar] [CrossRef]
- Goyal, A.; Bengio, Y. Inductive Biases for Deep Learning of Higher-Level Cognition. Proc. R. Soc. A Math. Phys. Eng. Sci. 2022, 478, 20210068. [Google Scholar] [CrossRef]
- Rein, D.; Hou, B.L.; Stickland, A.C.; Petty, J.; Pang, R.Y.; Dirani, J.; Michael, J.; Bowman, S.R. GPQA: A Graduate-Level Google-Proof Q&A Benchmark. arXiv 2023, arXiv:2311.12022. [Google Scholar] [CrossRef]








| Anomaly Type | Manifestation | Example | Detection Method |
|---|---|---|---|
| Factual hallucination | Fabrication of facts | Inventing citations | Semantic entropy [39] |
| Self-contradiction | Inconsistent claims | Stating X is true and then false | SelfCheckGPT [40] |
| Unfaithful CoT | Trace ≠ computation | Post hoc rationalization | Faithfulness probing [18] |
| Semantic nullification | Reverts to defaults | Ignoring redefinitions | Consistency tests [22] |
| Snowball error | Errors amplify in chains | Arithmetic carry errors | Process reward [32] |
| Susceptibility to distractors | Irrelevant context disrupts | Unrelated data in math | Robustness benchmark [11] |
| Sycophancy bias | Agrees regardless | Changing correct answer | Bias-Augmented training [41] |
| Phenomenon | Data Quality Impact | Reasoning Impact | Source |
|---|---|---|---|
| Loss in the middle | >30% drop by position | N/A | [6] |
| Irrelevant context | 1 distractor sentence | Accuracy < 30% | [11] |
| Knowledge vs. reasoning | 56.9% knowledge (HuatuoGPT-o1) | 44.8% reasoning (HuatuoGPT-o1) | [13] |
| RAG poisoning | 90% ASR with 5 texts | Models do not detect | [8] |
| Module [26] | Function | Anomaly Detection Role | Implementation |
|---|---|---|---|
| Selective attention | Information filtering | Context relevance scoring | RAG re-ranking |
| Working memory | Temporary storage | Reasoning trace buffering | Extended context |
| Introspective representation | Internal state | Confidence estimation | Semantic entropy [46] |
| Reasoning system | Inference engine | Multi-path generation | Self-consistency [29] |
| Execution monitor | Performance evaluation | Step-by-step verification | Process reward [32] |
| Reporting system | Diagnostics | Anomaly report generation | Explainable AI |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Arévalo-Royo, J.; Martín, Ó.; Martínez-Cámara, E.; Flor-Montalvo, F.-J.; Blanco-Fernández, J. Anomalies in AI Outputs Beyond Input Data Quality: The Significance of Reasoning. Appl. Sci. 2026, 16, 5491. https://doi.org/10.3390/app16115491
Arévalo-Royo J, Martín Ó, Martínez-Cámara E, Flor-Montalvo F-J, Blanco-Fernández J. Anomalies in AI Outputs Beyond Input Data Quality: The Significance of Reasoning. Applied Sciences. 2026; 16(11):5491. https://doi.org/10.3390/app16115491
Chicago/Turabian StyleArévalo-Royo, Javier, Óscar Martín, Eduardo Martínez-Cámara, Francisco-Javier Flor-Montalvo, and Julio Blanco-Fernández. 2026. "Anomalies in AI Outputs Beyond Input Data Quality: The Significance of Reasoning" Applied Sciences 16, no. 11: 5491. https://doi.org/10.3390/app16115491
APA StyleArévalo-Royo, J., Martín, Ó., Martínez-Cámara, E., Flor-Montalvo, F.-J., & Blanco-Fernández, J. (2026). Anomalies in AI Outputs Beyond Input Data Quality: The Significance of Reasoning. Applied Sciences, 16(11), 5491. https://doi.org/10.3390/app16115491

