Large Language Models and Retrieval-Augmented Generation in Natural Language Processing, Human–Robot Interaction and Quantum Computing

A Special Issue of AI (ISSN 2673-2688) belonging to the section "AI Systems: Theory and Applications".

Deadline for manuscript submissions: closed (30 June 2026) | Viewed by 16331

Editor


E-Mail Website
Guest Editor
Institute of High Performance Computing and Networking (ICAR), National Research Council of Italy (CNR), Via Pietro Bucci, 8-9 C, 87036 Rende, CS, Italy
Interests: artificial intelligence; machine learning; natural language processing; information retrieval; human–robot interaction; intelligent document processing; question answering; knowledge graphs; quantum computing

Special Issue Information

Dear Colleagues,

Recent advances in Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) have transformed Natural Language Processing and adjacent fields. This Special Issue explores the synergistic relationship between these technologies across three key domains—Natural Language Processing, Human–Robot Interaction, and Quantum Computing—paying particular attention to evaluation methodologies, reliability assessment, explainability, and knowledge graph integration.

LLMs have demonstrated remarkable capabilities in understanding and generating human-like text but face limitations, including hallucinations, knowledge cutoffs, and contextual constraints. RAG has emerged as a powerful framework that enhances LLMs by dynamically retrieving relevant external information—often presented as knowledge graphs—to ground model outputs in factual knowledge. This integration represents a significant frontier in AI research, combining the generative power of LLMs with the precision of retrieval systems.

In Natural Language Processing, RAG addresses the fundamental challenges of achieving factuality and up-to-date knowledge, enabling more reliable and contextually appropriate text generation with improved explainability. Knowledge graphs serve as structured repositories that can be queried to verify and supplement LLM outputs. In Human–Robot Interaction, LLM-RAG systems are revolutionizing how robots understand and respond to natural language instructions. These "agentic computing" systems enable robots to process open-ended commands, execute multi-step processes, and function more intuitively. However, challenges remain in grounding robot actions in the physical world, mitigating potentially unsafe behaviors caused by hallucinations, and developing robust error-handling mechanisms that ensure reliability.

Quantum Computing offers transformative opportunities for enhancing RAG systems, particularly in vector search optimization for information retrieval. The computationally intensive similarity calculations critical to document retrieval represent a bottleneck for large-scale implementations. Quantum-enhanced information retrieval could potentially revolutionize the efficiency of these operations. Additionally, RAG-enhanced LLMs show significant promise for quantum code generation, making quantum programming more accessible to non-specialists by providing contextually relevant code examples and documentation.

This Special Issue will supplement the existing literature by bridging these rapidly evolving fields, which have primarily been studied in isolation. While substantial research exists on LLMs, knowledge graphs, RAG, and their individual applications, their synergistic integration—especially in specialized domains like Human–Robot Interaction and Quantum Computing—remains underexplored. By fostering interdisciplinary dialogue, this Special Issue will accelerate innovation in areas where traditional approaches face limitations, placing a particular emphasis on evaluation methodologies and explainability to ensure these advanced systems are not only powerful but also trustworthy and interpretable.

We welcome the submission of original research papers, reviews, and case studies investigating theoretical foundations, methodological innovations, practical implementations, and ethical considerations in this emerging technology ecosystem. By bringing together diverse perspectives and cutting-edge research, we aim to advance our understanding of how these technologies can be effectively integrated to develop more capable, reliable, and trustworthy AI systems.

The topics of interest include but are not limited to the following:

RAG Architectures and Foundations:

- Novel RAG architectures and retrieval mechanisms optimized for various LLM frameworks.

- The integration of knowledge graphs with LLMs for enhanced retrieval and reasoning.

- Techniques for improving retrieval quality, relevance assessment, and contextual integration.

- Approaches to knowledge management, updates, and verification in RAG systems.

- Strategies for handling multi-modal information retrieval and generation.

Evaluation, Reliability, and Explainability:

- Methods for evaluating RAG performance and achieving hallucination reduction.

- Reliability assessment frameworks for LLM-RAG systems.

- Explainability techniques for understanding and interpreting RAG-based outputs.

- Standardizing protocols for reliable LLM and RAG evaluations.

- Hybrid evaluation frameworks combining automated and human annotations.

- User-centric evaluations, subjective assessments, and the personalization of RAG systems.

- Ethical considerations, bias mitigation, and transparency in combined LLM-RAG systems.

Human–Robot Interaction Applications:

- LLM-RAG systems for enhanced natural language understanding in human–robot interaction.

- Grounding robot actions through knowledge retrieval and contextual understanding.

- Vision–Language–Action Models (VLAs) that fuse vision, language, and actions.

- Safety mechanisms and reliability assessment for LLM-powered robotic systems.

- Multi-agent collaboration frameworks for robotic systems.

- Knowledge graph utilization for robotic task planning and execution.

Quantum Computing Integration:

- Quantum-enhanced information retrieval for RAG systems.

- Vector search optimization using quantum computing approaches.

- Theoretical frameworks and experimental implementations of quantum algorithms for document retrieval.

- Hybrid quantum–classical approaches for large-scale information retrieval.

- Performance comparisons between classical and quantum-enhanced RAG systems.

- Quantum code generation and assistance using LLMs and RAG.

- RAG frameworks for domain-specific quantum programming using frameworks like PennyLane and Qiskit.

Domain-Specific Applications:

- Domain-specific RAG applications in healthcare, legal and scientific research, and other specialized fields.

- Knowledge graph construction and utilization in specialized domains.

- Dataset curation and annotation methodologies for training LLMs in specialized domains.

- Explainable AI approaches for domain-specific applications requiring interpretability.

Dr. Ermelinda Oro
Guest Editor

Manuscript Submission Information

Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.

Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. AI is an international peer-reviewed open access monthly journal published by MDPI.

Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 1800 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.

Keywords

  • large language models
  • retrieval-augmented generation
  • natural language processing
  • knowledge graphs
  • knowledge grounding
  • human–robot interaction
  • agentic computing
  • vision–language–action models
  • quantum computing
  • quantum-enhanced information retrieval
  • evaluation methodologies
  • reliability assessment
  • explainability
  • hallucination mitigation
  • vector search optimization
  • quantum code generation
  • multi-modal RAG

Benefits of Publishing in a Special Issue

  • Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
  • Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
  • Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
  • External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
  • Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

Further information on MDPI's Special Issue policies can be found here.

Published Papers (5 papers)

Order results
Result details
Select all
Export citation of selected articles as:

Research

19 pages, 357 KB  
Article
Retrieval Granularity as Evidence Design in Small-Model RAG Question Answering: A Diagnostic HotpotQA Study
by Weimao Ke, Lixiao Yang and Mengyang Xu
AI 2026, 7(8), 320; https://doi.org/10.3390/ai7080320 - 19 Aug 2026
Viewed by 370
Abstract
Retrieval-Augmented Generation (RAG) has become a practical approach for question answering over external corpora, particularly when answers should be grounded in source documents rather than generated only from model parameters. While recent large language models can process increasingly long contexts, they do not [...] Read more.
Retrieval-Augmented Generation (RAG) has become a practical approach for question answering over external corpora, particularly when answers should be grounded in source documents rather than generated only from model parameters. While recent large language models can process increasingly long contexts, they do not remove the need for selecting, organizing, and auditing evidence, especially when systems rely on smaller local models for privacy, cost, or deployment constraints. In this paper, we frame retrieval granularity as an evidence-design variable for answer grounding in small-model RAG question answering. After a brief exploratory NewsQA phase that motivates the error categories, the main study uses the HotpotQA distractor validation split with 7405 hard multi-hop questions and sentence-level supporting-fact annotations. With Qwen3-8B as the fixed generator, we compare closed-book, fixed-budget whole-context, retrieved-context, gold-document, and gold-supporting-fact conditions while varying retrieval granularity, retriever type, and context budget. Retrieved context substantially outperforms closed-book answering and the 1024-token fixed-budget whole-context condition but remains below gold-document and gold-supporting-fact upper bounds, indicating that retrieval, generation, and evaluation limitations should be analyzed separately. Sentence-level retrieval under-recovers multi-hop evidence, especially for questions with three or more supporting facts, while paragraph-level and moderate token-level chunks recover substantially more complete evidence. In the full condition matrix, hybrid retrieval with 256-token chunks and no overlap achieves an F1 of 0.6816 with a supporting-fact recall of 0.9609, compared with an F1 of 0.6166 and supporting-fact recall of 0.7801 for BM25 sentence retrieval. Additional ablations show that fixed-budget whole-context performance is strongly affected by truncation, that overlap has little practical effect under the tested 1024-token budget, and that a stronger BGE dense retriever improves the best retrieved-context F1 to 0.7027. These results align with a diagnostic perspective on chunking: using evidence at a task-appropriate level of granularity can improve grounding, auditability, and answer quality, but the observed patterns should be interpreted within the HotpotQA distractor setting, fixed generator, and tested context budgets. Full article
Show Figures

Figure 1

22 pages, 1659 KB  
Article
FroLineR: Front-Line Response with Retrieval-Augmented Prompt-Engineered Reply Generation for IT Help Desks
by Alexandru Dima, Maria-Elena Mihăilescu, Darius Mihai, Mihai Carabaș and Mihai Dascalu
AI 2026, 7(8), 317; https://doi.org/10.3390/ai7080317 - 19 Aug 2026
Viewed by 454
Abstract
IT help desks at large organizations face a high volume of recurrent, well-documented user requests that nevertheless require human-written replies, creating a persistent staff workload that is repetitive in content but non-trivial in tone and procedural correctness. We present FroLineR, short for Front-Line [...] Read more.
IT help desks at large organizations face a high volume of recurrent, well-documented user requests that nevertheless require human-written replies, creating a persistent staff workload that is repetitive in content but non-trivial in tone and procedural correctness. We present FroLineR, short for Front-Line Response, a system that drafts the initial staff reply to such tickets in the login and account-activation category and integrates into a human-in-the-loop ticketing workflow on a Romanian-language ticketing platform. The generator is an unmodified instruct model augmented with retrieval from a small set of hand-curated guide documents, using a Romanian system prompt refined over several rounds of staff review. To evaluate and refine the prompt without manual labeling, we cluster the first user message of every historical thread with both BERTopic and Semantic Signal Separation (S3), score configurations along coherence and lexical-diversity axes, and extract a 200-message evaluation set from the winning model. Prompt convergence was certified by several rounds of manual review by support staff. The production system is quantized to Q4_K_M GGUF, served through llama-cpp-python behind a small Flask API, and deployed with GPU offloading on the target server, reducing end-to-end per-answer latency from approximately 830 s on the server’s CPU to roughly 61 s once layers are offloaded to the GPU, with no observable degradation in answer quality. Full article
Show Figures

Figure 1

28 pages, 1439 KB  
Article
Cross-LLM Paraphrase Laundering: A Register-Controlled Evaluation of Fake News Detectors
by Jalal Mehdiyev and Ramiz Aliguliyev
AI 2026, 7(8), 297; https://doi.org/10.3390/ai7080297 - 3 Aug 2026
Viewed by 847
Abstract
Detectors built on transformer language models report near-perfect accuracy on standard fake news benchmarks, which suggests the task is almost solved. We argue that much of this accuracy reflects a confound between writing register and veracity: in common benchmarks, the real class is [...] Read more.
Detectors built on transformer language models report near-perfect accuracy on standard fake news benchmarks, which suggests the task is almost solved. We argue that much of this accuracy reflects a confound between writing register and veracity: in common benchmarks, the real class is human-written while the fake class is machine-generated or machine-rewritten, so a detector can separate the classes by recognizing AI writing style rather than by judging truth. To evaluate this, we designed a two-regime evaluation. Phase 1 is the laundering regime, comparing untouched human-real articles against laundered fake articles. Phase 2 is register-controlled: the real class is passed through the same cross-LLM laundering chains as the fake class, so both classes are read in one machine register, and the register cue is no longer available to the detector. We train seven detectors on three datasets and evaluate each frozen detector under both regimes. Under Phase 1, detectors appear robust; under Phase 2, detection on WELFake collapses from about 99% to about 62% AUROC and the largest models approach chance. The effect is benchmark dependent, large on WELFake, mild on IFND and near zero on GossipCop, and it is confirmed by bootstrap testing with false discovery rate control. We recommend register-controlled evaluation as standard reporting practice. Full article
Show Figures

Figure 1

30 pages, 980 KB  
Article
Hallucination Mitigation in Large Language Model-Based Tool Recommendation: A Cross-Provider Architectural Ablation Study Across Two Model Generations
by Lavdim Menxhiqi and Galia Marinova
AI 2026, 7(7), 273; https://doi.org/10.3390/ai7070273 - 22 Jul 2026
Viewed by 3053
Abstract
In a closed-inventory large language model (LLM) system such as Online-CADCOM, which recommends engineering tools from a verified inventory, we measure inventory non-compliance, that is, a mention-level event in which the model recommends a tool not present in the verified inventory. We use [...] Read more.
In a closed-inventory large language model (LLM) system such as Online-CADCOM, which recommends engineering tools from a verified inventory, we measure inventory non-compliance, that is, a mention-level event in which the model recommends a tool not present in the verified inventory. We use this inventory-relative sense of hallucination throughout: an out-of-inventory mention may be a fabricated tool or a real commercial tool absent from the curated inventory, so the metric reports inventory non-compliance rather than factual fabrication. We evaluate a three-mechanism mitigation stack consisting of database-grounded context injection, fixed vocabulary constraints, and enforced JavaScript Object Notation (JSON) output across three commercial LLM providers (OpenAI, Anthropic, Google), two model generations, and two output modes (standard and reasoning), totaling 6912 Application Programming Interface (API) calls over 12 configurations. Under a recall-equalized detector adopted as the primary metric, the inventory non-compliance rate, which we denote the hallucination rate (HR) following common usage, decreases from roughly 69–80% to 4–13% under the full architecture. The cross-provider average is similar across the two generations tested (8.5% Generation 1 (Gen1), 6.9% Generation 2 (Gen2)), although per-provider directions diverge. We also examine the C3 configuration, in which only JSON output enforcement is active without grounding. A naive detector reports a large hallucination increase over the unconstrained baseline (+10.1 percentage points (pp) Gen1, +15.1 pp Gen2), but we show this gap is largely a detection-format artifact: structured JSON fields make out-of-inventory tools easy to extract, whereas the same real tools are frequently missed in free text. Under a recall-equalized detector the gap narrows to +2.6 pp (Gen1) and +4.8 pp (Gen2) and remains statistically significant only for two current-generation models, indicating a small, current-generation effect rather than a universal one. Reasoning-mode models provide no statistically significant improvement under architectural constraints. A frequency-weighted audit shows that the majority of remaining out-of-inventory mentions correspond to real engineering tools absent from the platform’s inventory. Under the full architecture, roughly half of responses (pooled Pany49.5%) still contain at least one such mention, indicating that handling unseen tools remains an open challenge for closed-inventory recommendation systems. Our evidence comes from a single engineering platform with four related electronic-design and power-electronics domains, so the findings characterize this setting rather than recommendation domains in general. Full article
Show Figures

Figure 1

30 pages, 2591 KB  
Article
Prompt Optimization with Two Gradients for Classification in Large Language Models
by Anthony Jethro Lieander, Hui Wang and Karen Rafferty
AI 2025, 6(8), 182; https://doi.org/10.3390/ai6080182 - 8 Aug 2025
Cited by 1 | Viewed by 9415
Abstract
Large language models (LLMs) generally perform well in common tasks, yet are often susceptible to errors in sophisticated natural language processing (NLP) on classification applications. Prompt engineering has emerged as a strategy to enhance their performance. Despite the effort required for manual prompt [...] Read more.
Large language models (LLMs) generally perform well in common tasks, yet are often susceptible to errors in sophisticated natural language processing (NLP) on classification applications. Prompt engineering has emerged as a strategy to enhance their performance. Despite the effort required for manual prompt optimization, recent advancements highlight the need for automation to reduce human involvement. We introduced the PO2G (prompt optimization with two gradients) framework to improve the efficiency of optimizing prompts for classification tasks. PO2G demonstrates improvement in efficiency, reaching almost 89% accuracy after just three iterations, whereas ProTeGi requires six iterations to achieve a comparable level. We evaluated PO2G and ProTeGi on a benchmark of nine NLP tasks, three tasks from the original ProTeGi study, and six non-domain-specific tasks. We also evaluated both frameworks on seven legal-domain classification tasks. These results provide broader insights into the efficiency and effectiveness of prompt optimization frameworks for classification across diverse NLP scenarios. Full article
Show Figures

Figure 1

Back to TopTop