Next Issue
Volume 7, September
Previous Issue
Volume 7, July
 
 

AI, Volume 7, Issue 8 (August 2026) – 52 articles

  • Issues are regarded as officially published after their release is announced to the table of contents alert mailing list.
  • You may sign up for e-mail alerts to receive table of contents of newly released issues.
  • PDF is the official format for papers published in both, html and pdf forms. To view the papers in pdf format, click on the "PDF Full-text" link, and use the free Adobe Reader to open them.
Cover Story (view full-size image):
Order results
Result details
Section
Select all
Export citation of selected articles as:
26 pages, 4251 KB  
Article
Evaluating Feature-Based Machine-Learning Models with Post Hoc Explainability for Eye-Tracking-Based Task Type and Workload Inference
by Tomi Božak, Shivalika Goyal, Marc Langheinrich, Martin Gjoreski and Gašper Slapničar
AI 2026, 7(8), 325; https://doi.org/10.3390/ai7080325 - 21 Aug 2026
Viewed by 352
Abstract
Eye tracking is a valuable behavioral signal for human-centered AI, yet the reliability of feature-based machine-learning models for inferring task type and workload across users and tasks remains uncertain, because experimentally defined workload labels may reflect task type and visual structure as much [...] Read more.
Eye tracking is a valuable behavioral signal for human-centered AI, yet the reliability of feature-based machine-learning models for inferring task type and workload across users and tasks remains uncertain, because experimentally defined workload labels may reflect task type and visual structure as much as cognitive demand. The practical problem is that designers of gaze-adaptive systems need to know which inferences are dependable enough to act on, and reported accuracies alone do not answer this, because the choice of prediction target and validation split can determine the result. This study systematically evaluates feature-based machine-learning models with post hoc explainability across three prediction targets: task type, binary load-versus-rest, and three-level workload. Eye-movement features derived from fixations, saccades, pupils, and blinks were extracted from short temporal windows collected from 54 participants performing attention, visual-spatial, and memory tasks under rest, easy, and difficult conditions, and evaluated using leave-one-subject-out (LOSO) and leave-one-group-out (LOGO) validation. Task type was classified most reliably (85.9% LOSO, 83.4% LOGO), binary load-versus-rest showed moderate, validation-sensitive robustness (81.4% LOSO, 63.9% LOGO), and three-level workload classification was substantially more challenging (56.4% LOSO, 44.3% LOGO). SHAP and statistical analyses consistently identified fixation dispersion, pupil-related measures, and subject-normalized features as the strongest contributors across all three targets. These findings show that prediction target definition, validation strategy, and post hoc explainability jointly determine what can be reliably inferred from gaze-based machine-learning models. Eye tracking alone therefore appears promising for task-type recognition and may support coarse engagement-related inference when the deployment task family is represented during model development, whereas task-independent fine-grained workload estimation remains unsupported by the present evidence. Full article
(This article belongs to the Special Issue Human-Computer Interaction and Human-Centered AI)
Show Figures

Figure 1

34 pages, 4998 KB  
Perspective
From Empowerment to Vulnerability: The Computation–Energy Paradox of AI-Enabled Power-Transport Systems
by Chenxuan Zhang, Peixiao Fan, Siqi Bu and Yuxin Wen
AI 2026, 7(8), 324; https://doi.org/10.3390/ai7080324 - 21 Aug 2026
Viewed by 603
Abstract
The transition towards smart megacities has deeply integrated Artificial Intelligence (AI) with power–transport networks. While AI empowers complex operations like multi-network coordinated dispatch and emergency rescue, current algorithm-centric perspectives largely ignore its massive physical energy costs. Accordingly, this Perspective examines the dual role [...] Read more.
The transition towards smart megacities has deeply integrated Artificial Intelligence (AI) with power–transport networks. While AI empowers complex operations like multi-network coordinated dispatch and emergency rescue, current algorithm-centric perspectives largely ignore its massive physical energy costs. Accordingly, this Perspective examines the dual role of AI, considering it not only as an intelligent decision-support tool but also as a potential source of additional stress on physical infrastructure. First, through a structured synthesis of the representative literature, we deconstruct the functional dependencies between algorithms and physical infrastructures, identifying how AI reshapes the operational paradigms of power, ground transport, and aerial networks under routine and emergency scenarios. We then introduce the concept of the “Computation–Energy Paradox.” Integrating conceptual analysis with a quantitative case study of a typical community, we illustrate a plausible failure mechanism: during extreme disasters, intensified AI invocation for emergency management generates surging computational loads, which paradoxically exacerbate power shortages and reduce the operating margin of already weakened systems. In addition, we analyze core engineering bottlenecks, including spatiotemporal computation–energy mismatches and physical constraints in extreme edge environments. To address these challenges, we outline a prospective roadmap encompassing lightweight emergency AI and computation–power-coordinated offloading mechanisms. Finally, the sustainable development of such systems suggests a paradigm shift: AI must evolve from a purely virtual algorithm into a physical component of an integrated compute–power–transport system. Full article
Show Figures

Figure 1

25 pages, 4188 KB  
Article
Systematic Comparison of Electroencephalography Feature Domains for Visual Stimuli Decoding with EEGNet and EEG Conformer
by Cesar Agustin Corona-Patricio, Carolina Reta and Jose Antonio Cantoral-Ceballos
AI 2026, 7(8), 323; https://doi.org/10.3390/ai7080323 - 20 Aug 2026
Viewed by 360
Abstract
Electroencephalography-based visual decoding has important applications in brain–computer interfaces and cognitive neuroscience, yet the relative effectiveness of different feature extraction methods for sustained visual paradigms remains unclear due to the absence of standardized, multi-dataset comparative evaluations. This study systematically compares eight feature extraction [...] Read more.
Electroencephalography-based visual decoding has important applications in brain–computer interfaces and cognitive neuroscience, yet the relative effectiveness of different feature extraction methods for sustained visual paradigms remains unclear due to the absence of standardized, multi-dataset comparative evaluations. This study systematically compares eight feature extraction methods across three public EEG datasets: MindBigData MNIST, MindBigData MNIST-8B for digit recognition, and MSS for natural image classification. The methods include coherence, Granger causality, directed transfer function, partially directed coherence, transfer entropy, discrete wavelet transform, empirical wavelet transform (EWT), and wavelet scattering transform. Two deep learning architectures, EEGNet and EEG Conformer, were trained using two pre-processing pipelines, with and without artifact removal. EWT achieved the highest classification accuracy, reaching 97.83% for digit-vs-blank and 77.10% for within-session natural image classification. Connectivity-based methods consistently underperformed, with the best connectivity method (coherence) reaching up to 91.67%, suggesting that spectral power information is more discriminative than inter-channel relationships. Cross-subject generalization remained challenging, with best accuracies near 68%. The findings establish wavelet-based adaptive spectral decomposition as a strong baseline for EEG visual decoding and highlight the need for domain adaptation techniques to address cross-subject variability. Full article
Show Figures

Figure 1

52 pages, 4148 KB  
Review
The Governance Gap in Contemporary LLM-Based Agentic Systems: A Structural Diagnostic Review
by Christopher Valdez-Cantú, Jose Antonio Cantoral-Ceballos and Joanna Alvarado-Uribe
AI 2026, 7(8), 322; https://doi.org/10.3390/ai7080322 - 20 Aug 2026
Viewed by 653
Abstract
Large Language Models (LLMs) are increasingly integrated into agentic workflows that require extended reasoning, persistent state management, coordinated tool use, and controlled execution. As this operational scope expands, a central question emerges: whether probabilistic generation alone can reliably support coherent behavior across interacting [...] Read more.
Large Language Models (LLMs) are increasingly integrated into agentic workflows that require extended reasoning, persistent state management, coordinated tool use, and controlled execution. As this operational scope expands, a central question emerges: whether probabilistic generation alone can reliably support coherent behavior across interacting system components. This paper addresses that question through a structural diagnostic review of contemporary agentic systems. Starting from LLM-based tutoring as an analytically demanding entry point and extending toward structurally related agent architectures, the paper draws on a five-phase review of N=145 research records. The analysis is organized through the Agentic Structure Taxonomy (AST), which structures the literature across four dimensions: Cognition, Interaction, Orchestration, and Governance. The review identifies five recurrent empirical problem patterns and uses them as abductive diagnostic cues for formulating seven cross-dimensional transition gaps that capture recurrent discontinuities at the boundaries between reasoning, state, control, and execution. From these gaps, fourteen structural constraints are derived across three control domains: state isolation, control alignment, and execution governance. These constraints are interpreted not as prescriptive design mandates, but as analytically derived conditions associated with reducing error propagation across subsystem transitions. The paper argues that reliability in agentic systems is shaped not only by model performance or prompt design, but also by whether the boundaries linking probabilistic reasoning to persistent state, orchestration, and execution are governed by explicit structural conditions. Full article
Show Figures

Figure 1

37 pages, 668 KB  
Article
Benchmarking Normative AI Assistants Under Inconsistent Evidence with Paraconsistent Trace Semantics
by Maksim V. Ulizko, Aleksandr V. Chernikov, Ivan V. Tomilov, Natalia F. Gusarova and Aleksandra S. Vatian
AI 2026, 7(8), 321; https://doi.org/10.3390/ai7080321 - 20 Aug 2026
Viewed by 507
Abstract
Normative AI assistants are increasingly used in domains governed by duties, permissions, prohibitions, exceptions, priorities, and institutional policies. Existing retrieval-augmented generation (RAG) and legal AI benchmarks evaluate answer accuracy, retrieval quality, citation grounding, natural-language inference, clause extraction, or general legal reasoning ability. These [...] Read more.
Normative AI assistants are increasingly used in domains governed by duties, permissions, prohibitions, exceptions, priorities, and institutional policies. Existing retrieval-augmented generation (RAG) and legal AI benchmarks evaluate answer accuracy, retrieval quality, citation grounding, natural-language inference, clause extraction, or general legal reasoning ability. These dimensions are necessary but insufficient when supplied evidence is incomplete, mutually inconsistent, or defeasible. The objective of this study is to introduce ParaTraceBench, a paraconsistent trace-based benchmarking framework for post-retrieval normative reasoning over fixed evidence packages. Each scenario contains a query, evidence fragments, extracted facts, defeasible rules, typed attack edges, priority relations, an expected conclusion status, and a gold diagnostic trace. The formalism uses evidence-grounded arguments, a single edge-based attack representation, explicit attack-licensing rules, acyclic priority bases with a transitive closure, grounded argument labeling, trace-normal-form alignment, and deterministic scoring. The operational NER metric is explicitly interpreted as inconsistency-conditioned unsupported-conclusion avoidance rather than proof of logical non-explosion. We evaluated the framework using 140 scenarios, external validation on 567 anonymized Russian-language cases from Russian Federation and EAEU-related materials, reasoning-oriented baseline adaptations, five-run prompt-fairness and stability controls, and a deterministic component-dependency audit. On the full external set, the trace-based configuration reached 85.7% answer-status accuracy, 85.5% contradiction-localization accuracy, 94.2% operational NER, 84.1% priority-handling accuracy, and 83.7% belief-revision accuracy. These results indicate that contradiction-aware trace evaluation provides diagnostic information beyond final-answer accuracy under the evaluated fixed-evidence conditions, while not establishing causal architectural superiority, logical non-triviality, or end-to-end RAG performance. Full article
Show Figures

Figure 1

19 pages, 357 KB  
Article
Retrieval Granularity as Evidence Design in Small-Model RAG Question Answering: A Diagnostic HotpotQA Study
by Weimao Ke, Lixiao Yang and Mengyang Xu
AI 2026, 7(8), 320; https://doi.org/10.3390/ai7080320 - 19 Aug 2026
Viewed by 293
Abstract
Retrieval-Augmented Generation (RAG) has become a practical approach for question answering over external corpora, particularly when answers should be grounded in source documents rather than generated only from model parameters. While recent large language models can process increasingly long contexts, they do not [...] Read more.
Retrieval-Augmented Generation (RAG) has become a practical approach for question answering over external corpora, particularly when answers should be grounded in source documents rather than generated only from model parameters. While recent large language models can process increasingly long contexts, they do not remove the need for selecting, organizing, and auditing evidence, especially when systems rely on smaller local models for privacy, cost, or deployment constraints. In this paper, we frame retrieval granularity as an evidence-design variable for answer grounding in small-model RAG question answering. After a brief exploratory NewsQA phase that motivates the error categories, the main study uses the HotpotQA distractor validation split with 7405 hard multi-hop questions and sentence-level supporting-fact annotations. With Qwen3-8B as the fixed generator, we compare closed-book, fixed-budget whole-context, retrieved-context, gold-document, and gold-supporting-fact conditions while varying retrieval granularity, retriever type, and context budget. Retrieved context substantially outperforms closed-book answering and the 1024-token fixed-budget whole-context condition but remains below gold-document and gold-supporting-fact upper bounds, indicating that retrieval, generation, and evaluation limitations should be analyzed separately. Sentence-level retrieval under-recovers multi-hop evidence, especially for questions with three or more supporting facts, while paragraph-level and moderate token-level chunks recover substantially more complete evidence. In the full condition matrix, hybrid retrieval with 256-token chunks and no overlap achieves an F1 of 0.6816 with a supporting-fact recall of 0.9609, compared with an F1 of 0.6166 and supporting-fact recall of 0.7801 for BM25 sentence retrieval. Additional ablations show that fixed-budget whole-context performance is strongly affected by truncation, that overlap has little practical effect under the tested 1024-token budget, and that a stronger BGE dense retriever improves the best retrieved-context F1 to 0.7027. These results align with a diagnostic perspective on chunking: using evidence at a task-appropriate level of granularity can improve grounding, auditability, and answer quality, but the observed patterns should be interpreted within the HotpotQA distractor setting, fixed generator, and tested context budgets. Full article
Show Figures

Figure 1

36 pages, 739 KB  
Article
Detecting AI-Generated Text and Code: An Empirical Study of Cross-Generator and Cross-Domain Generalization
by Neethika Alluri, Pardha Saradhi Varma Gottumukkala and Hemalatha Indukuri
AI 2026, 7(8), 319; https://doi.org/10.3390/ai7080319 - 19 Aug 2026
Viewed by 786
Abstract
Large language models (LLMs) now generate fluent natural language and source code, creating challenges for authorship attribution, academic integrity, and software supply-chain security. Most existing detectors for AI-generated content are evaluated separately on natural language or source code, often under matched train–test conditions [...] Read more.
Large language models (LLMs) now generate fluent natural language and source code, creating challenges for authorship attribution, academic integrity, and software supply-chain security. Most existing detectors for AI-generated content are evaluated separately on natural language or source code, often under matched train–test conditions that can overestimate real-world reliability. We present a paired-prompt benchmark for human-versus-machine detection across English text, Python code, and mixed text–code documents. The benchmark includes 22,141 instances from HC3, CodeSearchNet, MBPP, and HumanEval across training, validation, and test partitions, plus Mix-Eval, a mixed-content set of 997 Jupyter-notebook-style samples. We evaluate RoBERTa-large for text, GraphCodeBERT and CodeBERT-base for code, a unified RoBERTa-base detector trained on both modalities, and zero-shot baselines. Fine-tuned detectors achieve near-perfect in-distribution performance, with AUROC 1.0000±0.0000 and accuracy above 99.5%. Across five instruction-tuned generator families of varying size (3.8B–7B) and architecture, with the human and problem distributions held fixed, cross-generator transfer causes negligible degradation (AUROC spread 0.0002; drops of at most 0.0003). In contrast, domain shift is the main failure mode: on MBPP+HumanEval, GraphCodeBERT drops to 0.85±0.02 AUROC and CodeBERT-base to 0.67±0.02. On Mix-Eval, the unified detector outperforms a routed text–code pipeline by 21 AUROC points (0.96 vs. 0.75), largely because of router failures on mixed inputs. Training-time augmentation improves low-false-positive performance, while legacy supervised detectors show systematic class inversion on modern LLM outputs. These results show that reliable deployment requires cross-domain evaluation, mixed-content testing, and calibration beyond in-distribution accuracy. Full article
Show Figures

Figure 1

15 pages, 5795 KB  
Article
CBR-Enhanced ResNet50 for Five-Class Diabetic Retinopathy Grading: An Ablation-Based Study
by Samir Elouaham, Fatima Ezzahra Bouaaza, Ilyas Ait Ichou and Boujemaa Nassiri
AI 2026, 7(8), 318; https://doi.org/10.3390/ai7080318 - 19 Aug 2026
Viewed by 253
Abstract
Diabetic retinopathy (DR) is a common complication of diabetes and one of the leading causes of preventable vision loss worldwide. Because the manual grading of color fundus images is slow and depends on the availability of trained specialists, automated screening tools are needed. [...] Read more.
Diabetic retinopathy (DR) is a common complication of diabetes and one of the leading causes of preventable vision loss worldwide. Because the manual grading of color fundus images is slow and depends on the availability of trained specialists, automated screening tools are needed. This study proposes a lightweight channel-wise refinement strategy for automatic five-class DR grading, built on a ResNet50 backbone. Two custom blocks are evaluated: CBR, which applies a 3 × 3 convolution, batch normalization, and a ReLU activation to make the channel representation more compact, and CBS, which applies a 3 × 3 convolution, batch normalization, and a SiLU activation to reinforce local spatial features. On the Diabetic Retinopathy Balanced dataset, the baseline ResNet50 reached an accuracy of 90.77%, a precision of 90.60%, a recall of 90.79%, and an F1-score of 90.64%. In the ablation study, the best configuration was ResNet50 + CBR, with an accuracy of 91.85%, a precision of 91.75%, a recall of 91.88%, and an F1-score of 91.76%. The full CBR-CBS Hybrid ResNet50 was close behind, with an accuracy of 91.81% and an F1-score of 91.71%. The CBR block accounts for most of this improvement, which suggests that channel-wise refinement helps the model separate subtle lesion patterns. These results establish lightweight channel-wise refinement (CBR) as an effective, compact, and interpretable enhancement of ResNet50 for automated five-class DR grading, delivering a consistent multi-metric gain over the baseline and accuracy competitive with the literature, which makes it a promising solution for large-scale screening. Full article
(This article belongs to the Section Medical & Healthcare AI)
Show Figures

Figure 1

22 pages, 1659 KB  
Article
FroLineR: Front-Line Response with Retrieval-Augmented Prompt-Engineered Reply Generation for IT Help Desks
by Alexandru Dima, Maria-Elena Mihăilescu, Darius Mihai, Mihai Carabaș and Mihai Dascalu
AI 2026, 7(8), 317; https://doi.org/10.3390/ai7080317 - 19 Aug 2026
Viewed by 385
Abstract
IT help desks at large organizations face a high volume of recurrent, well-documented user requests that nevertheless require human-written replies, creating a persistent staff workload that is repetitive in content but non-trivial in tone and procedural correctness. We present FroLineR, short for Front-Line [...] Read more.
IT help desks at large organizations face a high volume of recurrent, well-documented user requests that nevertheless require human-written replies, creating a persistent staff workload that is repetitive in content but non-trivial in tone and procedural correctness. We present FroLineR, short for Front-Line Response, a system that drafts the initial staff reply to such tickets in the login and account-activation category and integrates into a human-in-the-loop ticketing workflow on a Romanian-language ticketing platform. The generator is an unmodified instruct model augmented with retrieval from a small set of hand-curated guide documents, using a Romanian system prompt refined over several rounds of staff review. To evaluate and refine the prompt without manual labeling, we cluster the first user message of every historical thread with both BERTopic and Semantic Signal Separation (S3), score configurations along coherence and lexical-diversity axes, and extract a 200-message evaluation set from the winning model. Prompt convergence was certified by several rounds of manual review by support staff. The production system is quantized to Q4_K_M GGUF, served through llama-cpp-python behind a small Flask API, and deployed with GPU offloading on the target server, reducing end-to-end per-answer latency from approximately 830 s on the server’s CPU to roughly 61 s once layers are offloaded to the GPU, with no observable degradation in answer quality. Full article
Show Figures

Figure 1

18 pages, 6309 KB  
Article
SmartMM: A Domain-Specific Large Language Model for Medical Microbiology
by Yongqiang Gong, Ruiqi Ma, Xicheng Wang, Ruixi Li, Han Dong, Yijin Liu, Xi Peng, Quanle Guo and Yin Liu
AI 2026, 7(8), 316; https://doi.org/10.3390/ai7080316 - 18 Aug 2026
Viewed by 294
Abstract
Background: Large language models (LLMs) show considerable promise for medical question answering and reasoning. Their use in medical microbiology, however, remains constrained by limited domain-specific knowledge and the risk of hallucinated outputs. Objective: To develop and evaluate Smart Medical Microbiology (SmartMM), a specialized [...] Read more.
Background: Large language models (LLMs) show considerable promise for medical question answering and reasoning. Their use in medical microbiology, however, remains constrained by limited domain-specific knowledge and the risk of hallucinated outputs. Objective: To develop and evaluate Smart Medical Microbiology (SmartMM), a specialized LLM for accurate, reliable, and context-aware responses in medical microbiology. Methods: SmartMM integrates domain-adaptive continual pretraining, supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), knowledge distillation, and retrieval-augmented generation (RAG). We constructed a high-quality microbiology corpus from textbooks, clinical guidelines, the scientific literature, case reports, and other authoritative sources. Model performance was assessed using objective examinations, subjective generation tasks, expert review, and real-world user preference evaluation. Results: SmartMM achieved accuracies of 0.897 and 0.563 on true-or-false and fill-in-the-blank questions, respectively. In subjective generation tasks, it obtained the highest ROUGE-L score (0.265) and BERTScore F1 score (0.771) among all compared models. Expert assessment showed excellent inter-rater reliability, with all ICC(C,3) values exceeding 0.970. In a user evaluation involving 20 participants and 100 real-world questions, SmartMM received the largest number of first-place rankings (33), placing it among the top-performing systems overall. Conclusions: SmartMM showed strong domain adaptability in medical microbiology knowledge organization, semantic generation, and retrieval-augmented reasoning. These findings support its potential use in educational support, infectious disease knowledge assistance, and retrieval-enhanced medical question answering. Full article
Show Figures

Figure 1

24 pages, 9905 KB  
Article
Artificial Intelligence Framework for Respiratory Disease Classification Using Multi-Spectral-Feature-Driven and Deep Neural Architectures
by Vijayalakshmi Sankaran, Paramasivam Alagumariappan, Sumendra Yogarayan, Thayananth Caran Varshana and Balaguru Ramana
AI 2026, 7(8), 315; https://doi.org/10.3390/ai7080315 - 18 Aug 2026
Viewed by 389
Abstract
Globally, respiratory diseases such as asthma, chronic obstructive pulmonary disease (COPD) and pneumonia affect populations significantly, requiring early and accurate diagnosis for effective clinical management. Manual auscultation and expert interpretation are the common shortcomings in conventional diagnostic approaches, as they lead to time-consuming [...] Read more.
Globally, respiratory diseases such as asthma, chronic obstructive pulmonary disease (COPD) and pneumonia affect populations significantly, requiring early and accurate diagnosis for effective clinical management. Manual auscultation and expert interpretation are the common shortcomings in conventional diagnostic approaches, as they lead to time-consuming and inconsistent analysis. To address these limitations, an artificial intelligence-driven framework for respiratory disease classification using multi-spectral feature extraction and deep learning architectures is proposed to classify four different respiratory conditions: Asthma, COPD, Pneumonia and Healthy. The dataset is collected from Kaggle’s respiratory sound database and the COUGHVID V3 database, which together contain 322 Asthma signals, 746 COPD signals, 323 Pneumonia signals and 174 Healthy signals. Subsequently, the features are extracted using four different feature extraction techniques—Constant Q Transform (CQT), a Gammatone spectrogram, Mel-Frequency Cepstral Coefficients (MFCC) and Perceptual Linear Prediction (PLP)—and these extracted spectral representations are provided as inputs to various deep learning models such as a Deep Convolutional Neural Network (Deep CNN), a Temporal Attention Network (TAN) and an Autoencoder for automated feature learning and disease classification. The proposed framework is evaluated using several performance metrics, and the experimental results clearly indicate that the performance of the proposed classification framework strongly depends on the selection of spectral feature extraction techniques and deep learning models. Among all the evaluated combinations, it is evident that the Autoencoder model integrated with CQT features exhibited the best classification performance, with an accuracy of 98.72%, precision of 98.74%, recall of 98.72%, Matthews correlation coefficient (MCC) of 98.11%, Cohen’s kappa value of 98.10% and the least log loss of 0.025. The proposed artificial intelligence (AI)-enabled respiratory disease classification framework has demonstrated the ability to produce a reliable computer-aided diagnostic system which is suitable for smart healthcare applications and automated pulmonary disease screening. Full article
Show Figures

Figure 1

18 pages, 2540 KB  
Review
Application of Artificial Intelligence in Perinatal Mental Health: A Review
by Sheikh Mohammed Shariful Islam, Alan W. Gemmill, Yafit Hirshler, Michaela Pascoe and Jeannette Milgrom
AI 2026, 7(8), 314; https://doi.org/10.3390/ai7080314 - 14 Aug 2026
Viewed by 648
Abstract
Perinatal mental health remains a critical global challenge, with maternal mortality, preterm birth, and persistent disparities in care contributing to adverse outcomes for mothers. In addition, mental health difficulties in the perinatal period are associated with poorer developmental outcomes for young children and [...] Read more.
Perinatal mental health remains a critical global challenge, with maternal mortality, preterm birth, and persistent disparities in care contributing to adverse outcomes for mothers. In addition, mental health difficulties in the perinatal period are associated with poorer developmental outcomes for young children and impose an economic burden on societies. Addressing these issues requires innovative approaches that can complement traditional clinical practices. Artificial intelligence (AI) has emerged as a powerful tool with the potential to transform perinatal care by enabling early risk prediction, personalised interventions, and scalable support systems. However, there are no existing reviews on use of AI across different stages of perinatal mental health. We conclude with a call to action for clinicians, researchers, policymakers, and technology developers to collaborate on a consensus framework that ensures ethical, safe, and equitable integration of AI into perinatal care. Full article
(This article belongs to the Special Issue Digital Health: AI-Driven Personalized Healthcare and Applications)
Show Figures

Figure 1

21 pages, 8105 KB  
Article
Bidirectional Cross-Level Feature Interaction and Context-Aware Multi-Scale Attention for Crowd Counting
by Zhifan Jin, Lin Zhou, He Wang, Sijia Chen, Liman Liu and Wenbing Tao
AI 2026, 7(8), 313; https://doi.org/10.3390/ai7080313 - 13 Aug 2026
Viewed by 345
Abstract
Crowd counting estimates the number and spatial distribution of people in images and videos, supporting smart city management and public safety. Existing methods often rely on intra-level feature refinement and simple cross-scale fusion, such as concatenation or addition, which limits interaction between fine-grained [...] Read more.
Crowd counting estimates the number and spatial distribution of people in images and videos, supporting smart city management and public safety. Existing methods often rely on intra-level feature refinement and simple cross-scale fusion, such as concatenation or addition, which limits interaction between fine-grained spatial details and high-level semantic representations. In addition, the limited receptive field of convolutional networks restricts global context modeling in scenes with heavy occlusion and extreme scale variation. To address these challenges, we propose a Hierarchical Context-Aware Multi-Scale Attention Network (HCMA). Its bidirectional cross-level interaction is realized through two complementary top-down decoding streams, where an attention-gating stream provides spatial guidance for the counting-oriented representations carried by a density-feature stream. HCMA includes three modules: the Selective Context-Aware Attention Module (SCAM), which performs context-dependent multi-scale filtering; Dynamic Positional Pooling (DPP), which introduces an image-level mean token and stochastic global-relation aggregation; and the Multi-Scale Enhancement Attention Module (MSEA), which refines high-level semantic features under scale variation. Experiments on ShanghaiTech, UCF-QNRF, and NWPU-Crowd show competitive counting accuracy across scenes with different density ranges, scale variation, and occlusion. In particular, HCMA achieves an MAE of 73.2 on NWPU-Crowd, 17.2% lower than that of DM-Count. Full article
(This article belongs to the Special Issue AI and Computer Vision in Real-World and Industrial Applications)
Show Figures

Figure 1

72 pages, 12684 KB  
Article
iCert-Fair: A Human-Preference-Guided Two-Layer Framework for Multi-Objective Fairness Assessment and Harm Recovery in Credit Scoring
by Rashed Bahlool and Nabil Hewahi
AI 2026, 7(8), 312; https://doi.org/10.3390/ai7080312 - 13 Aug 2026
Viewed by 396
Abstract
As regulatory requirements increasingly shape automated lending decisions, fairness remains a critical challenge in high-stakes domains, particularly credit scoring. Although artificial intelligence models can achieve strong predictive performance, they may also reproduce biased outcomes that reduce financial inclusion or transfer harm to overlooked [...] Read more.
As regulatory requirements increasingly shape automated lending decisions, fairness remains a critical challenge in high-stakes domains, particularly credit scoring. Although artificial intelligence models can achieve strong predictive performance, they may also reproduce biased outcomes that reduce financial inclusion or transfer harm to overlooked protected groups. Existing fairness interventions commonly operate at a single stage of the decision-making pipeline, despite bias often propagating across representational and decision layers. This study proposes iCert-Fair, a two-layer framework for technical fairness assessment and harm recovery in credit scoring. The first layer adopts a fairness-through-explainability paradigm, using SHAP-based explanations to identify direct and proxy dependence on protected attributes and guide structural dataset repair, while the second layer applies targeted threshold-policy adjustments to recover residual harm while preserving decision utility. Experiments on the German and Taiwanese credit datasets show that fairness gains are model- and dataset-specific and may be collective, concentrated, transferred, or recovered unevenly across protected attributes. The direct comparison with representative pre-processing, in-processing, and post-processing methods revealed that baseline methods targeting one protected attribute at a time frequently transferred residual harm to other monitored attributes. In contrast, the fairness-focused recommendations generated by iCert-Fair achieved larger collective fairness improvements across all considered protected attributes while avoiding residual harm. These gains were obtained while preserving predictive utility on the German dataset and with utility degradation remaining below 5% across the evaluated performance metrics on the Taiwanese dataset, alongside consistently lower false-negative risk. The empirical findings support the use of complementary structural and policy-level interventions and demonstrate the importance of jointly evaluating aggregate disparity, worst-case attribute-level harm, cross-attribute transfer, and predictive utility. Full article
(This article belongs to the Special Issue Human-Computer Interaction and Human-Centered AI)
Show Figures

Figure 1

49 pages, 669 KB  
Article
Discourse Structure as an Interpretable Signal for Detecting Hallucinated Chain-of-Thought Reasoning in Large Language Models
by Boris Galitsky
AI 2026, 7(8), 311; https://doi.org/10.3390/ai7080311 - 11 Aug 2026
Viewed by 849
Abstract
Large language models can generate fluent chain-of-thought (CoT) reasoning that appears coherent while exhibiting systematic distortions in evidence weighting and hypothesis comparison. This paper studies hallucinated CoT as a discourse-structural phenomenon, not only a factual one. We introduce a diagnostic reasoning benchmark with [...] Read more.
Large language models can generate fluent chain-of-thought (CoT) reasoning that appears coherent while exhibiting systematic distortions in evidence weighting and hypothesis comparison. This paper studies hallucinated CoT as a discourse-structural phenomenon, not only a factual one. We introduce a diagnostic reasoning benchmark with paired grounded and hallucinated explanations, where traces differ in how they organize evidence, alternatives, and defeaters. We extract discourse tree features that summarize evidence allocation, contrast preservation, commitment timing, and evidence integration, and combine them with the Joint Knowledge–Reasoning Hallucination Measure (JKRHM). Experiments on the synthetic diagnostic dataset and preliminary external validation on HaluBench suggest that discourse structure provides an interpretable signal for detecting reasoning hallucinations and complements existing factuality and uncertainty-based hallucination detectors. Because the HaluBench reasoning rationales are generated as an intermediate representation, these results should not be interpreted as definitive external proof of generalization. The results support a cautious conclusion: discourse analysis does not replace factual verification, but it helps expose reasoning paths that are structurally unsupported even when they are fluent and persuasive. Full article
Show Figures

Figure 1

27 pages, 6435 KB  
Article
Investigation into the Spectral Completion Algorithm Leveraging Dense Connection Autoencoders
by Yepeng Shi, Shengliang Fang, Shunhu Hou, Yuhai Li, You Fu and Qichen Wang
AI 2026, 7(8), 310; https://doi.org/10.3390/ai7080310 - 11 Aug 2026
Viewed by 338
Abstract
Radio Environment Map (REM) construction is frequently constrained by sparse and unevenly distributed spectrum measurements. While existing completion methods primarily target Power Spectral Density (PSD) data under random missing patterns, the reconstruction of Reference Signal Received Power (RSRP) maps under structured data loss [...] Read more.
Radio Environment Map (REM) construction is frequently constrained by sparse and unevenly distributed spectrum measurements. While existing completion methods primarily target Power Spectral Density (PSD) data under random missing patterns, the reconstruction of Reference Signal Received Power (RSRP) maps under structured data loss remains underexplored. This study addresses this gap by proposing a fully convolutional densely connected autoencoder(AE) for RSRP map completion. The encoder stacks dense blocks and transition layers, a bottleneck preserves the latent representation, and the decoder restores spatial resolution through transposed convolution. Both global and local skip connections are incorporated to fuse large-scale structure with fine-grained details. A composite loss function supervises observed and missing regions separately, which preserves the fidelity of known measurements while improving inference over unobserved grid points. Experiments on the public DeepREM dataset under random, spatial, and strip-wise missing patterns show that the method achieves the best or comparable completion accuracy in most tested settings, with the most pronounced performance gains over mainstream baselines under the challenging spatial block-missing case. Full article
Show Figures

Figure 1

28 pages, 2619 KB  
Article
AI as a Practice Partner: A Feasibility Study of MentaClassAI, a Conversational LLM Tool for Training Educators’ Mentalizing Responses to Child Dysregulation
by Gali Chelouche-Dwek and Peter Fonagy
AI 2026, 7(8), 309; https://doi.org/10.3390/ai7080309 - 8 Aug 2026
Viewed by 478
Abstract
Background: Teachers routinely encounter children whose behaviour reflects emotional distress and dysregulation, yet they have limited opportunities to practise the relational skills required to respond effectively. These challenges are particularly pronounced in Alternative Provision (AP), which serves children who frequently present with histories [...] Read more.
Background: Teachers routinely encounter children whose behaviour reflects emotional distress and dysregulation, yet they have limited opportunities to practise the relational skills required to respond effectively. These challenges are particularly pronounced in Alternative Provision (AP), which serves children who frequently present with histories of trauma, neurodevelopmental differences, and complex emotional and behavioural needs. Mentalization, the capacity to understand behaviour in terms of underlying mental states, is central to effective relational practice in such contexts. Conversational Artificial Intelligence (AI) may offer a scalable means of supporting this form of skills development, but its feasibility as a teacher-training modality remains largely unexplored. Methods: This mixed-methods proof-of-concept feasibility study evaluated MentaClassAI, a novel AI-based training tool in which educators engaged in simulated voice conversations with AI child characters portraying classroom dysregulation and subsequently received individualised, mentalization-informed feedback. Eleven staff members from a single AP school (four teachers and seven teaching assistants) completed a single training session and were allocated to either a psychoeducation video condition (n = 6) or a no-video condition (n = 5). The video condition received a brief introduction to mentalization and epistemic trust prior to engaging with the simulation. Pre- and post-engagement measures included the Reflective Functioning Questionnaire (RFQ-8) and a Teacher Self-Efficacy Scale. Post-engagement measures included an 18-item acceptability questionnaire, a Technology Acceptance Model scale, and open-ended questions analysed using thematic analysis. Results: Acceptability was high, with 84.8% of questionnaire responses falling within the positive range (overall M = 5.60/7). Feedback accuracy (M = 6.55) and clarity (M = 6.36) received the highest ratings. Participants reported higher teacher self-efficacy after the session than before (d = 1.20, p = 0.003), with 10 of 11 participants demonstrating improvement. Self-reported hypomentalizing was lower after the session (d = −0.86, p = 0.017). Between-condition differences (video versus no-video) were not statistically significant. The video condition scored numerically higher on the directional indicators. Qualitative analysis identified five themes: the value of consequence-free rehearsal; the specificity and usefulness of feedback; appreciation of the focus on the child’s emotional experience; limitations in the ecological diversity of AI child characters; and a desire for more naturalistic interaction. Conclusions: These findings provide preliminary support for the feasibility and acceptability of AI-based mentalization practice for AP staff. The principal value of the tool appears to lie not only in the simulation itself but in the quality of the reflective feedback generated. Although based on a small sample, the observed pre–post changes provide an encouraging signal that may justify a controlled trial. The contribution of pre-session psychoeducation to training outcomes remains an important question for future research. Full article
Show Figures

Figure 1

31 pages, 924 KB  
Article
Reinforcement Learning for Warehouse Management Using a Scenario-Based Simulation Testbed
by Laura Acosta García, Julen Cestero Portu, Ander García Gangoiti and Marco Quartulli
AI 2026, 7(8), 308; https://doi.org/10.3390/ai7080308 - 8 Aug 2026
Viewed by 634
Abstract
Warehouse operations involve dynamic item flows, fluctuating demand, and heterogeneous layouts, making adaptive decision-making essential for efficient storage and order fulfillment. In this context, reinforcement learning (RL) provides a promising approach for learning adaptive warehouse control policies under stochastic environments. However, evaluating RL-based [...] Read more.
Warehouse operations involve dynamic item flows, fluctuating demand, and heterogeneous layouts, making adaptive decision-making essential for efficient storage and order fulfillment. In this context, reinforcement learning (RL) provides a promising approach for learning adaptive warehouse control policies under stochastic environments. However, evaluating RL-based solutions in real warehouse settings is often costly and time-consuming, motivating the need for realistic and reproducible simulation environments. In this paper, we introduce a configurable warehouse simulation environment modeling stochastic item arrivals, order generation, and internal logistics operations across diverse layouts and workload conditions. Based on this environment, we construct a reproducible experimental testbed composed of multiple scenarios ranging from low-load to highly congested settings. The testbed is publicly released to support reproducible research and comparative evaluation within the research community. We formulate the warehouse management problem as a Markov decision process (MDP) and apply a Maskable Proximal Policy Optimization (Maskable PPO) agent to learn adaptive control policies. The RL-based approach is evaluated across the defined scenarios and compared against heuristic baseline strategies. Experimental results show that the proposed solution achieves performance comparable to a strong greedy first-in, first-out (FIFO) heuristic while improving order fulfillment by up to 13.5 percentage points under challenging workload conditions. These results demonstrate the ability of RL to learn robust warehouse control policies that adaptively optimize performance and maintain operational stability across a wide spectrum of distinct scenarios. Full article
Show Figures

Figure 1

26 pages, 1890 KB  
Article
Beyond Aggregate Sentiment: Machine Learning-Driven Discourse Indicators for AI News at Scale
by Oleksandra Topal, Inna Novalija, Joao Pita Costa and Dumitru Roman
AI 2026, 7(8), 307; https://doi.org/10.3390/ai7080307 - 7 Aug 2026
Viewed by 434
Abstract
This study deploys a scalable machine learning pipeline: combining a transformer-based classifier applied to 2.01 million English-language AI-related news headlines (July 2022–July 2024) with large-language-model and human-annotator validation (three annotators, Fleiss’ κ=0.80) on stratified subsamples, to extract six interpretable, bias-linked [...] Read more.
This study deploys a scalable machine learning pipeline: combining a transformer-based classifier applied to 2.01 million English-language AI-related news headlines (July 2022–July 2024) with large-language-model and human-annotator validation (three annotators, Fleiss’ κ=0.80) on stratified subsamples, to extract six interpretable, bias-linked discourse indicators computed at the AI-domain level: evaluative orientation (valence), loss salience, narrative drift, exposure-adjusted sentiment, cross-source divergence, and novelty-phase framing. Each operationalizes an established cognitive-psychology construct as a computable property of the information environment associated with biased risk–benefit reasoning. Results show systematic variation across domains: technical and methodological areas such as deep learning and natural language processing exhibit gain-salient framing, while safety-critical topics such as deepfakes (loss-to-gain headline ratio = 3.17) and facial recognition show strongly loss-salient profiles. Cross-model validation using an LLM on a stratified sample of 1000 headlines confirms that domain-level indicator rankings are robust to classifier choice (Spearman ρ=0.83; p<0.001), establishing the rank stability of pipeline outputs independently of the specific classification architecture. As a contextual application, domain-level profiles are mapped to European Union AI governance instruments, documenting parallels between discourse patterns and regulatory risk tiers. The framework provides a scalable, reproducible methodology for monitoring evaluative conditions in technology news across domains, sources, and time. Full article
Show Figures

Figure 1

15 pages, 3314 KB  
Article
Diagnostic Performance of an Artificial Intelligence Cervical Spine Fracture Decision Support System at a Non-Trauma Community Hospital Setting
by Genaro Herrera Cano, Michal Dyrda, Youssef Beshay, David Baltrusaitis, Mitch Paro, Rafael Olivieri-Ortiz, Grigoriy Androsov, Antonio Medina Luna and Michael Baldwin
AI 2026, 7(8), 306; https://doi.org/10.3390/ai7080306 - 7 Aug 2026
Viewed by 508
Abstract
Traumatic cervical spine fractures (CSFxs) require timely diagnosis due to associated morbidity. Artificial intelligence (AI)-based decision support systems have been proposed to improve imaging workflow efficiency. However, their performance in non-trauma settings remains unclear. This study evaluated the diagnostic performance of the AIDOC [...] Read more.
Traumatic cervical spine fractures (CSFxs) require timely diagnosis due to associated morbidity. Artificial intelligence (AI)-based decision support systems have been proposed to improve imaging workflow efficiency. However, their performance in non-trauma settings remains unclear. This study evaluated the diagnostic performance of the AIDOC decision support system (DSS) for detecting CSFxs in a non-trauma academic community hospital using a retrospective analysis of 1812 cervical spine CT scans, with radiologist interpretation as the reference standard. Sensitivity, specificity, accuracy, positive predictive value (PPV), and negative predictive value (NPV) were calculated for fracture detection and heatmap-based localization. The AI system demonstrated a sensitivity of 72.2% and specificity of 98.1%, with an accuracy of 97.9%. In the context of low fracture prevalence (0.99%), PPV was low (27.7%), while NPV was high (99.7%). Heatmap-based localization showed reduced sensitivity (43.8%) despite high specificity (97.5%). These findings demonstrate high specificity and NPV, with lower sensitivity for localization and low PPV in a low-prevalence setting. Prospective multi-institutional studies are required to further validate these diagnostic performance metrics and assess generalizability across diverse clinical settings and imaging protocols. Full article
(This article belongs to the Section Medical & Healthcare AI)
Show Figures

Figure 1

31 pages, 5510 KB  
Article
EcoSortBin: Accuracy–Generalisation Trade-Offs in Open-Vocabulary Campus Waste Detection on Raspberry Pi 4
by Madhini Balasundaram and Supraja Perumal
AI 2026, 7(8), 305; https://doi.org/10.3390/ai7080305 - 7 Aug 2026
Viewed by 614
Abstract
Waste management on university campuses is complicated by the constant change in packaging types, which existing waste-sorting systems cannot recognise unless they are retrained. Open-vocabulary object detectors can identify objects from text descriptions instead of a fixed list of categories, offering a possible [...] Read more.
Waste management on university campuses is complicated by the constant change in packaging types, which existing waste-sorting systems cannot recognise unless they are retrained. Open-vocabulary object detectors can identify objects from text descriptions instead of a fixed list of categories, offering a possible solution to this problem. However, it is not known how well this ability survives when such a detector is fine-tuned and deployed on low-power hardware. This paper presents EcoSortBin, a waste-sorting system built on the YOLOE-26 detector and deployed on a Raspberry Pi 4. YOLOE-26 was first fine-tuned on a 1330-image campus waste dataset covering seven classes, with masks generated using the Segment Anything Model, establishing a baseline called WasteYOLOE26-S with 74% top-1 accuracy on known classes; however, this fine-tuning reduces the model’s ability to recognise the same seven classes when they appear in a different dataset or setting. RLPA (RepRTA-Compatible LoRA Prompt Adapters) addresses this by adapting only the text-embedding component of the model using a small set of additional parameters (16,384 parameters, rank 16), leaving the rest of the network unchanged; this restores cross-domain generalisation but reduces top-1 accuracy on known classes to only 15%, which is too low for practical use. To recover this accuracy without losing cross-domain generalisation, frozen-backbone neck fine-tuning (NeckFT) was added, which fine-tunes the feature-combining layers of the network while keeping the main backbone frozen, preserving its pretrained visual–text alignment. Combining RLPA with NeckFT achieved the best balance of the three approaches, with 73.5% top-1 accuracy and a Cross-Domain Generalisation Ratio (CDGR) of 0.2435. To test whether this ability extends to genuinely new categories, the model was further tested on 28 novel categories not seen during training, totalling 840 images. RLPA + NeckFT showed consistent zero-shot generalisation to novel objects with container-like shapes, such as bottles and jars. After quantisation for edge deployment, the model kept its full accuracy ranking and produced a compact 41.8 MB file suitable for the Raspberry Pi 4. These results show that RLPA + NeckFT gives a practical balance of accuracy and generalisation for campus waste detection on low-power hardware. Full article
(This article belongs to the Section AI in Autonomous Systems)
Show Figures

Figure 1

27 pages, 7413 KB  
Article
GINet-DGC: Structural Inductive Biases and Dynamic Generalization Control for High-Dimensional Small-Sample Tabular Data
by Xinran Zhang, Yang Sheng, Sijie Shen, Dongjie Fan and Lizhuang Liu
AI 2026, 7(8), 304; https://doi.org/10.3390/ai7080304 - 6 Aug 2026
Viewed by 393
Abstract
Learning from high-dimensional, low-sample-size (HDLSS) data remains a persistent challenge in machine learning, as models must infer reliable patterns from limited observations while handling an excessive number of variables—a scenario particularly prevalent in biomedical applications. Such data structures render predictive modeling highly vulnerable [...] Read more.
Learning from high-dimensional, low-sample-size (HDLSS) data remains a persistent challenge in machine learning, as models must infer reliable patterns from limited observations while handling an excessive number of variables—a scenario particularly prevalent in biomedical applications. Such data structures render predictive modeling highly vulnerable to erratic optimization and overfitting. To address this challenge, we propose the Global Interaction Network with Dynamic Generalization Control (GINet-DGC), an artificial intelligence (AI) framework that integrates feature-wise structural priors with dynamic generalization monitoring. Rather than directly learning an unconstrained first-layer weight matrix, GINet-DGC generates task-specific weights from multi-view feature descriptors, encompassing latent semantic, global distributional, local topological, and hierarchical representations. This structure-constrained weight generation strategy effectively narrows the feature-interaction search space and acts as an inductive regularizer against noise and redundant molecular features. Furthermore, we introduce an Overfitting-aware Index (OFI) to monitor the training trajectory and effectively identify the generalization saturation point for adaptive termination. Empirical evaluations on eight public real-world biomedical HDLSS gene-expression datasets, using a repeated stratified 5 × 5 cross-validation protocol, demonstrate that GINet-DGC achieves competitive and stable performance against 17 baselines. These findings support the effectiveness of the proposed framework within the evaluated public biomedical HDLSS benchmark setting. Full article
(This article belongs to the Special Issue AI in Bioinformatics: The Next Frontier in Health Discovery)
Show Figures

Figure 1

26 pages, 6488 KB  
Article
Socratic Mediation Patterns in AI–Student Interactions: A Content Analysis of a Conversational Agent in Distance Higher Education
by Camilo Aurelio Velandia, Nelson Iván Bedoya, Andrés Chiappe and David Muñoz-Ballier
AI 2026, 7(8), 303; https://doi.org/10.3390/ai7080303 - 6 Aug 2026
Viewed by 665
Abstract
This study identifies and characterises the Socratic mediation patterns enacted by MIA, an AI-based conversational agent used in distance higher education. A deductive content analysis was conducted on 737 conversations using six categories: exploration of prior knowledge, contextual adjustment, linkage to experiences, autonomy-oriented [...] Read more.
This study identifies and characterises the Socratic mediation patterns enacted by MIA, an AI-based conversational agent used in distance higher education. A deductive content analysis was conducted on 737 conversations using six categories: exploration of prior knowledge, contextual adjustment, linkage to experiences, autonomy-oriented prompts, dialogic progression, and verification prompts. The categorical framework achieved full expert content-validity agreement (S-CVI/Ave = 1.00). Contextual Adjustment (77%), Verification Prompts (76%), and Autonomy-Oriented Prompts (74%) were the most frequently observed categories. Sixty of the 64 theoretically possible category combinations occurred in the corpus, and Linkage to Experiences appeared more frequently in personal conversations (33.7%) than in academic conversations (18.3%). The distribution of categories also varied according to conversation length, with longer exchanges containing a broader range of coded dialogic moves. These findings describe the conversational repertoire through which MIA operationalised features associated with Socratic mediation. Because the study did not include independent measures of student satisfaction, learning, engagement, or self-regulation, the results should not be interpreted as evidence of educational effectiveness or causal effects. The study contributes an operational framework for analysing Socratic features in AI–student interactions and identifies directions for outcome-based research. Full article
(This article belongs to the Topic AI Trends in Teacher and Student Training)
Show Figures

Figure 1

22 pages, 38201 KB  
Article
ACBE-CroFuseNet: An Optical and SAR Cross-Fusion Semantic Segmentation Network for Paddy Rice Extraction
by Xinru Guo and Linze Bai
AI 2026, 7(8), 302; https://doi.org/10.3390/ai7080302 - 6 Aug 2026
Viewed by 485
Abstract
Accurate mapping of paddy rice is essential for agricultural monitoring, yield estimation, and food security assessment. However, optical imagery is often affected by clouds and spectral confusion, while SAR imagery suffers from speckle noise and weak spatial detail representation. Simple optical and SAR [...] Read more.
Accurate mapping of paddy rice is essential for agricultural monitoring, yield estimation, and food security assessment. However, optical imagery is often affected by clouds and spectral confusion, while SAR imagery suffers from speckle noise and weak spatial detail representation. Simple optical and SAR feature concatenation is therefore insufficient for complex agricultural landscapes. To address these limitations, this study proposes ACBE-CroFuseNet, an optical and SAR cross-fusion semantic segmentation network for paddy rice extraction using Sentinel-1 SAR and Sentinel-2 optical imagery in Yancheng, Jiangsu Province. ACBE-CroFuseNet introduces two task-oriented designs for paddy rice mapping. First, an attention cross-fusion module is developed to adaptively model modality contributions and spatial responses between optical spectral–textural features and SAR scattering–structural features. Second, a boundary enhancement module with boundary supervision is introduced to strengthen the delineation of fragmented paddy fields and field edges. Multimodal feature aggregation and multi-scale deep supervision are further used to improve feature utilization and segmentation stability. Compared with UNet++, Swin-Unet, CroFuseNet, and CMFFNet under five-fold cross-validation, ACBE-CroFuseNet achieves the best overall performance. The extracted paddy rice area in Yancheng in 2025 demonstrates the applicability of the proposed method for large-scale crop mapping. Full article
(This article belongs to the Special Issue AI-Powered Remote Sensing for Agriculture)
Show Figures

Figure 1

13 pages, 6422 KB  
Article
Towards Automating Junctional Hemorrhage Control Using AI for Interpretation of Human Tissue
by Sofia I. Hernandez Torres, Jennifer Achay, Scotty Bolleter, James A. Bynum and Eric J. Snider
AI 2026, 7(8), 301; https://doi.org/10.3390/ai7080301 - 4 Aug 2026
Viewed by 463
Abstract
Junctional hemorrhage has a high fatality rate due to how difficult it is to control rapid bleeding from major vessels. The available methods to stop junctional blood loss are prone to placement errors as well as failure during transport and during prolonged field [...] Read more.
Junctional hemorrhage has a high fatality rate due to how difficult it is to control rapid bleeding from major vessels. The available methods to stop junctional blood loss are prone to placement errors as well as failure during transport and during prolonged field care. On the battlefield, medical imaging with a portable ultrasound can be leveraged for visualization of the underlying tissue and application of compression at the anatomical junction to effectively stop blood flow. In this work, we developed AI models for anatomical landmark tracking using a perfused human cadaver model. These AI models were paired with an end-user clinical application to guide proper placement and compression, improving junctional hemorrhage control on the future battlefield. The trained U-Net semantic segmentation model demonstrated strong performance across predictions for both validation and hold-out, blind subjects. Overall pixel accuracy across the dataset was 98.9% for training and 98.6% for blind subjects. The artery and vein predictions achieved the highest class-specific training intersection-over-union scores, both at 0.73. This segmentation model trained to interpret human tissue provides evidence that ultrasound visualization can help guide compression at anatomical junctions. Future work will focus on improving blind performance for implementation of this AI model into closed-loop control of hardware prototypes, delivering real-time predictions and control. Full article
(This article belongs to the Special Issue Applications of Artificial Intelligence in Medicine)
Show Figures

Figure 1

16 pages, 1328 KB  
Article
DR-Transformer: A Dual-Regularized Transformer Combining Sparse Attention and Supervised Contrastive Learning for Interpretable Stress Detection in Social Media Text
by Mehdi Chrifi Alaoui, Nour-Eddine Joudar and Mohamed Ettaouil
AI 2026, 7(8), 300; https://doi.org/10.3390/ai7080300 - 4 Aug 2026
Viewed by 507
Abstract
Automatic detection of stress in social media text holds promise for supporting digital mental health, but most existing Transformer-based approaches are opaque and computationally demanding. This work presents DR-Transformer, a Dual-Regularized Transformer that combines two complementary mechanisms: (i) a group sparsity penalty ( [...] Read more.
Automatic detection of stress in social media text holds promise for supporting digital mental health, but most existing Transformer-based approaches are opaque and computationally demanding. This work presents DR-Transformer, a Dual-Regularized Transformer that combines two complementary mechanisms: (i) a group sparsity penalty (L2,1/L2 elastic net) applied to the query and key projection matrices of every attention head, which encourages whole-row sparsity, producing more concentrated and inspectable attention patterns; (ii) a supervised contrastive loss on the [CLS] projection, which organizes the latent space according to the stress label. The architecture is intentionally lightweight (six layers, eight heads, 256-dim embeddings; ∼9.5 M parameters) and runs entirely on consumer-grade hardware (NVIDIA GTX 1660, 6 GB). Experiments on the publicly available Dreaddit dataset (binary stress classification, 2838 train/715 test segments) compare DR-Transformer against Logistic Regression, BiLSTM, a Standard Transformer of identical architecture, and MentalBERT. Across five seeded runs, DR-Transformer (Full) reaches F1=0.876 (bootstrap 95% CI 0.8520.898), outperforming the Standard Transformer (F1=0.842; McNemar p<0.001 with Bonferroni correction) and performing comparably to the much larger MentalBERT (F1=0.879; p=0.421). Sparse regularization increases the fraction of near-zero attention weights (below 0.01) from 0.215 to 0.682, while the supervised contrastive loss improves the silhouette score of [CLS] embeddings from 0.312 to 0.483. Dual regularization thus combines accuracy, efficiency, and structurally induced attention concentration in a single model which can be trained without specialized infrastructure. We use the term “interpretable” throughout in this restricted, structural sense—to refer to concentrated and inspectable attention—rather than in the sense of established causal or mechanistic faithfulness; this is only partially and indirectly supported by our token deletion analysis. Full article
Show Figures

Figure 1

26 pages, 7047 KB  
Review
Embodied Intelligence for Safer Power-System Field Operations: A Critical Review of Technologies, Applications, and Challenges
by Yuxin Wen, Peixiao Fan, Zhiyu Mao, Fang Chi, Chenxuan Zhang and Yuhong Lu
AI 2026, 7(8), 299; https://doi.org/10.3390/ai7080299 - 4 Aug 2026
Viewed by 585
Abstract
Modern power grids require safer and more reliable field operations, yet conventional robots often face limitations in unstructured environments because of rigid pre-programming and weak perception–action coupling. This review examines Embodied Intelligence (EI) as an emerging direction for enhancing power-system field operations. We [...] Read more.
Modern power grids require safer and more reliable field operations, yet conventional robots often face limitations in unstructured environments because of rigid pre-programming and weak perception–action coupling. This review examines Embodied Intelligence (EI) as an emerging direction for enhancing power-system field operations. We first evaluate the environmental adaptability of morphological carriers, including quadrupeds, humanoids, and unmanned aerial vehicles, and then define the perception–cognition–execution closed-loop architecture used in this review. Three application domains are then examined. Intelligent inspection focuses on active perception and potential open-vocabulary object detection. Live-line maintenance emphasizes Sim-to-Real methods and shared autonomy, while disaster-response applications involve heterogeneous air–ground robotic coordination. The review also discusses the potential for EI to reduce human exposure to hazardous tasks and influence labor structures, while a regional text-based proxy illustrates differences in policy attention to digital infrastructure. Finally, we analyze major constraints, including hardware endurance under extreme climates, edge-computing latency, foundation-model uncertainty and hallucination, cybersecurity, and safety certification. Overall, EI should not be interpreted as a mature replacement for current utility practice; it is a developing technological direction whose safe deployment will require field validation, standardized evaluation, cybersecurity assurance, and continued human supervisory authority. Full article
Show Figures

Figure 1

31 pages, 1906 KB  
Review
Agentic AI Safety: A Structured Review of Open Problems and Their Regulatory Anchoring
by Tomáš Valenta, Ondřej Rozinek and Josef Horálek
AI 2026, 7(8), 298; https://doi.org/10.3390/ai7080298 - 4 Aug 2026
Viewed by 1418
Abstract
The shift from passive predictive models to autonomous agents capable of tool use and multi-step planning moves the AI safety landscape from prediction error to control failure: small misjudgements become irreversible actions, and risks compound across long horizons and populations of interacting systems. [...] Read more.
The shift from passive predictive models to autonomous agents capable of tool use and multi-step planning moves the AI safety landscape from prediction error to control failure: small misjudgements become irreversible actions, and risks compound across long horizons and populations of interacting systems. We present a structured review and taxonomy of open scientific problems in agentic AI safety, mapped explicitly onto the EU AI Act and the NIST AI Risk Management Framework. The corpus follows a PRISMA-ScR scoping review, assembled through anchor-based citation chaining and curated reading lists across arXiv, the major machine-learning conferences, and selected security and fairness venues, with a primary March 2026 search cut-off (extended to May 2026 during revision for a small number of high-relevance governance and agentic-safety sources), explicit eligibility criteria, and an analytical distinction between open scientific problems and deployment risks. The taxonomy identifies eight problem families spanning reinforcement-learning policies and language-model planners: goal specification, inner alignment, safe learning and robustness, scalable oversight, interpretability, tool-use security, multi-agent safety, and evaluation and assurance. Mapping these onto the two frameworks shows close alignment for some families and notable absences for others, with multi-agent safety surfacing as a regulatory gap. We add a per-family research roadmap with concrete milestones and a practitioner-facing deployment-posture triage, arguing that progress on inner alignment, interpretability for deceptive-alignment detection, and multi-agent safety would most directly reduce compliance uncertainty. Full article
Show Figures

Figure 1

28 pages, 1439 KB  
Article
Cross-LLM Paraphrase Laundering: A Register-Controlled Evaluation of Fake News Detectors
by Jalal Mehdiyev and Ramiz Aliguliyev
AI 2026, 7(8), 297; https://doi.org/10.3390/ai7080297 - 3 Aug 2026
Viewed by 492
Abstract
Detectors built on transformer language models report near-perfect accuracy on standard fake news benchmarks, which suggests the task is almost solved. We argue that much of this accuracy reflects a confound between writing register and veracity: in common benchmarks, the real class is [...] Read more.
Detectors built on transformer language models report near-perfect accuracy on standard fake news benchmarks, which suggests the task is almost solved. We argue that much of this accuracy reflects a confound between writing register and veracity: in common benchmarks, the real class is human-written while the fake class is machine-generated or machine-rewritten, so a detector can separate the classes by recognizing AI writing style rather than by judging truth. To evaluate this, we designed a two-regime evaluation. Phase 1 is the laundering regime, comparing untouched human-real articles against laundered fake articles. Phase 2 is register-controlled: the real class is passed through the same cross-LLM laundering chains as the fake class, so both classes are read in one machine register, and the register cue is no longer available to the detector. We train seven detectors on three datasets and evaluate each frozen detector under both regimes. Under Phase 1, detectors appear robust; under Phase 2, detection on WELFake collapses from about 99% to about 62% AUROC and the largest models approach chance. The effect is benchmark dependent, large on WELFake, mild on IFND and near zero on GossipCop, and it is confirmed by bootstrap testing with false discovery rate control. We recommend register-controlled evaluation as standard reporting practice. Full article
Show Figures

Figure 1

42 pages, 3016 KB  
Review
AI- and Generative AI-Driven Digital Therapeutics: A Critical Narrative Review of Emerging Evidence
by Daniele Giansanti and Andrea Lastrucci
AI 2026, 7(8), 296; https://doi.org/10.3390/ai7080296 - 3 Aug 2026
Viewed by 886
Abstract
Background: Artificial intelligence-driven digital therapeutics (AI-DTx) are rapidly emerging as a transformative paradigm in healthcare, integrating machine learning, deep learning, and generative AI into digital interventions across diverse clinical domains. Despite rapid growth, the evidence landscape remains fragmented, with heterogeneous methodologies, diverse application [...] Read more.
Background: Artificial intelligence-driven digital therapeutics (AI-DTx) are rapidly emerging as a transformative paradigm in healthcare, integrating machine learning, deep learning, and generative AI into digital interventions across diverse clinical domains. Despite rapid growth, the evidence landscape remains fragmented, with heterogeneous methodologies, diverse application contexts, and limited cross-domain synthesis. Aim: This narrative view aims to provide an evidence-informed narrative synthesis of the available secondary literature on AI-driven digital therapeutics, primarily focusing on systematic reviews and meta-analyses, to identify emerging patterns, cross-cutting trends, and future directions across clinical and technological domains. Methods: A narrative synthesis of secondary evidence was conducted, focusing on 23 systematic reviews, meta-analyses, and relevant review articles addressing AI-driven digital therapeutics. The identified literature was analyzed to explore recurring themes across clinical domains, technological approaches, and implementation challenges. Findings were further contextualized through selected recent randomized controlled trials and translational studies to provide insights into emerging clinical applications and real-world perspectives. Results: Across the available literature, AI-driven digital therapeutics demonstrate a broad and rapidly evolving expansion across mental health, chronic disease management, rehabilitation, and behavioral health. The field is characterized by a progressive shift from static, rule-based interventions toward more adaptive systems supported by machine learning, deep learning, and generative AI. A key emerging theme is the role of AI as an enabling layer for personalization, adaptation, and dynamic intervention delivery rather than as a standalone therapeutic modality. Mental health represents the most extensively studied domain, particularly through conversational agents and cognitive behavioral therapy-informed interventions, while other clinical areas are progressively expanding their translational potential. Persistent challenges include methodological heterogeneity, limited long-term validation, and incomplete integration into routine clinical workflows. Discussion: The current evidence suggests a transition toward hybrid human–AI models of care, in which digital systems may support and augment clinical practice through adaptive and data-driven approaches. However, the field remains characterized by fragmented evidence, evolving evaluation approaches, and challenges related to standardization, validation, and real-world implementation. Conclusions: AI-driven digital therapeutics are evolving toward increasingly adaptive and clinically oriented healthcare solutions. Future progress will depend on improving methodological consistency, strengthening long-term evaluation, and supporting responsible integration into clinical pathways to ensure safe, scalable, and meaningful impact. Full article
(This article belongs to the Special Issue Digital Health: AI-Driven Personalized Healthcare and Applications)
Show Figures

Figure 1

Previous Issue
Next Issue
Back to TopTop