Journal Description
AI
AI
is an international, peer-reviewed, open access journal on artificial intelligence (AI), including broad aspects of cognition and reasoning, perception and planning, machine learning, intelligent robotics, and applications of AI, published monthly online by MDPI.
- Open Access— free for readers, with article processing charges (APC) paid by authors or their institutions.
- High Visibility: indexed within ESCI (Web of Science), Scopus, EBSCO, and other databases.
- Journal Rank: JCR - Q1 (Computer Science, Interdisciplinary Applications) / CiteScore - Q2 (Artificial Intelligence)
- Rapid Publication: manuscripts are peer-reviewed and a first decision is provided to authors approximately 20.4 days after submission; acceptance to publication is undertaken in 5.6 days (median values for papers published in this journal in the first half of 2026).
- Recognition of Reviewers: APC discount vouchers, optional signed peer review, and reviewer names published annually in the journal.
- Journal Cluster of Artificial Intelligence: AI, AI in Medicine, Algorithms, BDCC, MAKE, MTI, Stats, Virtual Worlds, Computers and Journal of Superintelligence.
Impact Factor:
6.5 (2025);
5-Year Impact Factor:
5.6 (2025)
Latest Articles
Advances in Multimodal Deep Learning for Drug Repurposing
AI 2026, 7(9), 335; https://doi.org/10.3390/ai7090335 (registering DOI) - 28 Aug 2026
Abstract
Computational drug repurposing increasingly integrates chemical, biological, omics, network, text, and clinical data through deep learning. This structured narrative review examines how such modalities are encoded, aligned, and fused. We organize representative studies into four mechanism-centered families: heterogeneous-graph neural networks, multimodal knowledge-graph embeddings,
[...] Read more.
Computational drug repurposing increasingly integrates chemical, biological, omics, network, text, and clinical data through deep learning. This structured narrative review examines how such modalities are encoded, aligned, and fused. We organize representative studies into four mechanism-centered families: heterogeneous-graph neural networks, multimodal knowledge-graph embeddings, pretrained language/sequence model-based cross-modal alignment, and multi-view or reconstruction-based fusion. Direct drug–disease association and repurposing studies form the core evidence; drug–target interaction, drug–drug interaction, target-identification, molecular-pretraining, and drug–microbe studies are treated as adjacent methodological evidence. We compare architectures, evaluation settings, failure modes, and evidence levels across oncology, neurology, infectious, and rare diseases. Practical guidance covers leakage-aware random, cold-start, temporal, and cluster-based evaluation; an actionable reproducibility checklist; and a scenario-based model-selection framework. We distinguish computational prioritization, docking, preclinical, retrospective clinical, and prospective evidence, and examine data sparsity, uncertain negatives, missing or noisy modalities, interpretability, and translational limitations. Future priorities include temporal and causal evaluation, external and multi-center validation, federated learning, and emerging therapeutic modalities. Multimodal fusion can improve complementary representation, but its value depends on task definition, data quality, evaluation design, and independent validation.
Full article
Open AccessArticle
SIRModel: Learning Spatial Intermediate Representation to Parameter-Efficiently Fine-Tune a Vision Language Model for Manipulation
by
Li Lin, Minghao Shi and Tenglong Wang
AI 2026, 7(9), 334; https://doi.org/10.3390/ai7090334 (registering DOI) - 28 Aug 2026
Abstract
Long-horizon robotic manipulation requires a policy to bridge task-level semantic reasoning with metric three-dimensional interaction geometry. Existing vision–language–action policies usually acquire geometry implicitly from visual tokens or introduce deterministic intermediate variables only in the image plane, which rely on expensive human annotations and
[...] Read more.
Long-horizon robotic manipulation requires a policy to bridge task-level semantic reasoning with metric three-dimensional interaction geometry. Existing vision–language–action policies usually acquire geometry implicitly from visual tokens or introduce deterministic intermediate variables only in the image plane, which rely on expensive human annotations and training cost. This article presents a spatial Gaussian-guided hierarchical framework that uses ordered 3D Gaussian interaction regions as an explicit planning interface between vision–language reasoning and action generation. The proposed framework enables efficient adaptation of a pretrained vision–language model for robotic manipulation tasks. First, an automatic geometric enhancement pipeline converts raw robot demonstration videos into near-, mid-, and late-stage Gaussian supervision through foreground extraction, metric depth estimation, stable camera aggregation, end-effector localization, 3D lifting, and temporal grouping, without requiring manual 3D interaction annotation. The generated Gaussian representations provide structured spatial guidance, where their covariance characterizes interaction-region extent and variability rather than fully calibrated physical uncertainty. Second, a shared vision–language backbone predicts structured subtasks and Gaussian interaction regions, while a conditional diffusion executor generates future action chunks under these semantic and geometric conditions. A trajectory-to-Gaussian likelihood objective explicitly encourages consistency between generated motions and the predicted spatial interaction plan. Experiments on a mixed real-robot dataset derived from LHManip and RH20T show that our method improves trajectory tracking success from 55.7% to 70.8% over a same-backbone direct VLA baseline. Closed-loop simulation evaluation on LIBERO with 80% backbone parameter frozen achieves 85.3% average task success, demonstrating the effectiveness of explicit 3D interaction representations for spatial reasoning and long-horizon manipulation.
Full article
Open AccessArticle
Causal Machine Learning for Heterogeneous Cost Effects in Mutual Funds: A Double Machine Learning and Causal Forest Approach
by
László Vancsura
AI 2026, 7(9), 333; https://doi.org/10.3390/ai7090333 (registering DOI) - 28 Aug 2026
Abstract
The cost–performance relationship in mutual funds is a longstanding open question in financial economics, particularly when costs are assumed to exert a single, linear effect on returns. This study proposes an integrated causal machine learning framework to revisit this question using a panel
[...] Read more.
The cost–performance relationship in mutual funds is a longstanding open question in financial economics, particularly when costs are assumed to exert a single, linear effect on returns. This study proposes an integrated causal machine learning framework to revisit this question using a panel of Hungarian open-ended public investment funds across all major asset classes—equity, bond, absolute yield, misc, money market, real estate, and commodity—covering 2017–2024. Six machine learning algorithms are benchmarked for return prediction, and Double Machine Learning, with fund-level cluster-robust inference and year fixed effects, is applied to estimate the effect of the Total Expense Ratio (TER) on next-year returns, under the identifying assumptions stated in the paper, while flexibly controlling for a set of observed fund-level confounders (size, NAV dynamics, volatility, past and cumulative performance, and fund age) without imposing a linear functional form. To move beyond average effects, a Causal Forest model—tuned using an out-of-fold, effect size-neutral selection criterion—estimates heterogeneous treatment effects across funds, and SHAP-based interpretation uncovers the mechanisms underlying this heterogeneity. The results show that, once the outcome is measured in the year following the one in which TER is observed and panel dependence is properly accounted for, the average TER effect is not robustly different from zero at the full-sample level; where a statistically robust effect emerges, it is negative rather than positive, concentrated in equity and absolute-yield funds, and largely confined to the period after 2022, which coincided with the war in Ukraine, rising interest rates, and heightened market volatility, although the research design does not identify which, if any, of these developments drove the change. Average-effect models are shown to conceal this heterogeneity, and the results are further shown to be sensitive to two methodological choices that might otherwise appear secondary—the timing convention linking cost and return, and the criterion used to select among competing heterogeneous-effects specifications—underscoring the importance of making such choices explicit. These findings demonstrate the added value of combining predictive and causal machine learning, together with identification-robust and panel-robust inference, for uncovering heterogeneity that conventional econometric approaches overlook and offer a transferable methodological template for causal machine learning applications in finance and other high-dimensional decision-making domains.
Full article
(This article belongs to the Section AI Systems: Theory and Applications)
►▼
Show Figures

Figure 1
Open AccessArticle
Research on Driver Mental Fatigue Detection Based on Improved Stripe Attention Mechanism and Deep Residual Shrinking Network
by
Xinyuan Zhang, Rui Zhao, Tianyue Sun and Yonghong Xu
AI 2026, 7(9), 332; https://doi.org/10.3390/ai7090332 - 27 Aug 2026
Abstract
Driving-fatigue-induced attentional decline and response retardation are critical contributors to traffic accidents. However, stably and precisely identifying fatigue states from noisy electroencephalogram (EEG) signals remains a challenging issue in intelligent driving safety. To address the dual deficiencies of traditional methods in fatigue feature
[...] Read more.
Driving-fatigue-induced attentional decline and response retardation are critical contributors to traffic accidents. However, stably and precisely identifying fatigue states from noisy electroencephalogram (EEG) signals remains a challenging issue in intelligent driving safety. To address the dual deficiencies of traditional methods in fatigue feature extraction precision and noise robustness, this paper innovatively constructs a collaborative recognition framework that integrates an Improved Strip Attention Mechanism (ISAM) with a Deep Residual Shrinkage Network (DRSN). The core innovations of this framework are twofold: ISAM achieves precise localization and focused enhancement of fatigue-related rhythmic bands in EEG signals via row–column separable adaptive pooling and channel-wise attention augmentation; concurrently, the DRSN module introduces an improved soft-thresholding function, which adaptively generates filtering thresholds through channel attention to effectively suppress noise and artifact interference in physiological signals. The deep fusion of these two modules forms a closed-loop optimization chain of “targeted feature reinforcement–adaptive noise suppression,” enabling the model to stably extract highly discriminative fatigue representations from complex non-stationary EEG signals. Validation on two public datasets, SEED-VIG and SADT, demonstrates that the proposed method achieves recognition accuracies of 98.86% and 97.38%, respectively, outperforming mainstream methods such as the convolutional spatial-frequency network and multi-scale convolutional neural network by 17.38% and 17.76%. These results confirm the significant advantages of the proposed dual-module collaborative architecture in precise fatigue characterization and anti-interference capability, offering a highly reliable technical solution for real-time driver mental fatigue monitoring in real-world road scenarios.
Full article
(This article belongs to the Special Issue Advances in AI-Driven Perception and Intelligent Control for Autonomous Vehicles)
►▼
Show Figures

Figure 1
Open AccessArticle
A Concept-Bottleneck Explainable AI Framework for Diagnosing Agile Delivery Outcomes
by
Ali Akbar ForouzeshNejad and Alexander Gegov
AI 2026, 7(9), 331; https://doi.org/10.3390/ai7090331 - 26 Aug 2026
Abstract
Agile outcome models commonly map Jira variables directly to a retrospective label and then explain the prediction through fragmented feature attributions; they rarely separate domain concepts, team clustering, unresolved work, and concept-label coupling. This study evaluates a domain-informed, concept-bottleneck-style explainable AI architecture for
[...] Read more.
Agile outcome models commonly map Jira variables directly to a retrospective label and then explain the prediction through fragmented feature attributions; they rarely separate domain concepts, team clustering, unresolved work, and concept-label coupling. This study evaluates a domain-informed, concept-bottleneck-style explainable AI architecture for retrospective diagnosis of Agile Epic outcomes. A frozen Jira export of 10,000 unique issue-level records was linked to a pre-specified analytical cohort of 180 Epics across 14 teams. Six experts rated efficiency, effectiveness, sustainability, and contextual risk, while outcomes were recorded as Successful, Challenged, or Unsuccessful. Because the outcome labels and concept ratings were informed by the same Jira evidence, the models estimate consistency with an expert labelling procedure, rather than independent project success. Under five-fold group-aware cross-validation, the fixed-configuration flat LightGBM achieved macro-F1 = 0.864 ± 0.053 and the fixed-configuration HMXAI/CBM-style model achieved 0.843 ± 0.084. These descriptive primary scores are not a joint nested-model-selection comparison. The proposed method, therefore does, not demonstrate a performance improvement; its contribution is an inspectable diagnostic structure. Performance fell materially on the resolved-only subset (LightGBM macro-F1 = 0.645), and model-specific nested, leave-one-team-out, calibration, uncertainty, correlation, and intervention analyses further bound the claims. Concept interventions were not uniformly monotone, so the concept layer is domain-interpretable in form but not yet user-validated as actionable. The study contributes a transparent audit of when concept-level diagnosis can complement flat classification and when circularity, censoring, and shortcut learning restrict interpretation.
Full article
(This article belongs to the Topic Theories, Techniques, and Real-World Applications for Advancing Explainable AI)
►▼
Show Figures

Figure 1
Open AccessArticle
Monthly PM2.5 Forecasting with Temporally Constrained Rolling Decomposition and DenseMamba
by
Hongbin Dai, Chen Wu, Qinqin Zhang and Huibin Zeng
AI 2026, 7(9), 330; https://doi.org/10.3390/ai7090330 - 26 Aug 2026
Abstract
The reliable monthly forecasting of fine particulate matter (PM2.5) requires artificial intelligence (AI) models that are accurate, temporally valid, and transparent. We develop a history-only rolling-decomposition framework with lightweight Mamba-inspired selective state-space backbones for 475 city-level administrative units in China. Complete Ensemble Empirical
[...] Read more.
The reliable monthly forecasting of fine particulate matter (PM2.5) requires artificial intelligence (AI) models that are accurate, temporally valid, and transparent. We develop a history-only rolling-decomposition framework with lightweight Mamba-inspired selective state-space backbones for 475 city-level administrative units in China. Complete Ensemble Empirical Mode Decomposition with Adaptive Noise (CEEMDAN) and three alternative decomposition strategies use only the PM2.5 data available before each decomposition cutoff, thereby avoiding future-information leakage and yielding inspectable multi-scale predictors. In the primary seven-model benchmark, CEEMDAN-DenseMamba achieved the lowest mean root mean squared error and mean absolute error (6.931 and 4.833 μg m−3, respectively). Equal-optimization reruns, paired moving-block bootstrap intervals, component-count sensitivity, city-wise diagnostics, and a parameter-matched gated recurrent unit baseline were then used to examine performance attribution. Under common optimization, the dense-connection contrasts showed paired error reductions with confidence intervals below zero, whereas the incremental CEEMDAN effect within a fixed backbone was smaller and its paired confidence intervals crossed zero. The recurrent baseline remained competitive. These findings support transparent, temporally valid multi-scale forecasting while limiting inference to future-month prediction for the known cities and the present experimental setting.
Full article
(This article belongs to the Special Issue Advances in Deep Learning for Air Pollution and Climate Science Research)
►▼
Show Figures

Figure 1
Open AccessArticle
Benchmarking Deep Learning Against Statistical Baselines and a Physical Climate-Model Comparator for Station-Scale Meteorological Forecasting: A 100-Station Study from the Western Balkans
by
Dalibor Nikolić, Ivica Djalović, Ivan Vitezović, Dejan B. Stojanović, Sara Pavkov, Rastislav Stojsavljević and Mlađen Jovanović
AI 2026, 7(9), 329; https://doi.org/10.3390/ai7090329 - 26 Aug 2026
Abstract
►▼
Show Figures
Benchmarking deep learning forecasters against classical and physically based numerical baselines remains uncommon in the time-series forecasting literature. Meteorological station networks offer an under-exploited evaluation environment, uniquely providing a physically based climate-model comparator alongside standard baselines. We evaluated eight forecasting approaches—climatology, SARIMA, Random
[...] Read more.
Benchmarking deep learning forecasters against classical and physically based numerical baselines remains uncommon in the time-series forecasting literature. Meteorological station networks offer an under-exploited evaluation environment, uniquely providing a physically based climate-model comparator alongside standard baselines. We evaluated eight forecasting approaches—climatology, SARIMA, Random Forest, and five deep learning architectures (TFT, N-HiTS, PatchTST, TiDE, xLSTM)—against bias-corrected output from a five-member CMIP6 ensemble, on 100 meteorological stations across four Western Balkan countries (monthly temperature and precipitation, 1961–2020), using non-parametric significance testing, a rolling-origin backtest (five windows, 2011–2020), and a five-seed robustness check. For temperature, all five deep learning architectures achieved lower MAE than the classical baselines (p < 10−99), though PatchTST’s advantage over climatology was not significant; the best-performing architecture varied across seeds and evaluation windows, so we characterise a leading cluster (N-HiTS, TFT, TiDE, PatchTST) rather than a single winner. The primary temperature advantage was geographically broad-based, while the comparison against the physical-model baseline was robust to the choice of comparator GCM. For precipitation, by contrast, a simple climatological-mean baseline outperformed all five deep learning architectures with no exception across all five rolling-origin windows. The deep learning advantage over classical and physical baselines is thus variable-specific rather than universal. Meteorological station networks, combined with a physically based climate-model comparator, constitute a well-suited evaluation environment for the broader time series forecasting community.
Full article

Figure 1
Open AccessArticle
The AI Literacy Leadership Framework (AILLF): A Framework for AI-Enabled Leadership in Higher Education
by
Alaa Mohasseb, Ronel Beukman and Andreas Kanavos
AI 2026, 7(9), 328; https://doi.org/10.3390/ai7090328 - 26 Aug 2026
Abstract
The integration of Artificial Intelligence (AI) in higher education is reshaping institutional decision-making, governance, policy development, and educational innovation. Effective leadership in AI-enabled environments requires more than technical competence, extending to strategic awareness, ethical judgement, and the ability to critically evaluate the broader
[...] Read more.
The integration of Artificial Intelligence (AI) in higher education is reshaping institutional decision-making, governance, policy development, and educational innovation. Effective leadership in AI-enabled environments requires more than technical competence, extending to strategic awareness, ethical judgement, and the ability to critically evaluate the broader implications of AI technologies. Drawing on survey and interview data from 52 academic leaders across UK higher education institutions, this study examines current levels of AI literacy and explores how AI capability relates to institutional readiness and leadership practice. The findings reveal variation in participants’ self-reported AI literacy and engagement, identify technical, strategic, ethical, and organisational capability gaps, and highlight structural barriers including limited time, fragmented professional development, and insufficient institutional support for institution-wide AI adoption. In response, the paper presents the AI Literacy Leadership Framework (AILLF), which conceptualises AI literacy as a multidimensional leadership capability comprising technical, strategic, ethical, and applied dimensions that support four interconnected leadership domains: innovation, decision-making, ethical governance, and policy development. Informed by leadership theory, international AI governance frameworks, and the study’s exploratory empirical findings, the AILLF is accompanied by a proposed capability progression model and role-differentiated leadership competency guide to provide implementation guidance for higher education institutions. The study contributes an empirically informed conceptual framework that integrates AI literacy, leadership theory, and AI governance, providing a foundation for future research and institutional approaches to AI leadership within higher education.
Full article
(This article belongs to the Special Issue Human–AI Collaboration and Augmented Intelligence: Bridging Theory, Design and Practice)
►▼
Show Figures

Figure 1
Open AccessSystematic Review
Credible Sovereignty: Operationalizing AI Governance Across Infrastructure, Data, and Models: A Systematic Review
by
Raghu Raman and Prema Nedungadi
AI 2026, 7(9), 327; https://doi.org/10.3390/ai7090327 - 24 Aug 2026
Abstract
Claims of AI sovereignty are increasingly invoked but operational control remains uneven. Claims to control are made through national models, sovereign clouds, data localization mandates, and procurement rules; however, whether such claims translate into demonstrable control over how AI systems are run, inspected,
[...] Read more.
Claims of AI sovereignty are increasingly invoked but operational control remains uneven. Claims to control are made through national models, sovereign clouds, data localization mandates, and procurement rules; however, whether such claims translate into demonstrable control over how AI systems are run, inspected, and contested remains poorly understood. This paper introduces credible sovereignty, the gap between declared and demonstrable control in deployment, as a conceptual lens for analyzing AI governance to examine how this gap is opened and closed across infrastructure, data, and model supply chains. Using a PRISMA-guided social-science corpus and machine learning-based BERTopic modeling, validated through topic diversity and topic separation diagnostics and triangulated through close reading, the analysis identifies four governance logics through which sovereignty is contested: data infrastructure and legitimacy frameworks; techno-bloc diplomacy and infrastructure politics; European regulatory sovereignty; and community-driven sovereignty in the Global South. Across these logics, sovereignty is enacted less through national capabilities than through proxy mechanisms—certification regimes, procurement clauses, cloud governance, and deployment architectures—each carrying trade-offs between autonomy, dependence, and accountability. Rereading the corpus through an Antecedents–Decisions–Outcomes lens yields a testable research agenda: antecedents that push actors toward sovereignty seeking; design and governance choices that translate ambition into implementation; and outcomes—resilience, inclusion, accountability—against which sovereign AI programs should be assessed. This paper reframes sovereignty as a layered operational capability rather than a discursive claim and links computational synthesis to a normative construct that applies across jurisdictions and scales.
Full article
(This article belongs to the Section AI Systems: Theory and Applications)
►▼
Show Figures

Figure 1
Open AccessReview
Ethics Before Algorithms: A Framework for AI-Driven Corporate Transparency
by
Nguyen Thi Thanh Binh
AI 2026, 7(9), 326; https://doi.org/10.3390/ai7090326 - 24 Aug 2026
Abstract
This review synthesizes theoretical and empirical insights from 1055 peer-reviewed articles on artificial intelligence (AI), corporate governance, and ethics. Situated in the corporate governance and accounting literature, it develops a computational framework to identify thematic patterns and conceptual links among AI, transparency, accounting,
[...] Read more.
This review synthesizes theoretical and empirical insights from 1055 peer-reviewed articles on artificial intelligence (AI), corporate governance, and ethics. Situated in the corporate governance and accounting literature, it develops a computational framework to identify thematic patterns and conceptual links among AI, transparency, accounting, governance, and ESG. Using latent Dirichlet allocation, co-occurrence network analysis, sentence-level semantic similarity, and exploratory regression, the study identifies three recurring configurations of conceptual association: (1) Ethics, Governance, and Transparency; (2) Machine Learning, Finance, Blockchain, and Accounting; and (3) Corporate, ESG, and Accounting. The findings indicate that these themes are repeatedly connected within the scholarly literature.
Full article
(This article belongs to the Topic AI-Driven Information Governance for Sustainable Decision Making and Innovation)
►▼
Show Figures

Figure 1
Open AccessArticle
Evaluating Feature-Based Machine-Learning Models with Post Hoc Explainability for Eye-Tracking-Based Task Type and Workload Inference
by
Tomi Božak, Shivalika Goyal, Marc Langheinrich, Martin Gjoreski and Gašper Slapničar
AI 2026, 7(8), 325; https://doi.org/10.3390/ai7080325 - 21 Aug 2026
Abstract
Eye tracking is a valuable behavioral signal for human-centered AI, yet the reliability of feature-based machine-learning models for inferring task type and workload across users and tasks remains uncertain, because experimentally defined workload labels may reflect task type and visual structure as much
[...] Read more.
Eye tracking is a valuable behavioral signal for human-centered AI, yet the reliability of feature-based machine-learning models for inferring task type and workload across users and tasks remains uncertain, because experimentally defined workload labels may reflect task type and visual structure as much as cognitive demand. The practical problem is that designers of gaze-adaptive systems need to know which inferences are dependable enough to act on, and reported accuracies alone do not answer this, because the choice of prediction target and validation split can determine the result. This study systematically evaluates feature-based machine-learning models with post hoc explainability across three prediction targets: task type, binary load-versus-rest, and three-level workload. Eye-movement features derived from fixations, saccades, pupils, and blinks were extracted from short temporal windows collected from 54 participants performing attention, visual-spatial, and memory tasks under rest, easy, and difficult conditions, and evaluated using leave-one-subject-out (LOSO) and leave-one-group-out (LOGO) validation. Task type was classified most reliably (85.9% LOSO, 83.4% LOGO), binary load-versus-rest showed moderate, validation-sensitive robustness (81.4% LOSO, 63.9% LOGO), and three-level workload classification was substantially more challenging (56.4% LOSO, 44.3% LOGO). SHAP and statistical analyses consistently identified fixation dispersion, pupil-related measures, and subject-normalized features as the strongest contributors across all three targets. These findings show that prediction target definition, validation strategy, and post hoc explainability jointly determine what can be reliably inferred from gaze-based machine-learning models. Eye tracking alone therefore appears promising for task-type recognition and may support coarse engagement-related inference when the deployment task family is represented during model development, whereas task-independent fine-grained workload estimation remains unsupported by the present evidence.
Full article
(This article belongs to the Special Issue Human-Computer Interaction and Human-Centered AI)
►▼
Show Figures

Figure 1
Open AccessPerspective
From Empowerment to Vulnerability: The Computation–Energy Paradox of AI-Enabled Power-Transport Systems
by
Chenxuan Zhang, Peixiao Fan, Siqi Bu and Yuxin Wen
AI 2026, 7(8), 324; https://doi.org/10.3390/ai7080324 - 21 Aug 2026
Abstract
►▼
Show Figures
The transition towards smart megacities has deeply integrated Artificial Intelligence (AI) with power–transport networks. While AI empowers complex operations like multi-network coordinated dispatch and emergency rescue, current algorithm-centric perspectives largely ignore its massive physical energy costs. Accordingly, this Perspective examines the dual role
[...] Read more.
The transition towards smart megacities has deeply integrated Artificial Intelligence (AI) with power–transport networks. While AI empowers complex operations like multi-network coordinated dispatch and emergency rescue, current algorithm-centric perspectives largely ignore its massive physical energy costs. Accordingly, this Perspective examines the dual role of AI, considering it not only as an intelligent decision-support tool but also as a potential source of additional stress on physical infrastructure. First, through a structured synthesis of the representative literature, we deconstruct the functional dependencies between algorithms and physical infrastructures, identifying how AI reshapes the operational paradigms of power, ground transport, and aerial networks under routine and emergency scenarios. We then introduce the concept of the “Computation–Energy Paradox.” Integrating conceptual analysis with a quantitative case study of a typical community, we illustrate a plausible failure mechanism: during extreme disasters, intensified AI invocation for emergency management generates surging computational loads, which paradoxically exacerbate power shortages and reduce the operating margin of already weakened systems. In addition, we analyze core engineering bottlenecks, including spatiotemporal computation–energy mismatches and physical constraints in extreme edge environments. To address these challenges, we outline a prospective roadmap encompassing lightweight emergency AI and computation–power-coordinated offloading mechanisms. Finally, the sustainable development of such systems suggests a paradigm shift: AI must evolve from a purely virtual algorithm into a physical component of an integrated compute–power–transport system.
Full article

Figure 1
Open AccessArticle
Systematic Comparison of Electroencephalography Feature Domains for Visual Stimuli Decoding with EEGNet and EEG Conformer
by
Cesar Agustin Corona-Patricio, Carolina Reta and Jose Antonio Cantoral-Ceballos
AI 2026, 7(8), 323; https://doi.org/10.3390/ai7080323 - 20 Aug 2026
Abstract
►▼
Show Figures
Electroencephalography-based visual decoding has important applications in brain–computer interfaces and cognitive neuroscience, yet the relative effectiveness of different feature extraction methods for sustained visual paradigms remains unclear due to the absence of standardized, multi-dataset comparative evaluations. This study systematically compares eight feature extraction
[...] Read more.
Electroencephalography-based visual decoding has important applications in brain–computer interfaces and cognitive neuroscience, yet the relative effectiveness of different feature extraction methods for sustained visual paradigms remains unclear due to the absence of standardized, multi-dataset comparative evaluations. This study systematically compares eight feature extraction methods across three public EEG datasets: MindBigData MNIST, MindBigData MNIST-8B for digit recognition, and MSS for natural image classification. The methods include coherence, Granger causality, directed transfer function, partially directed coherence, transfer entropy, discrete wavelet transform, empirical wavelet transform (EWT), and wavelet scattering transform. Two deep learning architectures, EEGNet and EEG Conformer, were trained using two pre-processing pipelines, with and without artifact removal. EWT achieved the highest classification accuracy, reaching 97.83% for digit-vs-blank and 77.10% for within-session natural image classification. Connectivity-based methods consistently underperformed, with the best connectivity method (coherence) reaching up to 91.67%, suggesting that spectral power information is more discriminative than inter-channel relationships. Cross-subject generalization remained challenging, with best accuracies near 68%. The findings establish wavelet-based adaptive spectral decomposition as a strong baseline for EEG visual decoding and highlight the need for domain adaptation techniques to address cross-subject variability.
Full article

Figure 1
Open AccessReview
The Governance Gap in Contemporary LLM-Based Agentic Systems: A Structural Diagnostic Review
by
Christopher Valdez-Cantú, Jose Antonio Cantoral-Ceballos and Joanna Alvarado-Uribe
AI 2026, 7(8), 322; https://doi.org/10.3390/ai7080322 - 20 Aug 2026
Abstract
►▼
Show Figures
Large Language Models (LLMs) are increasingly integrated into agentic workflows that require extended reasoning, persistent state management, coordinated tool use, and controlled execution. As this operational scope expands, a central question emerges: whether probabilistic generation alone can reliably support coherent behavior across interacting
[...] Read more.
Large Language Models (LLMs) are increasingly integrated into agentic workflows that require extended reasoning, persistent state management, coordinated tool use, and controlled execution. As this operational scope expands, a central question emerges: whether probabilistic generation alone can reliably support coherent behavior across interacting system components. This paper addresses that question through a structural diagnostic review of contemporary agentic systems. Starting from LLM-based tutoring as an analytically demanding entry point and extending toward structurally related agent architectures, the paper draws on a five-phase review of research records. The analysis is organized through the Agentic Structure Taxonomy (AST), which structures the literature across four dimensions: Cognition, Interaction, Orchestration, and Governance. The review identifies five recurrent empirical problem patterns and uses them as abductive diagnostic cues for formulating seven cross-dimensional transition gaps that capture recurrent discontinuities at the boundaries between reasoning, state, control, and execution. From these gaps, fourteen structural constraints are derived across three control domains: state isolation, control alignment, and execution governance. These constraints are interpreted not as prescriptive design mandates, but as analytically derived conditions associated with reducing error propagation across subsystem transitions. The paper argues that reliability in agentic systems is shaped not only by model performance or prompt design, but also by whether the boundaries linking probabilistic reasoning to persistent state, orchestration, and execution are governed by explicit structural conditions.
Full article

Figure 1
Open AccessArticle
Benchmarking Normative AI Assistants Under Inconsistent Evidence with Paraconsistent Trace Semantics
by
Maksim V. Ulizko, Aleksandr V. Chernikov, Ivan V. Tomilov, Natalia F. Gusarova and Aleksandra S. Vatian
AI 2026, 7(8), 321; https://doi.org/10.3390/ai7080321 - 20 Aug 2026
Abstract
►▼
Show Figures
Normative AI assistants are increasingly used in domains governed by duties, permissions, prohibitions, exceptions, priorities, and institutional policies. Existing retrieval-augmented generation (RAG) and legal AI benchmarks evaluate answer accuracy, retrieval quality, citation grounding, natural-language inference, clause extraction, or general legal reasoning ability. These
[...] Read more.
Normative AI assistants are increasingly used in domains governed by duties, permissions, prohibitions, exceptions, priorities, and institutional policies. Existing retrieval-augmented generation (RAG) and legal AI benchmarks evaluate answer accuracy, retrieval quality, citation grounding, natural-language inference, clause extraction, or general legal reasoning ability. These dimensions are necessary but insufficient when supplied evidence is incomplete, mutually inconsistent, or defeasible. The objective of this study is to introduce ParaTraceBench, a paraconsistent trace-based benchmarking framework for post-retrieval normative reasoning over fixed evidence packages. Each scenario contains a query, evidence fragments, extracted facts, defeasible rules, typed attack edges, priority relations, an expected conclusion status, and a gold diagnostic trace. The formalism uses evidence-grounded arguments, a single edge-based attack representation, explicit attack-licensing rules, acyclic priority bases with a transitive closure, grounded argument labeling, trace-normal-form alignment, and deterministic scoring. The operational NER metric is explicitly interpreted as inconsistency-conditioned unsupported-conclusion avoidance rather than proof of logical non-explosion. We evaluated the framework using 140 scenarios, external validation on 567 anonymized Russian-language cases from Russian Federation and EAEU-related materials, reasoning-oriented baseline adaptations, five-run prompt-fairness and stability controls, and a deterministic component-dependency audit. On the full external set, the trace-based configuration reached 85.7% answer-status accuracy, 85.5% contradiction-localization accuracy, 94.2% operational NER, 84.1% priority-handling accuracy, and 83.7% belief-revision accuracy. These results indicate that contradiction-aware trace evaluation provides diagnostic information beyond final-answer accuracy under the evaluated fixed-evidence conditions, while not establishing causal architectural superiority, logical non-triviality, or end-to-end RAG performance.
Full article

Figure 1
Open AccessArticle
Retrieval Granularity as Evidence Design in Small-Model RAG Question Answering: A Diagnostic HotpotQA Study
by
Weimao Ke, Lixiao Yang and Mengyang Xu
AI 2026, 7(8), 320; https://doi.org/10.3390/ai7080320 - 19 Aug 2026
Abstract
Retrieval-Augmented Generation (RAG) has become a practical approach for question answering over external corpora, particularly when answers should be grounded in source documents rather than generated only from model parameters. While recent large language models can process increasingly long contexts, they do not
[...] Read more.
Retrieval-Augmented Generation (RAG) has become a practical approach for question answering over external corpora, particularly when answers should be grounded in source documents rather than generated only from model parameters. While recent large language models can process increasingly long contexts, they do not remove the need for selecting, organizing, and auditing evidence, especially when systems rely on smaller local models for privacy, cost, or deployment constraints. In this paper, we frame retrieval granularity as an evidence-design variable for answer grounding in small-model RAG question answering. After a brief exploratory NewsQA phase that motivates the error categories, the main study uses the HotpotQA distractor validation split with 7405 hard multi-hop questions and sentence-level supporting-fact annotations. With Qwen3-8B as the fixed generator, we compare closed-book, fixed-budget whole-context, retrieved-context, gold-document, and gold-supporting-fact conditions while varying retrieval granularity, retriever type, and context budget. Retrieved context substantially outperforms closed-book answering and the 1024-token fixed-budget whole-context condition but remains below gold-document and gold-supporting-fact upper bounds, indicating that retrieval, generation, and evaluation limitations should be analyzed separately. Sentence-level retrieval under-recovers multi-hop evidence, especially for questions with three or more supporting facts, while paragraph-level and moderate token-level chunks recover substantially more complete evidence. In the full condition matrix, hybrid retrieval with 256-token chunks and no overlap achieves an F1 of 0.6816 with a supporting-fact recall of 0.9609, compared with an F1 of 0.6166 and supporting-fact recall of 0.7801 for BM25 sentence retrieval. Additional ablations show that fixed-budget whole-context performance is strongly affected by truncation, that overlap has little practical effect under the tested 1024-token budget, and that a stronger BGE dense retriever improves the best retrieved-context F1 to 0.7027. These results align with a diagnostic perspective on chunking: using evidence at a task-appropriate level of granularity can improve grounding, auditability, and answer quality, but the observed patterns should be interpreted within the HotpotQA distractor setting, fixed generator, and tested context budgets.
Full article
(This article belongs to the Special Issue Large Language Models and Retrieval-Augmented Generation in Natural Language Processing, Human–Robot Interaction and Quantum Computing)
►▼
Show Figures

Figure 1
Open AccessArticle
Detecting AI-Generated Text and Code: An Empirical Study of Cross-Generator and Cross-Domain Generalization
by
Neethika Alluri, Pardha Saradhi Varma Gottumukkala and Hemalatha Indukuri
AI 2026, 7(8), 319; https://doi.org/10.3390/ai7080319 - 19 Aug 2026
Abstract
Large language models (LLMs) now generate fluent natural language and source code, creating challenges for authorship attribution, academic integrity, and software supply-chain security. Most existing detectors for AI-generated content are evaluated separately on natural language or source code, often under matched train–test conditions
[...] Read more.
Large language models (LLMs) now generate fluent natural language and source code, creating challenges for authorship attribution, academic integrity, and software supply-chain security. Most existing detectors for AI-generated content are evaluated separately on natural language or source code, often under matched train–test conditions that can overestimate real-world reliability. We present a paired-prompt benchmark for human-versus-machine detection across English text, Python code, and mixed text–code documents. The benchmark includes instances from HC3, CodeSearchNet, MBPP, and HumanEval across training, validation, and test partitions, plus Mix-Eval, a mixed-content set of 997 Jupyter-notebook-style samples. We evaluate RoBERTa-large for text, GraphCodeBERT and CodeBERT-base for code, a unified RoBERTa-base detector trained on both modalities, and zero-shot baselines. Fine-tuned detectors achieve near-perfect in-distribution performance, with AUROC and accuracy above . Across five instruction-tuned generator families of varying size (3.8B–7B) and architecture, with the human and problem distributions held fixed, cross-generator transfer causes negligible degradation (AUROC spread ; drops of at most ). In contrast, domain shift is the main failure mode: on MBPP+HumanEval, GraphCodeBERT drops to AUROC and CodeBERT-base to . On Mix-Eval, the unified detector outperforms a routed text–code pipeline by 21 AUROC points ( vs. ), largely because of router failures on mixed inputs. Training-time augmentation improves low-false-positive performance, while legacy supervised detectors show systematic class inversion on modern LLM outputs. These results show that reliable deployment requires cross-domain evaluation, mixed-content testing, and calibration beyond in-distribution accuracy.
Full article
(This article belongs to the Special Issue Trustworthy Large Language Models: Advancing Reliability, Safety, Fairness, and Transparency)
►▼
Show Figures

Figure 1
Open AccessArticle
CBR-Enhanced ResNet50 for Five-Class Diabetic Retinopathy Grading: An Ablation-Based Study
by
Samir Elouaham, Fatima Ezzahra Bouaaza, Ilyas Ait Ichou and Boujemaa Nassiri
AI 2026, 7(8), 318; https://doi.org/10.3390/ai7080318 - 19 Aug 2026
Abstract
Diabetic retinopathy (DR) is a common complication of diabetes and one of the leading causes of preventable vision loss worldwide. Because the manual grading of color fundus images is slow and depends on the availability of trained specialists, automated screening tools are needed.
[...] Read more.
Diabetic retinopathy (DR) is a common complication of diabetes and one of the leading causes of preventable vision loss worldwide. Because the manual grading of color fundus images is slow and depends on the availability of trained specialists, automated screening tools are needed. This study proposes a lightweight channel-wise refinement strategy for automatic five-class DR grading, built on a ResNet50 backbone. Two custom blocks are evaluated: CBR, which applies a 3 × 3 convolution, batch normalization, and a ReLU activation to make the channel representation more compact, and CBS, which applies a 3 × 3 convolution, batch normalization, and a SiLU activation to reinforce local spatial features. On the Diabetic Retinopathy Balanced dataset, the baseline ResNet50 reached an accuracy of 90.77%, a precision of 90.60%, a recall of 90.79%, and an F1-score of 90.64%. In the ablation study, the best configuration was ResNet50 + CBR, with an accuracy of 91.85%, a precision of 91.75%, a recall of 91.88%, and an F1-score of 91.76%. The full CBR-CBS Hybrid ResNet50 was close behind, with an accuracy of 91.81% and an F1-score of 91.71%. The CBR block accounts for most of this improvement, which suggests that channel-wise refinement helps the model separate subtle lesion patterns. These results establish lightweight channel-wise refinement (CBR) as an effective, compact, and interpretable enhancement of ResNet50 for automated five-class DR grading, delivering a consistent multi-metric gain over the baseline and accuracy competitive with the literature, which makes it a promising solution for large-scale screening.
Full article
(This article belongs to the Section Medical & Healthcare AI)
►▼
Show Figures

Figure 1
Open AccessArticle
FroLineR: Front-Line Response with Retrieval-Augmented Prompt-Engineered Reply Generation for IT Help Desks
by
Alexandru Dima, Maria-Elena Mihăilescu, Darius Mihai, Mihai Carabaș and Mihai Dascalu
AI 2026, 7(8), 317; https://doi.org/10.3390/ai7080317 - 19 Aug 2026
Abstract
IT help desks at large organizations face a high volume of recurrent, well-documented user requests that nevertheless require human-written replies, creating a persistent staff workload that is repetitive in content but non-trivial in tone and procedural correctness. We present FroLineR, short for Front-Line
[...] Read more.
IT help desks at large organizations face a high volume of recurrent, well-documented user requests that nevertheless require human-written replies, creating a persistent staff workload that is repetitive in content but non-trivial in tone and procedural correctness. We present FroLineR, short for Front-Line Response, a system that drafts the initial staff reply to such tickets in the login and account-activation category and integrates into a human-in-the-loop ticketing workflow on a Romanian-language ticketing platform. The generator is an unmodified instruct model augmented with retrieval from a small set of hand-curated guide documents, using a Romanian system prompt refined over several rounds of staff review. To evaluate and refine the prompt without manual labeling, we cluster the first user message of every historical thread with both BERTopic and Semantic Signal Separation ( ), score configurations along coherence and lexical-diversity axes, and extract a 200-message evaluation set from the winning model. Prompt convergence was certified by several rounds of manual review by support staff. The production system is quantized to Q4_K_M GGUF, served through llama-cpp-python behind a small Flask API, and deployed with GPU offloading on the target server, reducing end-to-end per-answer latency from approximately 830 s on the server’s CPU to roughly 61 s once layers are offloaded to the GPU, with no observable degradation in answer quality.
Full article
(This article belongs to the Special Issue Large Language Models and Retrieval-Augmented Generation in Natural Language Processing, Human–Robot Interaction and Quantum Computing)
►▼
Show Figures

Figure 1
Open AccessArticle
SmartMM: A Domain-Specific Large Language Model for Medical Microbiology
by
Yongqiang Gong, Ruiqi Ma, Xicheng Wang, Ruixi Li, Han Dong, Yijin Liu, Xi Peng, Quanle Guo and Yin Liu
AI 2026, 7(8), 316; https://doi.org/10.3390/ai7080316 - 18 Aug 2026
Abstract
►▼
Show Figures
Background: Large language models (LLMs) show considerable promise for medical question answering and reasoning. Their use in medical microbiology, however, remains constrained by limited domain-specific knowledge and the risk of hallucinated outputs. Objective: To develop and evaluate Smart Medical Microbiology (SmartMM), a specialized
[...] Read more.
Background: Large language models (LLMs) show considerable promise for medical question answering and reasoning. Their use in medical microbiology, however, remains constrained by limited domain-specific knowledge and the risk of hallucinated outputs. Objective: To develop and evaluate Smart Medical Microbiology (SmartMM), a specialized LLM for accurate, reliable, and context-aware responses in medical microbiology. Methods: SmartMM integrates domain-adaptive continual pretraining, supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), knowledge distillation, and retrieval-augmented generation (RAG). We constructed a high-quality microbiology corpus from textbooks, clinical guidelines, the scientific literature, case reports, and other authoritative sources. Model performance was assessed using objective examinations, subjective generation tasks, expert review, and real-world user preference evaluation. Results: SmartMM achieved accuracies of 0.897 and 0.563 on true-or-false and fill-in-the-blank questions, respectively. In subjective generation tasks, it obtained the highest ROUGE-L score (0.265) and BERTScore F1 score (0.771) among all compared models. Expert assessment showed excellent inter-rater reliability, with all ICC(C,3) values exceeding 0.970. In a user evaluation involving 20 participants and 100 real-world questions, SmartMM received the largest number of first-place rankings (33), placing it among the top-performing systems overall. Conclusions: SmartMM showed strong domain adaptability in medical microbiology knowledge organization, semantic generation, and retrieval-augmented reasoning. These findings support its potential use in educational support, infectious disease knowledge assistance, and retrieval-enhanced medical question answering.
Full article

Figure 1
Highly Accessed Articles
Latest Books
E-Mail Alert
News
Topics
Topic in
Algorithms, Applied Sciences, Electronics, MAKE, AI, Software
Applications of NLP, AI, and ML in Software Engineering
Topic Editors: Affan Yasin, Javed Ali Khan, Lijie WenDeadline: 30 August 2026
Topic in
AI, Applied Sciences, Designs, Electronics, Mathematics, Remote Sensing, Sensors
The Future of Artificial Intelligence: Trends, Challenges, and Developments
Topic Editors: Yinbin Miao, Jun Feng, Wenqian DongDeadline: 30 September 2026
Topic in
Molecules, Biomimetics, Chemosensors, Life, AI, Sci
Recent Advances in Chemical Artificial Intelligence
Topic Editors: Pier Luigi Gentili, Jerzy Górecki, David C Magri, Pasquale StanoDeadline: 15 October 2026
Topic in
AI, Applied Sciences, Computers, Electronics, Entropy, Future Internet, Information, IoT, Sensors, Telecom
Advances in Sixth Generation and Beyond (6G&B)
Topic Editors: Luis Javier García Villalba, Ana Lucila Sandoval OrozcoDeadline: 31 October 2026
Conferences
Special Issues
Special Issue in
AI
Multi-Agent Modal Computing: Synergy, Scalability and Satellite-Earth Collaboration
Guest Editors: Qian Li, Yunfei Long, Cheng JiDeadline: 31 August 2026
Special Issue in
AI
Application of Artificial Intelligence on Structures Subjected to Natural Hazards
Guest Editors: Eden Bojórquez, Juan Bojórquez, Robespierre Chavez-LopezDeadline: 31 August 2026
Special Issue in
AI
Smart Educational Technologies: Integrating AI, Robotics, and AR for 21st Century Learning
Guest Editors: Yi-Chun Lin, Yen-Ting LinDeadline: 31 August 2026
Special Issue in
AI
Next-Generation Medical Imaging AI: Deep Learning Computer Vision with Multi-Task and Continual Learning
Guest Editors: Zahid Ullah, Mustaqeem KhanDeadline: 31 August 2026





