Next Issue
Volume 17, October
Previous Issue
Volume 17, August
 
 

Information, Volume 17, Issue 9 (September 2026) – 120 articles

Cover Story (view full-size image): Cybergrooming is an urgent and growing threat to children online, crossing both language and national borders. Protecting children and uncovering ongoing abuse effectively requires reliable detection approaches beyond English, yet most public datasets and automated methods remain English-centered. As a result, researchers and practitioners lack sufficient authentic labeled data to develop tools in other languages. Hence, our study investigates whether established English datasets can serve as a foundation for broader detection. Using German as a case study, we develop a reproducible multi-step machine translation pipeline and evaluate it using closed law-enforcement cases. How can machine-translated data help bridge the language gap in cybergrooming detection? What is required to move from translated benchmarks toward reliable real-world use? View this paper
  • Issues are regarded as officially published after their release is announced to the table of contents alert mailing list.
  • You may sign up for e-mail alerts to receive table of contents of newly released issues.
  • PDF is the official format for papers published in both, html and pdf forms. To view the papers in pdf format, click on the "PDF Full-text" link, and use the free Adobe Reader to open them.
Order results
Result details
Section
Select all
Export citation of selected articles as:
58 pages, 20169 KB  
Systematic Review
Detecting Contradictions in Bilingual Legislation: A Systematic Review and Conceptual Architecture
by Maxatbek Satymbekov, Arman Yeleussinov, Zholdas Buribayev, Nurbol Beisov, Nurlykhan Kalzhanov, Yerbol Alimkulov and Nazerke Serikzhankyzy Zhumadilova
Information 2026, 17(9), 931; https://doi.org/10.3390/info17090931 - 21 Sep 2026
Viewed by 490
Abstract
Ensuring legislative consistency is increasingly difficult as legal systems expand and change more frequently, and the problem is sharper where two language versions are equally binding. This systematic review examines artificial intelligence applied to legal and legislative texts, following the PRISMA 2020 guidelines. [...] Read more.
Ensuring legislative consistency is increasingly difficult as legal systems expand and change more frequently, and the problem is sharper where two language versions are equally binding. This systematic review examines artificial intelligence applied to legal and legislative texts, following the PRISMA 2020 guidelines. Searches of Scopus and the Web of Science Core Collection for English-language peer-reviewed journal articles (January 2019–May 2026) returned 497 records, of which 231 studies met the inclusion criteria; 102 supplied the core analytical evidence and 129 characterized the wider research landscape. Publication output rose sharply after 2024 without a matching improvement in reporting and methodological completeness. Within the reviewed corpus, legal summarization and judgment prediction are the most established directions, whereas natural language inference, legal knowledge representation and compliance verification remain underdeveloped. Only four studies released code or datasets, and although thirteen addressed multilingual legal texts, none of them examined bilingual legislative corpora in which both versions carry equal legal force. Work in the reviewed corpus therefore advances individual legal natural language processing (NLP) tasks but offers no integrated framework for detecting contradictions in bilingual legislation. We propose a conceptual four-layer architecture combining legislative text normalization, knowledge-graph representation, hybrid semantic reasoning and explainable human oversight, and identify research priorities for reliable and transparent AI-assisted legislative analysis. Full article
►▼ Show Figures

Figure 1

40 pages, 9036 KB  
Article
Uncertainty-Calibrated Residual Conformal Monitoring of Wind Turbine SCADA Data for Cross-Asset Anomaly Detection Under Distribution Shift
by Zalan Haneef, Muhammad Umar, Faisal Saleem, Ikram Ullah, Kamran Aqeel and Muhammad Farooq Siddique
Information 2026, 17(9), 930; https://doi.org/10.3390/info17090930 - 21 Sep 2026
Viewed by 349
Abstract
Wind turbine SCADA anomaly detectors are commonly calibrated using data from the same asset or development period, whereas deployment requires transfer across turbines with different operating distributions. This mismatch can produce optimistic thresholds, excessive false alarms, and misleading generalization estimates. This study proposes [...] Read more.
Wind turbine SCADA anomaly detectors are commonly calibrated using data from the same asset or development period, whereas deployment requires transfer across turbines with different operating distributions. This mismatch can produce optimistic thresholds, excessive false alarms, and misleading generalization estimates. This study proposes UC-RCF, an uncertainty-calibrated residual conformal framework integrating nonlinear multi-output normal behavior modeling, embargoed blocked cross-fitting, uncertainty-normalized residuals, operating support assessment, channel-wise conformal evidence, reflected cumulative criticality, and cross-asset alarm calibration. Evaluation followed a nested leave-one-turbine-out, asset-disjoint protocol on CARE v6, comprising 95 monitoring cases from 36 turbines across three wind farms, including 45 anomalous and 50 normal cases. UC-RCF achieved a pooled outer-fold CARE score of 0.5903, fault coverage of 0.3920, weighted earliness of 0.2423, event-level reliability of 0.5760, and normal-case accuracy of 0.8706. It detected 25 anomalous cases and generated alarms in 18 normal cases. A farm-stratified paired turbine-cluster bootstrap estimated a CARE improvement of 0.0349 over the initial full configuration, with a 95% percentile interval of 0.0055–0.0671 and a bootstrap probability of improvement of 0.991. These estimates remain exploratory because the effective resampling units comprise only 36 physical turbines from three wind farms. Mahalanobis monitoring achieved a higher CARE score of 0.6012, detecting 21 anomalous cases while generating alarms in nine normal cases. UC-RCF therefore provided broader fault coverage and four additional anomalous-case detections but incurred a higher false-alarm burden, demonstrating a sensitivity–reliability trade-off rather than universal detector dominance. Ablation analysis identified farm-normalized residual magnitude as the strongest case-level discriminator, with a receiver operating characteristic area under the curve of 0.854. Heteroscedastic uncertainty scaling and operating support adjustment did not consistently improve CARE or raw discrimination; the auxiliary framework components instead provide mechanisms for uncertainty characterization, score comparability, temporal persistence, and calibration auditing. At the CARE-optimal nominal case-level false-alarm budget of α=0.30, the empirical normal-case false-alarm rate was 0.36. Because temporal dependence and cross-asset distribution shift can violate exchangeability, α is interpreted as an operational calibration target rather than a theoretically guaranteed case-level error bound. Overall, UC-RCF provides an interpretable and auditable framework for investigating residual evidence, operating support shift, temporal persistence, and calibration reliability under cross-asset deployment. Full article
►▼ Show Figures

Graphical abstract

47 pages, 18995 KB  
Article
O-M-Flow: Ontology-Constrained Evidence Retrieval for RAG over Aviation Maintenance Documents
by Qiang Cui, Yitong Zhang, Shudong An and Yunwei Dong
Information 2026, 17(9), 929; https://doi.org/10.3390/info17090929 - 21 Sep 2026
Viewed by 258
Abstract
Aviation maintenance questions often require evidence scattered across several passages, including possible causes, exclusion statements, operating conditions, and manual references. This study presents O-M-Flow, an ontology-constrained evidence retrieval method for retrieval-augmented generation (RAG). A lightweight, task-oriented ontology defines the roles and attributes of [...] Read more.
Aviation maintenance questions often require evidence scattered across several passages, including possible causes, exclusion statements, operating conditions, and manual references. This study presents O-M-Flow, an ontology-constrained evidence retrieval method for retrieval-augmented generation (RAG). A lightweight, task-oriented ontology defines the roles and attributes of maintenance evidence. Document content is organized into an Episode–Facet–FacetPoint–Entity graph, and these attributes guide candidate selection, ranking, and evidence packaging. The ontology is manually defined by the research team and populated through deterministic rules. Evaluation uses three documents containing 457 pages and 100 question records, corresponding to 77 distinct questions and 46 groups of related or repeated questions. Under a shared text-embedding configuration, O-M-Flow with its original scoring procedure achieved a groundedness score of 0.8939, compared with 0.7366 for the keyword-retrieval baseline. An additional version implementing the explicit weighted score achieved 0.8667. Its source-page recall was 0.9809, while key-role coverage was 0.5352. Weight sensitivity, repeated reranking, and computational-cost measurements further characterize the method. The results support role-aware evidence organization on this benchmark and show that the additional explicit score does not consistently improve retrieval. Incomplete role coverage and limited independent human evaluation remain important constraints on interpretation. Full article
►▼ Show Figures

Figure 1

19 pages, 737 KB  
Article
Cross-Lingual Transfer for Mammography Report Classification in Low-Resource Settings
by Anuar Dosmaganbetov, Tomiris Zhaksylyk and Beibit Abdikenov
Information 2026, 17(9), 928; https://doi.org/10.3390/info17090928 - 21 Sep 2026
Viewed by 213
Abstract
Labeled clinical text is scarce in many non-English healthcare settings, limiting the development of robust clinical NLP systems. We tested whether supervision from Spanish mammography reports improved the classification of Russian-language reports from Kazakhstan. The study included 4279 Spanish and 495 Russian reports [...] Read more.
Labeled clinical text is scarce in many non-English healthcare settings, limiting the development of robust clinical NLP systems. We tested whether supervision from Spanish mammography reports improved the classification of Russian-language reports from Kazakhstan. The study included 4279 Spanish and 495 Russian reports mapped to three BI-RADS-derived operational classes (routine, follow-up, and suspicious). Explicit BI-RADS identifiers were removed from the input, and Russian performance was assessed using grouped five-fold out-of-fold evaluation. We compared Russian-only XLM-R fine-tuning, Spanish zero-shot transfer, sequential Spanish-to-Russian XLM-R fine-tuning, and a matched word/character TF–IDF logistic-regression baseline. Under the fixed split, label budget, and four-epoch schedule tested here, sequential transfer exceeded Russian-only XLM-R in all three paired training seeds, increasing the mean macro-F1 from 0.2178 to 0.3239; zero-shot performance was less stable (mean 0.1898). However, the matched TF–IDF model achieved the highest macro-F1 (0.5214) and better performance on the rare suspicious class (F1 0.3306 versus 0.1139 for transferred XLM-R). Thus, Spanish initialization was compatible with subsequent low-resource Russian adaptation in this experiment, while the target-language lexical model remained stronger. These findings are limited to the two datasets, grouped split, XLM-R configuration, and small, imbalanced target-data regime evaluated here. Full article
►▼ Show Figures

Figure 1

36 pages, 5552 KB  
Article
Time-Varying Privacy Scheduling for Communication-Efficient Differentially Private Federated Learning on Heterogeneous Educational Data
by Guanglei Li and Peng Chen
Information 2026, 17(9), 927; https://doi.org/10.3390/info17090927 - 21 Sep 2026
Viewed by 275
Abstract
Heterogeneous client distributions and differential-privacy noise can slow learning. This study proposes time-varying privacy-scheduled differentially private federated learning (TVP-DPFL), combining a public Gaussian multiplier schedule, bounded reliability-weighted signals, and control-variate drift correction. Inverse-square normalization gives the time-varying and fixed schedules identical final Gaussian [...] Read more.
Heterogeneous client distributions and differential-privacy noise can slow learning. This study proposes time-varying privacy-scheduled differentially private federated learning (TVP-DPFL), combining a public Gaussian multiplier schedule, bounded reliability-weighted signals, and control-variate drift correction. Inverse-square normalization gives the time-varying and fixed schedules identical final Gaussian Rényi differential-privacy costs. On the Open University Learning Analytics Dataset (OULAD) and the Predict Students’ Dropout and Academic Success dataset from the University of California, Irvine repository (UCI-Dropout), the three-seed endpoint results are close to a differentially private variant of Stochastic Controlled Averaging for Federated Learning (DP-SCAFFOLD) under the common 30-round aggregate-update budget of ε = 8. The ten-seed scheduling comparison shows earlier attainment of the 95% validation target: 11.1 versus 12.4 rounds on OULAD and 16.4 versus 18.7 rounds on UCI-Dropout (two-sided Wilcoxon p = 0.002). The normalized area under the validation learning curve over 30 rounds (AULC30) also favors the time-varying schedule. The privacy guarantee applies to aggregate updates conditional on fixed bounded weights and control states, with auxiliary metadata and control-state releases requiring separate protection and accounting. Communication efficiency is measured by coordinated rounds and model/control-vector payloads on retrospective public-data silos. Full article
►▼ Show Figures

Figure 1

22 pages, 414 KB  
Article
Source-Linked Retrieval for Context-Dependent Short Clauses in Digitized Chemical-Safety Standards
by Kaifeng Sun, Linmao Duan, Xuesheng Li and Yi Yang
Information 2026, 17(9), 926; https://doi.org/10.3390/info17090926 - 21 Sep 2026
Viewed by 222
Abstract
Digitizing standards does not by itself make their provisions reliably retrievable. Many requirements are expressed as short clauses whose regulated object, scope, list context, table value, or referenced condition is supplied elsewhere in the document. This study examines that indexing problem in a [...] Read more.
Digitizing standards does not by itself make their provisions reliably retrievable. Many requirements are expressed as short clauses whose regulated object, scope, list context, table value, or referenced condition is supplied elsewhere in the document. This study examines that indexing problem in a Neo4j deployment containing 9206 processed safety-related standards, 991,453 Sections, and 70,016 Tables. A repaired provenance layer maps 248,454 historical QAPair source records to unique Sections and exposes 1,207,059 question variants as attributable retrieval entries; 32,714 unresolved source records remain quarantined. We evaluate a source-linked design that searches Sections and QAPair fields in parallel and applies citation or dependency traversal only after query-time gating. On a frozen 120-query confirmation set, Section BM25 achieved Recall@10 of 0.792, deterministic context-enriched Section BM25 achieved 0.875, and QAPair question-plus-answer BM25 achieved 0.933. Question-only and answer-only Recall@10 were 0.833 and 0.900, respectively, showing that the gain cannot be attributed to question expansion alone. BGE-M3 Section retrieval, QAPair-question retrieval, and their dense fusion reached 0.800, 0.833, and 0.867, while BM25 top-100 retrieval followed by BGE reranking reached 0.908. On 120 ordinary queries not preselected for relations, the graph gate activated for 5 queries, improved none, and lowered one source rank; on 18 relation-dependent queries, the same gate increased complete two-clause recovery from 5/18 to 11/18. Two independent chemical-safety experts additionally reviewed selected context-restoration, query, graph, and end-to-end cases. These results support source-linked multi-entry retrieval as a bounded mechanism for this corpus while showing that deterministic context, answer fields, and selective graph activation must be evaluated separately. Full article
(This article belongs to the Special Issue Standards Digitisation and Digital Standardisation)
►▼ Show Figures

Figure 1

35 pages, 7286 KB  
Review
Foundation-Model-Assisted Reward Design for Reinforcement Learning: A Review of Reward Program Synthesis, Multimodal Feedback, and Trustworthiness
by Wei Zhu, Jinyin Bai, Rui Tang, Zehao Pang, Mingxi Wang, Chengjie Lu, Tianjin Ni, Hang Liu, Xiangchen Wang, Jinji Zhou, Yanlin Wu, Yongjun Peng, Zongzhe Nie, Shiluo Guo, Qinglin Xu, Kaiyang Kou and Yihao Zhong
Information 2026, 17(9), 925; https://doi.org/10.3390/info17090925 - 20 Sep 2026
Viewed by 241
Abstract
Reward functions determine what reinforcement learning agents ultimately optimize, yet reward design for complex tasks has traditionally relied on extensive domain expertise and iterative engineering. Recent large language models and vision–language foundation models have introduced new mechanisms for interpreting task intent, synthesizing reward [...] Read more.
Reward functions determine what reinforcement learning agents ultimately optimize, yet reward design for complex tasks has traditionally relied on extensive domain expertise and iterative engineering. Recent large language models and vision–language foundation models have introduced new mechanisms for interpreting task intent, synthesizing reward programs, evaluating states and trajectories, and refining rewards through policy feedback. This review organizes the emerging literature along three complementary directions: reward program synthesis, multimodal feedback, and feedback-driven reward optimization. We further propose a five-level trustworthiness framework spanning format validity, execution validity, semantic validity, behavioral validity, and structural assurance. Existing evidence shows that foundation models substantially broaden how rewards can be represented and acquired but do not eliminate grounding errors, proxy misalignment, reward hacking, selection bias, or reward-search costs. We therefore examine the field from an end-to-end perspective that jointly considers policy performance, reward fidelity, trustworthiness, computational and human cost, and transfer. Finally, we identify verifiable reward representations, process reward models, budget-aware reward search, and transferable reward knowledge across tasks and multi-agent systems as key directions for future research. Full article
►▼ Show Figures

Figure 1

33 pages, 2261 KB  
Article
Developing an AI-Assisted Methodology for Analyzing Documented Procedural Openness in Academic Recruitment Announcements
by Walery Okulicz-Kozaryn, Artem Artyukhov and Nadiia Artyukhova
Information 2026, 17(9), 924; https://doi.org/10.3390/info17090924 - 20 Sep 2026
Viewed by 332
Abstract
Academic recruitment announcements are the primary public source of information on competitive procedures. However, their unstructured nature significantly complicates the systematic analysis of procedural characteristics and comparison of academic recruitment announcements published by different universities and in different scientific disciplines. This study aims [...] Read more.
Academic recruitment announcements are the primary public source of information on competitive procedures. However, their unstructured nature significantly complicates the systematic analysis of procedural characteristics and comparison of academic recruitment announcements published by different universities and in different scientific disciplines. This study aims to develop an AI-assisted research methodology to analyze unstructured texts in academic recruitment announcements, grounded in procedural and ethical criteria. The methodology enables the formation of a standardized procedural profile for each academic recruitment announcement, the identification of procedural and ethical risks, and the presentation of the results in a unified analytical format. The assigned scores are verified by the researcher for consistency with the original text and the uniform application of the criteria. Additionally, independent expert validation of 40 analytical decisions from five disciplinary corpora showed that 38 of 40 classifications (95.0%) were confirmed without changes. To provide an integral characteristic of an individual academic recruitment announcement, the Procedural Openness Score (POS) indicator is proposed to reflect the proportion of criteria classified as low risk. The application of the methodology is demonstrated on five independent samples of academic recruitment announcements from various scientific disciplines. The empirical part demonstrates the applicability of a unified analytical architecture to various corpora of academic recruitment announcements. The interpretability of the results is ensured by comparing the original fragments of academic recruitment announcements with the assigned procedural risk scores. The proposed methodology expands the potential for using artificial intelligence in research practice by combining transparent analysis rules, a standardized procedural academic recruitment announcement profile, and a scalable procedural transparency metric. Although the methodology was tested on academic recruitment announcements, its architecture allows for application to the analysis of other types of organizational and regulatory documents. Full article
(This article belongs to the Special Issue Advancing Educational Innovation with Artificial Intelligence)
►▼ Show Figures

Graphical abstract

56 pages, 15896 KB  
Article
Governing Agentic AI in Enterprise Workflows: A Bounded-Autonomy Framework for Delegated Authority and Controlled Execution
by Bo Nørregaard Jørgensen and Zheng Grace Ma
Information 2026, 17(9), 923; https://doi.org/10.3390/info17090923 - 20 Sep 2026
Viewed by 446
Abstract
Agentic AI can interpret information, plan, make workflow decisions, and use enterprise tools. Yet technical capability does not establish authoritative meaning, legitimate process state, organisational permission, or accountable execution. The challenge is to preserve adaptability while ensuring that consequential actions remain governed. This [...] Read more.
Agentic AI can interpret information, plan, make workflow decisions, and use enterprise tools. Yet technical capability does not establish authoritative meaning, legitimate process state, organisational permission, or accountable execution. The challenge is to preserve adaptability while ensuring that consequential actions remain governed. This article develops a domain-independent conceptual framework for governed agentic AI in enterprise workflows based on bounded autonomy, delegated authority, and controlled execution. A consequential action or workflow decision selected by an agent is treated as a proposed action. It may change enterprise state only after independent controls confirm semantic validity, procedural admissibility, policy compliance, and delegated authority. Actions that pass these controls and remain within a task envelope may proceed automatically through controlled enterprise tools. Those exceeding thresholds for consequence, irreversibility, uncertainty, data sensitivity, value, or organisational policy are escalated to an accountable human. Human oversight is therefore risk-proportionate rather than required for every action. The framework integrates enterprise ontologies, governed knowledge graphs, Business Process Model and Notation (BPMN) orchestration, policy and decision services, controlled tool execution, and provenance within explicit responsibility and authority boundaries. A review-informed design-science process synthesises evidence into five connected control gaps and derives ten design requirements, operationalised through task envelopes, capability and authority relations, lifecycle states, exception paths, and conformance criteria. An illustrative online-shopping order exception demonstrates the control logic, while a control-loop walkthrough and ten failure and adversarial conditions trace requirements-to-control mappings, responsibility separation, and defined recovery paths. The framework provides a systematic basis for governing agentic AI as an adaptable enterprise participant. It supports risk-proportionate autonomy, auditability, accountability, and regulatory evidence. The analytical walkthrough supports conceptual coherence and design plausibility but does not establish deployed effectiveness or legal compliance. Full article
(This article belongs to the Special Issue Intelligent Agent and Multi-Agent System, 2nd Edition)
►▼ Show Figures

Figure 1

26 pages, 16066 KB  
Article
Real-Traffic Enrichment for Improved Minority Web Attack Detection in Network Intrusion Detection
by Zeyneb Berkat, Amina Fatima Zahra Yahiaoui, Mahfoud Aliouat, Emad Abd-Elrady, Aymen Bendjebbas, Kamel Eddine Haouari and Riyadh Bouddou
Information 2026, 17(9), 922; https://doi.org/10.3390/info17090922 - 20 Sep 2026
Viewed by 263
Abstract
Class imbalance severely limits Network Intrusion Detection Systems (NIDSs) for minority Web attack classes: CICIDS2017 contains only 21 SQL Injection instances among 2.27 million benign flows. This study enriches CICIDS2017 with authentic SQL Injection, Cross-Site Scripting (XSS), and Web Brute Force (WBF) traffic [...] Read more.
Class imbalance severely limits Network Intrusion Detection Systems (NIDSs) for minority Web attack classes: CICIDS2017 contains only 21 SQL Injection instances among 2.27 million benign flows. This study enriches CICIDS2017 with authentic SQL Injection, Cross-Site Scripting (XSS), and Web Brute Force (WBF) traffic captured from a controlled DVWA/XAMPP environment, processed with CICFlowMeter to match the original feature space. An anti-data-leakage protocol (stratified partitioning, post-split normalization, five-fold cross-validation, and a SHA-1 cryptographic membership audit of an 8881 –flow test sub-sample) found no hash collisions between this sub-sample and the evaluation partitions. The framework added 32,670 authentic flows, increasing SQL Injection from 21 to 10,678, XSS from 652 to 13,212, and WBF from 1507 to 10,960. Among four evaluated ensemble models, LightGBM performed best, achieving 99.85% Accuracy, 99.85% F1-score, 99.29% Balanced Accuracy, and 97.87 ± 1.88% in five-fold cross-validation, improving detection rates by 44.9% (XSS), 23.0% (WBF), and 16.6% (SQL Injection) over the original dataset. A volume-matched ablation study showed comparable aggregate accuracy to synthetic balancing methods (SMOTE, SMOTE-Tomek), while geometric diversity analysis confirmed that authentic traffic occupies feature-space regions unreachable by interpolation, and chronological holdout evaluation confirmed generalization to unseen traffic (F1: 98.53–99.90%). Real-traffic enrichment thus offers a practical, more realistic complement to synthetic balancing for minority Web-attack detection. Full article
(This article belongs to the Topic New Trends in Cybersecurity and Data Privacy)
►▼ Show Figures

Graphical abstract

30 pages, 562 KB  
Article
When Does Fine-Tuning Matter for a Low-Resource Language? Evaluating Open-Weight LLMs on Kazakh Benchmarks
by Aman Mussa, Zhanseit Tuimebayev, Madina Mansurova and Ahsan Habib Shihab
Information 2026, 17(9), 921; https://doi.org/10.3390/info17090921 - 20 Sep 2026
Viewed by 349
Abstract
Fine-tuning is often treated as a prerequisite for deploying large language models in low-resource languages, yet its value for recent open-weight models is unclear. We evaluate eleven checkpoints of 7–14 billion parameters before and after Low-Rank Adaptation on 5000 human-curated Kazakh instruction–response pairs. [...] Read more.
Fine-tuning is often treated as a prerequisite for deploying large language models in low-resource languages, yet its value for recent open-weight models is unclear. We evaluate eleven checkpoints of 7–14 billion parameters before and after Low-Rank Adaptation on 5000 human-curated Kazakh instruction–response pairs. Evaluation is performed on the Kazakh splits of KazMMLU, ARC, GPQA, MMLU-Pro, and GSM8K, with Russian GSM8K as a higher-resource comparison. Our central finding is a language asymmetry: the same Kazakh-trained adapters improved Russian GSM8K more than the Kazakh GSM8K they targeted. Excluding one checkpoint that collapsed after fine-tuning by up to 58 percentage points, the remaining ten gained +2.6 percentage points on average on Kazakh (range −6.1 to +13.6) against +6.9 on Russian, significant under a paired signed-rank test. A manual audit of 400 responses identifies a contributing mechanism: Kazakh outputs were more often truncated before a final answer, and correct Kazakh reasoning more often scored wrong by answer extraction, so part of the gap reflects response form, not reasoning. Adaptation produced no meaningful average change on the four multiple-choice benchmarks. GSM8K is the suite’s only free-form task, so generality is untested. These results suggest that low-resource fine-tuning should be evaluated per checkpoint, benchmark, and language rather than being adopted by default. Full article
►▼ Show Figures

Graphical abstract

20 pages, 289 KB  
Article
Evaluating Prompt Engineering Techniques for LLaMA-3: A Study of Zero-Shot, Few-Shot, and Chain-of-Thought Prompts Across Reasoning and Classification Tasks
by Darren Astle Travasso and Aboozar Taherkhani
Information 2026, 17(9), 920; https://doi.org/10.3390/info17090920 - 20 Sep 2026
Viewed by 274
Abstract
Prompt engineering has emerged as a practical and resource-efficient alternative to fine-tuning large language models (LLMs), particularly as these methods have a lower computation cost than fine-tuning. In this paper, three widely adopted prompting techniques—Zero-Shot, Few-Shot, and Chain-of-Thought (CoT)—were assessed. While these prompting [...] Read more.
Prompt engineering has emerged as a practical and resource-efficient alternative to fine-tuning large language models (LLMs), particularly as these methods have a lower computation cost than fine-tuning. In this paper, three widely adopted prompting techniques—Zero-Shot, Few-Shot, and Chain-of-Thought (CoT)—were assessed. While these prompting strategies are well established, practitioners still lack clear guidance on when each technique should be preferred, which types of tasks they fail to support reliably, and how performance trade-offs may affect the practical use of LLM-based systems. These techniques are tested across three benchmark tasks: sentiment classification (SST-2), multiple-choice questions (CommonsenseQA), and multi-step math problem solving (GSM8K) using Meta’s LLaMA-3 8B Instruct model. We provide a thorough performance comparison based on accuracy, F1 score, and solve rate. The solve rate is highlighted as a complementary metric for evaluating the usability of LLM outputs—a factor often overlooked in the existing literature. Experimental results showed that Few-Shot prompts are particularly effective in structured classification tasks, while CoT prompts excel in logic-heavy tasks that require multi-step reasoning. On the classification task, Few-Shot prompting improved the solve rate but achieved lower accuracy and F1 score than Zero-Shot prompting. On the multiple-choice questions, Zero-Shot, Few-Shot, and CoT prompting achieved a solve rate of 100%. On multi-step math problem solving, CoT improved the solve rate and interpretability compared to Zero-Shot but did not surpass Zero-Shot accuracy. Overall, the results demonstrate that the effectiveness of prompting strategies is task-dependent, with differences observed in both accuracy and output validity across the three benchmark tasks. Full article
24 pages, 2196 KB  
Article
Enhancing Rule-Based Explanations via Cognitive Bias Towards Semantic Relevance
by Parisa Mahya and Johannes Fürnkranz
Information 2026, 17(9), 919; https://doi.org/10.3390/info17090919 - 19 Sep 2026
Viewed by 251
Abstract
As interest in explainable artificial intelligence (XAI) continues to grow, a critical gap remains between the explanations generated by models and their interpretability by human users, particularly when cognitive biases and semantic relevance are not adequately addressed. This paper introduces CoRIfEE-Rel, a [...] Read more.
As interest in explainable artificial intelligence (XAI) continues to grow, a critical gap remains between the explanations generated by models and their interpretability by human users, particularly when cognitive biases and semantic relevance are not adequately addressed. This paper introduces CoRIfEE-Rel, a novel human-centered meta-XAI method aimed at bridging this gap by producing rule-based explanations that are both interpretable and semantically aligned with the target domain concept. CoRIfEE-Rel synthesizes outputs from a diverse pool of interpretable models and employs a knowledge graph-driven heuristic that combines semantic relevance with traditional rule learning metrics. This approach ensures that the resulting explanations are deeply tied to core domain concepts while maintaining the clarity and structure needed for human understanding. Empirical evaluations conducted across multiple datasets demonstrate that CoRIfEE-Rel achieves higher semantic relevance than random forest and JRip rule-based explanations without notable compromises in predictive accuracy. The results highlight the ability of CoRIfEE-Rel to generate rule-based explanations that are semantically related to the target concepts while maintaining predictive performance. Full article
(This article belongs to the Section Artificial Intelligence)
►▼ Show Figures

Figure 1

35 pages, 2651 KB  
Article
Comparative Evaluation of AI Programming Assistants: An Exploratory Longitudinal Case Study of GitHub Copilot, ChatGPT, and Cursor Configurations in Full-Stack Development
by Goran Đambić, Anton Maurovic, Ivana Ogrizek Biškupić and Aleksander Radovan
Information 2026, 17(9), 918; https://doi.org/10.3390/info17090918 - 19 Sep 2026
Viewed by 470
Abstract
The rapid adoption of artificial intelligence (AI) programming assistants has raised questions about the actual benefits they provide in professional software development. This exploratory longitudinal case study compares configurations of three AI programming assistants (GitHub Copilot, ChatGPT, and Cursor)—specific combinations of tool, underlying [...] Read more.
The rapid adoption of artificial intelligence (AI) programming assistants has raised questions about the actual benefits they provide in professional software development. This exploratory longitudinal case study compares configurations of three AI programming assistants (GitHub Copilot, ChatGPT, and Cursor)—specific combinations of tool, underlying model, interaction interface, and period of use—by reimplementing a full-stack thesis management application (a .NET Core 8.0 representational state transfer (REST) application programming interface (API) and a React.js client with 25 functionalities) that was first developed manually to establish a baseline. Each tool was evaluated using a seven-criteria framework covering code correctness, prompt complexity, context awareness, number of prompts, bug count, bug severity, and recorded implementation time. All three configurations significantly reduced the total recorded implementation time relative to manual implementation (by 66%, 63%, and 83% for Copilot, ChatGPT, and Cursor, respectively; p < 0.001). The Cursor configuration ranked best on four of the five evaluation criteria, with significantly higher context awareness and significantly fewer bugs than Copilot; because the tools were applied in a fixed order between January and November 2025, a period during which the underlying models were upgraded, and Cursor was both applied last and received the largest such upgrade, these results reflect tool–model–interface–time configurations rather than the tools in isolation. All tools performed significantly worse on the multi-layer API than on the client application. All 336 recorded bugs were organized into a fourteen-category taxonomy, in which hallucinated code elements and incomplete multi-file modifications were the most frequent failure modes, a pattern consistent with the tools’ limited ability to track context across files and architectural layers. The results indicate that AI programming assistants are most effective as pair programming tools whose output requires systematic human review before integration. Full article
►▼ Show Figures

Figure 1

33 pages, 4709 KB  
Review
Entity Resolution Using Transformer-Based Language Models: A Systematic Scoping Review
by Mohammad Beheshti, Maryam Seifaddini, Amir Erfan Zareei Shams Abadi, Karan Karthik, Tarun Mummidi Ramesh Kumar, Suzanne Austin Boren and Iris Zachary
Information 2026, 17(9), 917; https://doi.org/10.3390/info17090917 - 18 Sep 2026
Viewed by 573
Abstract
Entity resolution (ER) is fundamental to integrating heterogeneous data, which traditional approaches address through deterministic rule-based methods and probabilistic record linkage. We conducted a systematic scoping review following PRISMA-ScR to characterize the use of transformer-based language models for ER. Five databases were searched, [...] Read more.
Entity resolution (ER) is fundamental to integrating heterogeneous data, which traditional approaches address through deterministic rule-based methods and probabilistic record linkage. We conducted a systematic scoping review following PRISMA-ScR to characterize the use of transformer-based language models for ER. Five databases were searched, and 155 studies were included for synthesis. The literature expanded sharply after 2023, with 55% of included studies published in 2025 or 2026. General-purpose matching was the most common entity focus, followed by product/e-commerce. Healthcare applications were especially scarce, with only one study applying ER to the patient/healthcare domain. Among studies that performed blocking, embedding-based similarity was most common, followed by string-based approaches. Classification-head approaches remained the most common matching approach, followed by prompt-based approaches. Encoder-only models remained the most widely evaluated architecture, while decoder-only models grew increasingly prominent from 2023 onward. Full-parameter fine-tuning was the predominant learning strategy, followed by zero-shot and few-shot prompting. Among studies classified as general-purpose, nearly one-third were evaluated on only one or two entity types, limiting their generalizability. Reported best F1 scores varied across benchmarks, with no consistent advantage for encoder-only versus decoder-only architectures. These findings support the need for more diverse, end-to-end evaluations that consider efficiency and robustness alongside accuracy. Full article
►▼ Show Figures

Figure 1

32 pages, 3636 KB  
Article
From Digital Tools to Digital Treatment Ecosystems: A Closed-Loop Information Architecture for Precision Addiction Psychiatry
by Vincenzo Maria Romeo, Bruna Caridi, Agnese Tedeschi and Elisabetta Ratti
Information 2026, 17(9), 916; https://doi.org/10.3390/info17090916 - 18 Sep 2026
Viewed by 298
Abstract
Digital technologies are increasingly used in addiction care, yet telemedicine, ecological momentary assessment, mobile applications, wearables, electronic health records, digital therapeutics, and artificial intelligence commonly remain fragmented across separate platforms and clinical workflows. This Concept Paper addresses this information–integration gap by proposing the [...] Read more.
Digital technologies are increasingly used in addiction care, yet telemedicine, ecological momentary assessment, mobile applications, wearables, electronic health records, digital therapeutics, and artificial intelligence commonly remain fragmented across separate platforms and clinical workflows. This Concept Paper addresses this information–integration gap by proposing the Digital Treatment Ecosystem (DTE), a person-centered, closed-loop information architecture for precision addiction psychiatry. The framework was developed through an integrative narrative synthesis of addiction, digital-health, health-informatics, artificial-intelligence, implementation, and regulatory literature. Its originality lies not in any individual technology, but in specifying a governed information-to-action cycle in which heterogeneous longitudinal data are integrated with provenance and uncertainty, interpreted against population-level and within-person baselines, translated into explicitly owned clinician-supervised actions, and returned as outcome feedback for treatment adaptation and organizational learning. Five functional layers are proposed: multimodal data acquisition; integration and interoperability; adaptive intelligence; clinical decision support; and intervention delivery with outcome feedback. Addiction-specific requirements include dynamic craving and recurrence risk, treatment disengagement, polysubstance use, medication continuity, overdose and withdrawal risk, stigma, confidentiality, and fragmented service pathways. Six operationalized propositions define how the DTE can be prospectively tested. The DTE is therefore proposed as a falsifiable socio-technical architecture rather than an established or clinically validated treatment system. Full article
(This article belongs to the Special Issue Information Technology for Smart Healthcare)
►▼ Show Figures

Figure 1

28 pages, 2456 KB  
Article
From Global to Context-Specific Process Views: A Configurable Process Mining Framework for Hospital Billing Analysis
by Imane El Alama, Hanae Sbai and Soumaya El Mamoune
Information 2026, 17(9), 915; https://doi.org/10.3390/info17090915 - 18 Sep 2026
Viewed by 248
Abstract
Hospital billing event logs combine shared administrative routines with context-specific behavior, making a single global process model difficult to interpret. This study proposes a configurable process mining framework that complements a hospital-wide model with context-specific views derived from Hospital Billing data. Seventeen medical [...] Read more.
Hospital billing event logs combine shared administrative routines with context-specific behavior, making a single global process model difficult to interpret. This study proposes a configurable process mining framework that complements a hospital-wide model with context-specific views derived from Hospital Billing data. Seventeen medical specialties (97,753 cases; 437,444 events) were retained and split temporally into 70% Train and 30% Test cases. Train data were used for behavioral representation, similarity analysis, hierarchical clustering, process discovery, merging, and configuration. The selected two-cluster solution (mean silhouette = 0.8960) separated specialty K from the remaining 16 specialties. Global Train analysis showed a tie among IMf thresholds 0.40, 0.50, and 0.60; 0.60 was retained as the common configuration threshold because it reduced configuration points from 25 to 20. Cluster-specific Process Trees were merged into a configurable model with 41 unique nodes and 20 configurable relations, and Genetic search produced Derived Process Trees. On held-out Test data, the Derived K model achieved fitness/balanced precision of 0.9992/0.6690, while the Derived Others model achieved 0.8598/0.9305. Both remained competitive with independently mined local baselines while preserving shared and context-dependent behavior within one process family. Full article
(This article belongs to the Special Issue Machine Learning and Data Analytics for Business Process Improvement)
►▼ Show Figures

Figure 1

36 pages, 782 KB  
Review
Assessment Instruments for Artificial Intelligence Literacy in Healthcare Professionals: A Scoping Review
by Sergio Mies-Padilla, Claudio-Alberto Rodríguez-Suárez and Héctor González-de la Torre
Information 2026, 17(9), 914; https://doi.org/10.3390/info17090914 - 18 Sep 2026
Viewed by 285
Abstract
Artificial Intelligence is being integrated into clinical practice within regulatory frameworks that assign healthcare professionals responsibility for human oversight, making AI literacy a core professional competency; however, there is little consensus on how it should be measured. This scoping review maps the instruments [...] Read more.
Artificial Intelligence is being integrated into clinical practice within regulatory frameworks that assign healthcare professionals responsibility for human oversight, making AI literacy a core professional competency; however, there is little consensus on how it should be measured. This scoping review maps the instruments used to assess AI literacy among healthcare professionals, together with the domains they cover and the psychometric properties they report. Following the Joanna Briggs Institute framework and PRISMA-ScR, five databases were searched for primary studies published from 2017 onwards in English or Spanish. Screening was supported by an active-learning framework, and methodological quality was appraised with JBI tools. Thirty-nine studies were included, predominantly involving nurses and physicians. Methodological quality was heterogeneous, with the main weaknesses concentrated in the identification and handling of confounding factors. The instruments fell into validated standardized scales and study-specific ad hoc questionnaires that seldom report structural validation. Item-level deconstruction suggested a two-level thematic framework: three competence domains—cognitive-conceptual (82.1%), practical-clinical (69.2%), and ethical-regulatory (56.4%)—and two co-assessed domains corresponding to dispositions and organizational context (84.6% and 38.5%). Measurement in this field is fragmented, and reliance on unvalidated instruments limits international comparability. Assessment and training should prioritize critical appraisal of outputs, ethical data governance, and human oversight. Full article
(This article belongs to the Special Issue Artificial Intelligence-Based Digital Health Emerging Technologies)
►▼ Show Figures

Figure 1

22 pages, 671 KB  
Article
CNN Architecture Optimization for Multiplierless Inference
by Ivan Al Khayat, John Reuben and Dietmar Fey
Information 2026, 17(9), 913; https://doi.org/10.3390/info17090913 - 17 Sep 2026
Viewed by 203
Abstract
Many CNN architectures are primarily optimized for classification accuracy and software-level metrics, but they do not necessarily account for resource-constrained integer inference. This work addresses this gap by optimizing CNN architectures with subsequent In-Memory Computing (IMC) implementation in mind, using an analytical resource–cost [...] Read more.
Many CNN architectures are primarily optimized for classification accuracy and software-level metrics, but they do not necessarily account for resource-constrained integer inference. This work addresses this gap by optimizing CNN architectures with subsequent In-Memory Computing (IMC) implementation in mind, using an analytical resource–cost model for multiplierless integer inference. Motivated by the need to reduce data movement and avoid explicit multipliers, the proposed methodology constrains candidate CNN architectures toward integer computations that can be expressed through memory look-ups, additions, and shifts, rather than conventional multiply–accumulate operations. Distributed Arithmetic provides the computational basis for this multiplierless formulation, while the proposed HWCost model is used together with validation performance and validation efficiency to rank and select candidate architectures. Candidate architectures are generated by combining six redesign techniques: Global Average Pooling (GAP) to reduce classifier complexity, Quantization-Aware Training (QAT) to optimize weight precision, reduction in convolution filter size, selection of the downsampling strategy, power-of-two constraints on the feature map size before GAP, and training oriented toward integer-only inference. We evaluate the combined effect of these techniques on MNIST, FashionMNIST, CIFAR-10, and PneumoniaMNIST. The results show that, across MNIST and FashionMNIST, the selected compact models reduce the estimated HWCost by up to 2.52× with approximately one percentage point or less degradation in test accuracy. On CIFAR-10, the lowest-cost candidate reduces the estimated HWCost by 3.94×, relative to the highest-accuracy candidate, at the cost of a larger accuracy reduction. On PneumoniaMNIST, the efficiency-oriented candidate achieves comparable independent-test-balanced accuracy to the highest-validation-balanced accuracy candidate while requiring substantially lower estimated HWCost. Full article
(This article belongs to the Special Issue Deep Learning for Image, Video and Signal Processing, 2nd Edition)
►▼ Show Figures

Figure 1

18 pages, 1674 KB  
Article
A Social Media-Driven Public Participatory Emergency Decision-Making Method Considering Sentiment and Social Influence
by Qifeng Wan, Jing Han, Xiangyu Zhong and Xuanhua Xu
Information 2026, 17(9), 912; https://doi.org/10.3390/info17090912 - 17 Sep 2026
Viewed by 190
Abstract
In the data intelligence era, social media platforms have supplemented emergency decision-making with a wealth of timely data, offering new research paradigms for emergency response. It is crucial to identify and predict the emergency material demand for reducing secondary damage during emergencies. This [...] Read more.
In the data intelligence era, social media platforms have supplemented emergency decision-making with a wealth of timely data, offering new research paradigms for emergency response. It is crucial to identify and predict the emergency material demand for reducing secondary damage during emergencies. This study aims to fill the research gap in analysing the emergency material demand using real-time social media data. We propose a method for generating an emergency material demand index from social media that takes into account both sentiment intensity and social influence. Negative sentiment and social influence are integrated to generate review-level demand intensity and further aggregated into a dynamic MDI. Using mask demand during the early COVID-19 outbreak as a case study, 3,323,151 Weibo reviews were collected, with a 20% temporally stratified sample used for MDI generation. The domain-adapted RoBERTa achieved a negative F1-score of 0.79 and a Macro-F1 of 0.83, while sensitivity analysis confirmed the robustness of the MDI. External comparison suggests a potential lagged association with subsequent material distribution, and rolling-origin forecasting demonstrates its applicability to short-term demand forecasting. Our study provides a tool for dynamically monitoring and forecasting emergency materials demand to ensure sufficient time for the production or distribution of emergency materials. Full article
(This article belongs to the Special Issue Decision-Making Process in E-Commerce and Social Networks)
►▼ Show Figures

Figure 1

49 pages, 693 KB  
Review
A Review of Joint Unmanned Aerial Vehicle Trajectory and Camera Orientation Optimization
by Jakub Kůdela
Information 2026, 17(9), 911; https://doi.org/10.3390/info17090911 - 17 Sep 2026
Viewed by 260
Abstract
Camera-equipped Unmanned Aerial Vehicle (UAV) planning couples vehicle motion, camera pose, and scene-dependent sensing utility. The relevant literature is distributed across aerial reconstruction, inspection, target tracking, active perception, cinematography, and coverage planning, with substantial differences in vehicle models, camera mechanisms, visibility assumptions, and [...] Read more.
Camera-equipped Unmanned Aerial Vehicle (UAV) planning couples vehicle motion, camera pose, and scene-dependent sensing utility. The relevant literature is distributed across aerial reconstruction, inspection, target tracking, active perception, cinematography, and coverage planning, with substantial differences in vehicle models, camera mechanisms, visibility assumptions, and evaluation practice. A common vehicle–camera formulation is used here to compare two physical camera-realization mechanisms—independent gimbal actuation and body-coupled orientation—while treating viewpoint-first camera-pose planning as a separate representation that may defer physical realization. Mixed-integer formulations are first examined in detail because coverage, visibility, sequencing, assignment, and discrete camera modes introduce a logical structure; a representative time-expanded MILP is then provided for joint motion–view selection. Evolutionary methods are reviewed from early genetic and differential-evolution UAV planners through information-driven, constrained multiobjective, and hybrid formulations; and cooperative, surrogate-assisted, and transferability-aware methods from adjacent problem classes are then examined as possible extensions. In the frozen coded corpus, 13 direct studies use a physically independently actuated camera/gimbal, but none uses an evolutionary method as the primary optimizer for joint vehicle–gimbal motion; direct evolutionary joint vehicle–gimbal optimization therefore remains sparse. Simulation-to-reality transfer is analyzed as mismatch in dynamics, tracking, gimbal response, calibration, image formation, scene geometry, perception, and timing, with corresponding discussion of randomization, adaptive models, HIL evaluation, robust optimization, and transferability-aware search. A structured reproducibility audit through 31 August 2026 freezes the evidence base at 124 coded sources (62 direct studies, 48 adjacent precedents, and 14 proposed-transfer sources) and confirms persistent fragmentation in scenes, sensors, metrics, and computational budgets. The remaining technical questions concern mechanism-aware camera realization, scalable visibility, multi-view sensing utility under feedback, solver decomposition for mixed discrete–continuous problems, and transfer-sensitive evaluation. By synthesizing these methodological differences, this review identifies persistent research gaps and formulates recommendations for algorithm design, hybrid optimization, benchmarking, and sim-to-real validation. Full article
►▼ Show Figures

Figure 1

32 pages, 4985 KB  
Article
State-Triggered Adaptive Microgrid Dispatch for EV Charging: A Hybrid Deep Learning and Multi-Objective Optimization Framework
by Guanjian Zhu, Xin Ma, Yang Guo, Ke Wang, Jianyu Chen, Wei Li, Zujun Ding, Hui Huang, Xin Xia, Baolian Liu and Jie Ji
Information 2026, 17(9), 910; https://doi.org/10.3390/info17090910 - 17 Sep 2026
Viewed by 165
Abstract
The rapid growth of electric vehicle (EV) charging demand requires accurate short-term load forecasts and dispatch strategies that can respond to changing microgrid operating conditions. This study proposes a hybrid framework that combines a VMD-CNN-ABiLSTM-IGCRA forecasting model with a state-triggered adaptive scheduling strategy. [...] Read more.
The rapid growth of electric vehicle (EV) charging demand requires accurate short-term load forecasts and dispatch strategies that can respond to changing microgrid operating conditions. This study proposes a hybrid framework that combines a VMD-CNN-ABiLSTM-IGCRA forecasting model with a state-triggered adaptive scheduling strategy. Variational Mode Decomposition extracts multi-scale components from volatile charging profiles; the CNN and attention-based bidirectional LSTM learn local and long-range temporal features; and the Improved Giant Cane Rat Algorithm regulates the iterative search step and tunes the principal model hyperparameters. The scheduling layer constructs a five-dimensional state vector from electricity price, renewable-energy volatility, predicted EV load, grid carbon pressure, and available battery capacity. Historical thresholds activate four candidate operating cases: carbon-priority, price-driven, fluctuation-stabilized, and economic–environmental balanced scheduling. When several cases are triggered, the best feasible candidate is selected according to the current dispatch objective under common equipment and power-balance constraints. Three confidential real-world charging datasets from Shenzhen, each containing 744 hourly observations over 31 days, are evaluated using a chronological 70% training and 30% testing split. The proposed forecasting model obtains RMSE values of 1.7719–5.1042, MAE values of 0.8963–3.7562, and R2 values of 85.67–95.09% across the three datasets. In the station-level dispatch simulations, the balanced Case 4 reduces total cost relative to Case 1 by 16.95%, 54.53%, and 31.99% for Charging Stations 1–3, respectively. These results support the operational value of linking load forecasting with transparent state-dependent mode selection. Full article
►▼ Show Figures

Figure 1

26 pages, 2262 KB  
Article
From Prediction to Diagnostic Support: A Data-Driven System for Retail Demand Forecasting and Inventory Risk Assessment
by Gao Huan and Mohammad-Ali Sarvghadi
Information 2026, 17(9), 909; https://doi.org/10.3390/info17090909 - 17 Sep 2026
Viewed by 385
Abstract
The rapid digital transformation of the retail industry has generated large-scale, high-frequency data; yet many retailers still rely on siloed systems that decouple demand forecasting from operational inventory management. This separation often leads to structural inventory imbalances, including simultaneous stockouts and overstock situations. [...] Read more.
The rapid digital transformation of the retail industry has generated large-scale, high-frequency data; yet many retailers still rely on siloed systems that decouple demand forecasting from operational inventory management. This separation often leads to structural inventory imbalances, including simultaneous stockouts and overstock situations. To address this, we propose a unified, data-driven framework that integrates advanced sales forecasting with a diagnostic inventory health system, bridging predictive analytics with diagnostic decision support for proactive inventory-risk assessment. Utilizing a real-world dataset of approximately 500,000 product–store–day records from a Chinese e-commerce company, we evaluate six forecasting models across statistical, machine learning, and Deep Learning (DL) architectures. Results indicate that DL models achieved the lowest RMSE and MAPE values, whereas RF and XGBoost produced more favorable MASE values. Forecasting performance varied by evaluation criterion, with no model consistently outperforming all others. However, LSTM achieved the lowest testing RMSE (0.3996) and MAPE (21.63%). Building upon these predictive outputs, we introduce an inventory health diagnosis framework based on empirically calibrated thresholds for the Inventory Turnover Ratio (ITR) and Excess Inventory Rate (EIR), with the thresholds calibrated on the training and calibration period and evaluated on an independent temporal validation period. Application of this framework to the independent diagnostic validation period shows that 61.1% of SKU–store observations were classified as Potential Risk or Critical, including 16.3% classified as Critical, indicating inventory misalignment requiring managerial attention. By providing an interpretable and scalable system for proactive risk detection, this research provides a practical framework to support retailers in transitioning from demand forecasting to inventory diagnostic decision support. Full article
►▼ Show Figures

Figure 1

30 pages, 952 KB  
Review
Quantum Algorithms for Trading: A Survey of Speedups, Thresholds, and Dequantization
by Luigi Laura, Marco Parrillo, Alessio Pascucci, Marco Rossi and Valerio Rughetti
Information 2026, 17(9), 908; https://doi.org/10.3390/info17090908 - 17 Sep 2026
Viewed by 414
Abstract
Quantum algorithms for trading span computational tasks with different input models and standards of evidence. This structured critical review describes its search scope, selection criteria and limitations, and distinguishes reported findings from author assessments. Amplitude estimation estimates bounded expectations to additive error ϵ [...] Read more.
Quantum algorithms for trading span computational tasks with different input models and standards of evidence. This structured critical review describes its search scope, selection criteria and limitations, and distinguishes reported findings from author assessments. Amplitude estimation estimates bounded expectations to additive error ϵ with O(ϵ−1) oracle queries rather than the O(ϵ−2) samples of plain classical Monte Carlo at fixed confidence. This query advantage does not establish an end-to-end runtime advantage: state preparation, arithmetic, error correction and classical competitors must also be costed. Published resource estimates for benchmark exotics require thousands of logical qubits and demanding logical operation rates. An illustrative sensitivity analysis shows how the crossover depends on classical throughput, oracle depth and the comparison window; its numbers are scenarios, not calibrated hardware forecasts. Portfolio and trading-trajectory formulations have device demonstrations, but a 250-instance benchmark of discretized minimum-variance allocation finds classical mixed-integer programming and a tailored heuristic superior on the tested formulation. Specific low-rank quantum machine learning algorithms have been dequantized under analogous sampling-access assumptions. Separately, a controlled study of quantum kernels for Chinese equity returns finds no significant advantage and demonstrates sensitivity to evaluation design; this does not settle other learning tasks. Entanglement-assisted coordination offers advantages in specified nonlocal games without computational-complexity assumptions, while its financial implementation and economics remain open. We distinguish computational performance from economic value and identify input/output costs, structured algorithms, reproducible benchmarking and explicit financial evaluation as priorities. Full article
(This article belongs to the Special Issue Surveys in Information Systems and Applications)
►▼ Show Figures

Figure 1

30 pages, 1484 KB  
Article
Faster After a Weak Video: Platform Signals and Upload Timing on YouTube
by Tobias Ebbing, Paul C. Wolf and Jan Kratzer
Information 2026, 17(9), 907; https://doi.org/10.3390/info17090907 - 17 Sep 2026
Viewed by 254
Abstract
How do creators adapt when their work underperforms? In a within-channel panel of 9985 consecutive transitions from 57 technology channels, relative performance and upload timing are systematically related. A video that does worse than its channel’s recent average is followed by a shorter [...] Read more.
How do creators adapt when their work underperforms? In a within-channel panel of 9985 consecutive transitions from 57 technology channels, relative performance and upload timing are systematically related. A video that does worse than its channel’s recent average is followed by a shorter interval to the next release and by a semantically more distant release. The content response is specific to underperformance. Equally large positive deviations do not predict a comparable shift. The content results use the views-based signal alone. Neither association is detectable one transition further out, which is a descriptive horizon rather than a demonstrated stopping point. The timing association appears in three performance measures, but a shortfall and an equally large surplus cannot be told apart. The results survive rich past-only controls, flexible specifications, and a channel-specific forecast of publishing rhythm. Because performance was measured in a settled snapshot after both events, the data cannot separate a response to performance from reverse causality or a shared production state. The regularity and the identification template are both transferable. Full article
(This article belongs to the Section Information Applications)
►▼ Show Figures

Graphical abstract

27 pages, 21877 KB  
Article
A Multi-Scale Dual-Head YOLOv5 Framework for Hand Gesture Recognition via Spatial Relationship Modeling
by Yingying Zhao, Yizhi Wang, Bin Cai and Jinshui Miao
Information 2026, 17(9), 906; https://doi.org/10.3390/info17090906 - 16 Sep 2026
Viewed by 314
Abstract
Hand gesture recognition plays an important role in computer vision with broad applications in human–computer interaction, intelligent perception, and contactless interaction systems. However, existing detection-based methods mainly rely on hand appearance features, which are often insufficient for distinguishing gesture categories with similar local [...] Read more.
Hand gesture recognition plays an important role in computer vision with broad applications in human–computer interaction, intelligent perception, and contactless interaction systems. However, existing detection-based methods mainly rely on hand appearance features, which are often insufficient for distinguishing gesture categories with similar local hand shapes but different hand–head spatial configurations. To address this issue, a multi-scale dual-head hand gesture detection framework based on YOLOv5s is proposed. The framework introduces an auxiliary detection branch to explicitly detect hand and head regions, and the corresponding multi-scale feature representations are fused to incorporate hand–head spatial contextual information into gesture recognition. Furthermore, Spatial Pyramid Pooling (SPP) and a Convolutional Block Attention Module (CBAM) are employed to enhance multi-scale contextual representation and feature discrimination before final gesture prediction. To evaluate the proposed framework, experiments were conducted on a HaGRID-based Dataset, a UAV-Gesture Dataset, and a self-collected real-world zero-shot dataset. Experimental results demonstrate that the proposed framework consistently improves detection performance over the YOLOv5s baseline while maintaining relatively lightweight model complexity and real-time inference capability. In addition, ablation studies verify the effectiveness of the proposed dual-head architecture, multi-scale feature fusion, SPP, and CBAM, while cross-dataset and zero-shot evaluations further demonstrate the applicability of the proposed framework across different gesture datasets and unseen real-world scenarios. Full article
(This article belongs to the Section Artificial Intelligence)
►▼ Show Figures

Figure 1

35 pages, 657 KB  
Article
WinAPIReplay: Safely Re-Executing Win32 and NT-Native Malware API-Call Logs to Measure Behavioral Reproducibility and Generate Labeled Endpoint Telemetry
by Youji Fukuta, Yoshiaki Shiraishi, Masanori Hirotomo and Masami Mohri
Information 2026, 17(9), 905; https://doi.org/10.3390/info17090905 - 16 Sep 2026
Viewed by 222
Abstract
Behavioral malware analysis relies on dynamic-analysis logs—sequences of Windows API calls recorded by sandboxes such as CAPEv2—assumed to represent the malware’s effects faithfully. To the best of our knowledge, this assumption has never been tested by re-executing the recorded calls. We propose WinAPIReplay, [...] Read more.
Behavioral malware analysis relies on dynamic-analysis logs—sequences of Windows API calls recorded by sandboxes such as CAPEv2—assumed to represent the malware’s effects faithfully. To the best of our knowledge, this assumption has never been tested by re-executing the recorded calls. We propose WinAPIReplay, which re-executes each recorded Win32 and NT-native call as a real operating-system call, without the malware binary, using a unified handle map for cross-layer handle chains and a three-tier sandbox confining every side effect to disposable places. Across 500 WinMET samples from five families, 74.04% of modeled Layer-1 behavior-domain calls are reproduced (95% bootstrap CI [71.02, 76.75]); the Layer-1 domain covers 57.98% of all recorded calls, and a conservative rate excluding substituted calls is 71.15%. An ablation attributes this causally to the per-category executors—handle-validity falls from 94.9% to 10.2% without them—and no side effect escapes the sandbox on the channels the tool models and monitors. A four-way taxonomy assigns most of the residual to intrinsic, environment-dependent behavior; re-execution is near-deterministic (98.90% stable). Under Sysmon, it safely generates family-labeled file/registry telemetry for the reproduced subset (46,150 events) from static logs alone—while its result record recovers the malware’s process arguments and network destinations—and a classifier over it reaches 92.0% leave-one-out accuracy on 100 samples, statistically indistinguishable from an API-category baseline. WinAPIReplay thus provides, to the best of our knowledge, the first quantitative measurement of dynamic-log reproducibility within a demonstrated safety envelope, and a safe route to labeled, environment-consistent endpoint telemetry. Full article
(This article belongs to the Special Issue Information Security, Data Preservation and Digital Forensics)
►▼ Show Figures

Figure 1

17 pages, 1061 KB  
Article
Knowledge-Guided Multimodal Resource Identification in Low-Voltage Transformer Areas with Sparse Measurements
by Xiaoxing Lu, Xiaolong Xiao, Wenqiang Xie, Shuo Han, Jinyu Li, Chengjun Zhang and Wenbin Yu
Information 2026, 17(9), 904; https://doi.org/10.3390/info17090904 - 16 Sep 2026
Viewed by 218
Abstract
Identifying distributed photovoltaic generation, battery energy storage, and load-dominant behaviour remains challenging when transformer-area measurements are sparse, operating patterns overlap, and reference labels are incomplete. We develop a knowledge-guided multimodal learning framework that combines raw electrical sequences with temporal, statistical, frequency-domain, and contextual [...] Read more.
Identifying distributed photovoltaic generation, battery energy storage, and load-dominant behaviour remains challenging when transformer-area measurements are sparse, operating patterns overlap, and reference labels are incomplete. We develop a knowledge-guided multimodal learning framework that combines raw electrical sequences with temporal, statistical, frequency-domain, and contextual descriptors. A knowledge-guided label library reconciles archived operating records, expert rules, and clustering-based screening, while transfer learning and WGAN-based augmentation are used to improve learning under limited and imbalanced data. The proposed framework jointly encodes raw electrical sequences, engineered descriptors, and contextual information and integrates their complementary representations through attention-based multimodal fusion. On a held-out real-only test set with independently verified labels, the proposed framework achieved 93.8% accuracy and a macro-F1 of 93.8%, while retaining 90.4% accuracy when 30% of the inputs were randomly masked. These results support further evaluation of knowledge-guided multimodal learning for transformer-area resource identification, although broader cross-region and cross-utility validation is required before operational deployment. Full article
(This article belongs to the Special Issue Data Analytics and Machine Learning in Smart Energy Systems)
►▼ Show Figures

Figure 1

43 pages, 867 KB  
Article
Annotation-Budget Fairness Reliability: How Much Data Does a Trustworthy Fairness Audit of Hate-Speech Classifiers Require?
by Arjun Mukherjee, Thomas Mandl and Sukomal Pal
Information 2026, 17(9), 903; https://doi.org/10.3390/info17090903 - 15 Sep 2026
Viewed by 348
Abstract
Fairness audits of hate-speech classifiers are typically performed at a single annotation budget, leaving a basic question unanswered: how much labeled data does a trustworthy fairness audit require? We introduce ABFR (Annotation-Budget Fairness Reliability), a framework that sweeps the training budget and indexes [...] Read more.
Fairness audits of hate-speech classifiers are typically performed at a single annotation budget, leaving a basic question unanswered: how much labeled data does a trustworthy fairness audit require? We introduce ABFR (Annotation-Budget Fairness Reliability), a framework that sweeps the training budget and indexes two reliability thresholds—the Performance Reliability Threshold (PRT), the smallest budget at which model performance stabilizes across runs, and the Fairness Reliability Threshold (FRT), the smallest budget at which a probe-based fairness gap stabilizes. Across seven conditions—five from the HASOC hate-speech shared tasks spanning three languages (English, German, code-mixed Hindi), plus two independent English benchmarks, HateXplain and OLID—and a model ladder from classical TF-IDF logistic regression to a 2025-era multilingual encoder, comprising 34 condition × model cells in total, we find a consistent dissociation. Performance reliability resolves at modest budgets and improves with model capability, whereas fairness reliability does not resolve anywhere in the tested range, up to 90% of the available training pool, in nearly every condition and does not improve with capability. A near-deterministic linear model exhibits comparable fairness instability, showing the effect is not solely an artifact of transformer optimization. A controlled experiment holding the training subsample fixed and varying only the optimization seed shows that both data sampling and optimization stochasticity contribute materially. The single condition in which fairness reliability resolves is not explained by probe count as restricting other conditions to equally few groups does not recover reliability. Instead, it is explained by whether the specific audited groups exhibit stable per-group behavior under subsampling. We further show that probe-based worst-group findings are interpretable only where a model measurably fires on neutral probes, and we report exploratory results on caste, identifying a methodological obstacle: probe-neutrality assumptions do not transfer to identity categories, such as caste, whose mention is itself socially marked. Full article
(This article belongs to the Special Issue Natural Language Processing for Online Social Behavior)
►▼ Show Figures

Figure 1

23 pages, 1551 KB  
Article
White-Box Tweakable Block Cipher and Hardware-Binding Scheme Under Tweak-Unique Mode
by Jun Liu, Yanwei Zhou, Jie Chen, Xiaoli Dong, Feng Zhu and Bo Yang
Information 2026, 17(9), 902; https://doi.org/10.3390/info17090902 - 15 Sep 2026
Viewed by 288
Abstract
White-box cryptography protects secret keys in untrusted execution environments but remains vulnerable to code lifting attacks. Hardware-binding offers a practical countermeasure by coupling cryptographic implementations with specific devices, yet a unified theoretical foundation is lacking. This paper establishes such a foundation by formalizing [...] Read more.
White-box cryptography protects secret keys in untrusted execution environments but remains vulnerable to code lifting attacks. Hardware-binding offers a practical countermeasure by coupling cryptographic implementations with specific devices, yet a unified theoretical foundation is lacking. This paper establishes such a foundation by formalizing the white-box tweakable block cipher and introducing the tweak-binding property, which guarantees decryption failure upon tweak mismatch and underpins hardware-binding security. We prove that strong tweakable pseudo-random permutation security implies tweak-binding. A concrete construction is realized by embedding the WARX white-box block cipher into the CLRW14 framework, with provable security against key extraction. We further propose the tweak-unique mode as a mandatory operational paradigm, requiring a unique non-repeating tweak per encryption. Leveraging this mode with physical unclonable functions (PUFs), we design a hardware-binding scheme that stores only a single challenge-response pair per device, enabling deployment with commodity static random-access memory (SRAM) PUFs. Security proofs demonstrate provable resilience against code lifting. Performance evaluations on Intel server, Raspberry Pi 5, and Xilinx Zynq-7010 platforms show linear scaling with message size and minimal overhead, confirming suitability for both server-grade and resource-constrained internet of things (IoT) or mobile environments. Full article
(This article belongs to the Special Issue Cryptographic Protocols for Decentralized Security and Privacy)
►▼ Show Figures

Figure 1

Previous Issue
Next Issue
Back to TopTop