Editor’s Choice Articles

Editor’s Choice articles are based on recommendations by the scientific editors of MDPI journals from around the world. Editors select a small number of articles recently published in the journal that they believe will be particularly interesting to readers, or important in the respective research area. The aim is to provide a snapshot of some of the most exciting work published in the various research areas of the journal.

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
23 pages, 632 KB  
Review
AI-Driven Software Testing: A Review
by Guilherme Martins, Nelson Tenório and Jorge Bernardino
Big Data Cogn. Comput. 2026, 10(7), 233; https://doi.org/10.3390/bdcc10070233 - 10 Jul 2026
Cited by 2 | Viewed by 1719
Abstract
The rapid evolution of software complexity demands more efficient and autonomous testing mechanisms. Artificial intelligence (AI) has emerged as a solution to the limitations of traditional manual testing in software development, which is time-consuming, prone to human error, and unable to scale with [...] Read more.
The rapid evolution of software complexity demands more efficient and autonomous testing mechanisms. Artificial intelligence (AI) has emerged as a solution to the limitations of traditional manual testing in software development, which is time-consuming, prone to human error, and unable to scale with the increasing size and complexity of modern software systems. In this context, this paper presents an application-focused review of 35 selected empirical studies focusing on the use of AI during software testing, based on PRISMA guidelines. We introduce a comprehensive taxonomy categorizing current research into six core fields, including test case generation, defect prediction, and AI model verification. The analysis reveals that large language models, machine learning, and computer vision can significantly improve testing efficiency. Key findings demonstrate that AI can autonomously repair broken test scripts, generate robust synthetic data, enable codeless web testing, and accurately predict system defects before execution. Furthermore, advanced techniques such as reinforcement learning and deep learning successfully validate complex environments, including cloud robotics and quantum software. However, our qualitative and quantitative synthesis also highlights that challenges, such as generative AI “hallucinations” and the brittleness of Continuous Integration and Continuous Deployment (CI/CD) integration, persist. Ultimately, this review proposes a tailored research roadmap for robust industrial adoption, showing that AI is changing the way software is tested, shifting it from a predominantly reactive and static activity toward a proactive, intelligence-driven discipline. Full article
Show Figures

Figure 1

14 pages, 1635 KB  
Article
Temporal Patterns of Advanced Shot-Quality Metrics in Elite Men’s and Women’s European Football
by Blanca De-la-Cruz-Torres, Anselmo Ruiz-de-Alarcón-Quintero and Miguel Navarro-Castro
Big Data Cogn. Comput. 2026, 10(7), 214; https://doi.org/10.3390/bdcc10070214 - 1 Jul 2026
Viewed by 986
Abstract
The temporal dynamics of shot quality in elite football remain poorly understood, despite well-documented declines in physical and technical performance during matches. This study aimed to analyze the evolution of advanced shot metrics across match halves and 15 min intervals in elite men’s [...] Read more.
The temporal dynamics of shot quality in elite football remain poorly understood, despite well-documented declines in physical and technical performance during matches. This study aimed to analyze the evolution of advanced shot metrics across match halves and 15 min intervals in elite men’s and women’s international competitions. A total of 4074 shots from the UEFA European Championships were examined. To ensure methodological consistency among the three advanced shooting metrics (expected goals, xG; expected shot impact timing, xSIT; and expected goals on target, xGOT), analyses were restricted to shots on target (men: 775; women: 554), as xGOT can only be calculated for on-target attempts. Shot quality was assessed using xG, xSIT and xGOT. Differences between halves were evaluated using the Mann–Whitney U test, while temporal trends were analyzed through linear mixed-effects models. Results showed no significant differences between halves in shot distribution or quality metrics in either competition (all p > 0.05). Likewise, no significant temporal variations were found across the six match intervals for any metric. Women’s football exhibited a largely stable quality of shots on target throughout the match. In men’s competitions, although a significant difference in xG was observed (p = 0.04, effect sizes were trivial (d = −0.17), and no consistent patterns emerged in xSIT or xGOT. No sex-related differences were observed. Overall, shot quality remained stable despite match progression, suggesting that changes in goal frequency are not driven by variations in the intrinsic quality of shooting opportunities. Therefore, these findings suggest that optimizing the contextual conditions that facilitate the creation of shooting opportunities may be as important as, or more important than, focusing solely on shooting execution. Full article
(This article belongs to the Special Issue AI and Data Science in Sports Analytics)
Show Figures

Figure 1

32 pages, 10561 KB  
Article
Bio-Inspired Spiking Recurrent Networks with Evolutionary Optimization for Non-Stationary Cryptocurrency Forecasting
by Francis Noah Walugembe, Maciej Wielgosz, Matej Mertik and Matjaž Gams
Big Data Cogn. Comput. 2026, 10(7), 200; https://doi.org/10.3390/bdcc10070200 - 23 Jun 2026
Viewed by 814
Abstract
Forecasting cryptocurrency prices remains difficult because market dynamics are highly volatile, non-stationary, and regime-dependent. This study investigates whether combining a spiking-inspired recurrent architecture with the Grey Wolf Optimizer (GWO) can improve one-step-ahead Bitcoin forecasting within a controlled model family. We compare four configurations, [...] Read more.
Forecasting cryptocurrency prices remains difficult because market dynamics are highly volatile, non-stationary, and regime-dependent. This study investigates whether combining a spiking-inspired recurrent architecture with the Grey Wolf Optimizer (GWO) can improve one-step-ahead Bitcoin forecasting within a controlled model family. We compare four configurations, LSTM, SLSTM, GWO-LSTM, and GWO-SLSTM, on 4039 daily BTC–USD closing prices from 17 September 2014 to 9 October 2025 using Min–Max normalization, strict chronological splitting, windowed regime-based robustness analysis across three distinct market regimes, and repeated-run testing. The proposed SLSTM replaces the conventional hidden-state recurrence with leaky integrate-and-fire-inspired synaptic, membrane, and adaptive-threshold dynamics, functioning as a spiking-inspired recurrent model with thresholded event gating (reset = `none’, learnable threshold). On the primary hold-out split, GWO-SLSTM achieved a test RMSE of 1840.97 and a test MAPE of 1.76%, compared with 2217.24 and 2.46% for GWO-LSTM, 3501.48 and 3.86% for SLSTM, and 4030.10 and 4.40% for LSTM. Both GWO-optimized models exhibited substantial improvements over their non-optimized counterparts, while the SLSTM baseline also outperformed the plain LSTM, indicating gains from both spiking-inspired recurrence and evolutionary hyperparameter optimization. Both optimized models exhibited near-zero bias (PBIAS 0.11% for GWO-LSTM and 0.36% for GWO-SLSTM). Within the present implementation, GWO-SLSTM also trained faster than GWO-LSTM (39.71 s vs. 137.28 s), although this runtime difference should be interpreted as setup-specific because the model families were implemented in different frameworks and stopped after different numbers of epochs. Overall, within the expanded univariate BTC–USD setting, the results support GWO-SLSTM as a strong within-family candidate for one-step-ahead forecasting under non-stationary conditions. Full article
(This article belongs to the Special Issue Financial Time Series Analysis and Forecasting in the Big Data Era)
Show Figures

Figure 1

40 pages, 1511 KB  
Article
Quantum Hyperbolic Deep Learning for Foreign-Exchange Trading: A Hybrid Reinforcement-Learning Pipeline over Attractor-Aware Magnet-Price Manifolds
by Francesco Rundo
Big Data Cogn. Comput. 2026, 10(6), 191; https://doi.org/10.3390/bdcc10060191 - 11 Jun 2026
Viewed by 1305
Abstract
Foreign-exchange decisions rest on hierarchically organized evidence whose latent structure is inadequately captured by Euclidean representations. Reinforcement-learning agents trained on flat embeddings inherit stability guarantees that do not transfer to the manifold supporting the latent state. We address both limitations through a hybrid [...] Read more.
Foreign-exchange decisions rest on hierarchically organized evidence whose latent structure is inadequately captured by Euclidean representations. Reinforcement-learning agents trained on flat embeddings inherit stability guarantees that do not transfer to the manifold supporting the latent state. We address both limitations through a hybrid architecture in which a schema-constrained structured chain-of-thought is embedded into a Poincaré ball, transported to a qubit register via angle encoding, and processed by an L-layer hardware-efficient variational ansatz on a state-vector backend. The circuit exposes two read-outs to the policy, namely, a scalar Pauli-Z observable and a projected quantum kernel inducing a fidelity-based similarity over magnet-price attractors, the latter identified via kernel-weighted recurrence density and finite-time Lyapunov statistics. The Lipschitz constraint on the action-value function is lifted from the hyperbolic geodesic distance to a joint metric on Bκn×P(H). A stability theorem yields an explicit bound depending on the read-out operator norm, on the depth–width product of the ansatz, and on the curvature–Hilbert balance. The pipeline is evaluated on nine major FX crosses over a 2015–2025 out-of-sample window, with rolling-origin walk-forward retraining and broker-published transaction costs. The system attains 2.55% pair-averaged non-compounded monthly P&L and 8.83% maximum drawdown, with Sharpe 1.78, Calmar 3.43, and Probabilistic Sharpe Ratio exceeding 0.95 on every cross. The gain remains significant under a deflated-Sharpe-ratio test with Ntrials=42 correction. Block-wise ablations exhibit strictly monotone degradation: removing the projected kernel costs 4.15 p.p. on annualized P&L, the joint Lipschitz penalty 6.42 p.p., the attractor module 7.64 p.p., and the hyperbolic embedding 8.40 p.p. The quantum block thereby instantiates a structurally non-classical, geometry-aware regularizer identifiable through ablation rather than asymptotically advantageous. Full article
Show Figures

Figure 1

18 pages, 4742 KB  
Article
The Container Market in Baltic Ports: Market Share Development and Trend Forecasting
by Diana Šateikienė and Jurga Kučinskienė
Big Data Cogn. Comput. 2026, 10(6), 187; https://doi.org/10.3390/bdcc10060187 - 6 Jun 2026
Viewed by 709
Abstract
This study examines the evolution of container throughput and competitive market-share dynamics in the three principal Baltic container ports—Klaipeda, Riga, and Tallinn—during the period 2005–2024 and provides baseline forecasts to 2030. The proposed analytical framework combines descriptive statistical analysis, normalized market-share assessment, growth-rate [...] Read more.
This study examines the evolution of container throughput and competitive market-share dynamics in the three principal Baltic container ports—Klaipeda, Riga, and Tallinn—during the period 2005–2024 and provides baseline forecasts to 2030. The proposed analytical framework combines descriptive statistical analysis, normalized market-share assessment, growth-rate analysis, and ordinary least-squares trend estimation with prediction intervals to distinguish aggregate market fluctuations from port-specific competitive realignments. The results indicate increasing market concentration in Klaipeda, a gradual decline in Riga’s relative position, and long-term stagnation and volatility in Tallinn. Common regional shocks are observed during the 2009 global financial crisis and the COVID-19 disruption in 2020, while atypical positive deviations in Klaipeda suggest competitive redistribution effects associated with changes in regional logistics flows and shipping-network configurations. Forecast results indicate continued medium-term growth in Klaipeda and Riga, whereas Tallinn demonstrates weaker trend stability and greater forecast uncertainty. The study contributes a transparent and reproducible baseline decision-support framework that can be implemented using routinely available throughput statistics for medium-term infrastructure assessment and capacity evaluation, infrastructure prioritisation, and risk monitoring. The findings also highlight the limitations of deterministic linear forecasting in volatile port systems and support future integration with higher-frequency operational data and machine-learning forecasting approaches. Full article
(This article belongs to the Topic Data Intelligence and Computational Analytics)
Show Figures

Figure 1

40 pages, 2487 KB  
Systematic Review
Human-Centered AI for Decision Support Systems: A Systematic Review of Application Domains, Architecture Designs, Current Trends and Future Directions
by Marco Fanfani, Luciano Alessandro Ipsaro Palesi and Paolo Nesi
Big Data Cogn. Comput. 2026, 10(6), 186; https://doi.org/10.3390/bdcc10060186 - 5 Jun 2026
Cited by 5 | Viewed by 2642
Abstract
Artificial Intelligence is increasingly used to support decision-making across many domains. However, concerns related to transparency, reliability and human oversight indicate the need for improved human-centered AI (HCAI) approaches in decision support systems (DSSs). In this paper, a systematic review was conducted in [...] Read more.
Artificial Intelligence is increasingly used to support decision-making across many domains. However, concerns related to transparency, reliability and human oversight indicate the need for improved human-centered AI (HCAI) approaches in decision support systems (DSSs). In this paper, a systematic review was conducted in accordance with PRISMA 2020 in the Web of Science database: ninety research articles published between 2015 and 2025 were analyzed to investigate how HCAI is applied within DSSs in multiple application domains. HCAI + DSS research outcomes were analyzed and explored, first identifying the main architectural designs and discussing the involved components integrating human interaction, generative AI models, data and knowledge management, decision logic, and orchestration mechanisms, then focusing on specific domains and highlighting impact achieved, technologies used, and validation strategies employed. In addition, alignment with United Nations’ Sustainable Development Goals (SDG) was considered, and the temporal evolution of the most relevant topics was studied to identify more interesting trends and less investigated areas. Finally, findings were summarized, current limitations were discussed, and future research directions for helping researchers and practitioners in developing more reliable, explainable, and human-aware decision support systems were outlined. Full article
Show Figures

Figure 1

24 pages, 5714 KB  
Article
Functional Assessment of Neonatal Hypoxic–Ischemic Encephalopathy Using Long-Duration EEG and Interpretable Deep Learning Models
by Athira Chandran, Lekshmi Chandrika Reghunath, Claudio Tomazzoli, Christian Napoli and Cristian Randieri
Big Data Cogn. Comput. 2026, 10(6), 175; https://doi.org/10.3390/bdcc10060175 - 1 Jun 2026
Viewed by 945
Abstract
Neonatal hypoxic–ischemic encephalopathy (HIE) remains a critical neurological emergency resulting from perinatal asphyxia, often leading to lifelong neurodevelopmental disabilities or mortality. The accurate and timely grading of HIE severity is paramount for initiating therapeutic interventions such as therapeutic hypothermia. This work proposes a [...] Read more.
Neonatal hypoxic–ischemic encephalopathy (HIE) remains a critical neurological emergency resulting from perinatal asphyxia, often leading to lifelong neurodevelopmental disabilities or mortality. The accurate and timely grading of HIE severity is paramount for initiating therapeutic interventions such as therapeutic hypothermia. This work proposes a diagnostic framework that uses long-duration electroencephalogram (EEG) recordings through a hierarchical classification strategy and advanced sequence modeling. A Hybrid Mamba-inspired architecture was developed to effectively capture long-range temporal dependencies in multi-channel neonatal EEG while maintaining computational efficiency. In order to enhance clinical consistency and initialize the models appropriately, a Self-Supervised Learning step based on Masked Signal Modeling is implemented with a mask ratio of 30%. The model structure takes into consideration clinically verified biomarkers, including the suppression ratio, Delta–Alpha Ratio, Spectral Edge Frequency, and Rhythmicity Index, extracted from signals at a microvolt level prior to normalization for physiological interpretability purposes. These features are combined with waveforms using feature gating. In an experiment conducted on a dataset of 169 records using 5-fold subject-wise cross-validation, the designed Hybrid Mamba-based model achieves significant stability and generalizability, achieving an accuracy score of 90%, with an average accuracy of 88.45% ± 6.8% per hierarchical level. Full article
Show Figures

Figure 1

26 pages, 786 KB  
Article
Output Correction of Recurrence-Aware Long-Term Cognitive Network Classifiers
by Gonzalo Nápoles, Isel Grau and Yamisleydi Salgueiro
Big Data Cogn. Comput. 2026, 10(6), 178; https://doi.org/10.3390/bdcc10060178 - 1 Jun 2026
Viewed by 676
Abstract
Recurrence-Aware Long-Term Cognitive Network (rLTCN) classifiers have reported comparable performance to mainstream black-box models, including tree ensembles and support vector machines, in tabular pattern classification tasks. These classifiers use a two-step learning algorithm to address issues that arise during the training of recurrent [...] Read more.
Recurrence-Aware Long-Term Cognitive Network (rLTCN) classifiers have reported comparable performance to mainstream black-box models, including tree ensembles and support vector machines, in tabular pattern classification tasks. These classifiers use a two-step learning algorithm to address issues that arise during the training of recurrent neural networks. While the weights in the recurrent block are computed using unsupervised learning, recurrence-aware weights are determined using a one-step learning rule based on the Moore-Penrose inverse. However, the related least-squares learning problem tends to favor easy instances and common patterns, particularly those associated with the majority class in imbalanced datasets. In such scenarios, a loss function that directly optimizes a robust metric, such as the F1 score, would lead to models with stronger generalization capabilities. Unfortunately, incorporating such a metric into the Moore-Penrose inverse learning procedure presents challenges from a mathematical viewpoint. In this paper, we propose four gradient-based correction methods that modify the output logits of rLTCN classifiers once the two-step training process is done. Inspired by procedures such as Platt or Beta scaling, the proposed post-optimization correction methods seek to maximize the F1 score rather than produce calibrated probabilities. The simulations using real-world datasets show that adding a correction layer to rLTCNs improves their performance significantly at the expense of occasional reductions in the precision metric. Full article
(This article belongs to the Section Cognitive System)
Show Figures

Figure 1

45 pages, 7500 KB  
Article
SemNet Explorer: An Evidence-Grounded Knowledge Graph–LLM Framework for Multi-Scale Mechanistic Reporting Across Biomedical Domains
by Xin He, David Camacho, Lama Moukheiber, Meghna Iyer, Benjamin Zhao, Christophe Ye, Batuhan Nursal, Xinyu Guo, Albert J. B. Lee and Cassie S. Mitchell
Big Data Cogn. Comput. 2026, 10(6), 171; https://doi.org/10.3390/bdcc10060171 - 25 May 2026
Viewed by 1049
Abstract
Background: Mechanistic reporting from large-scale biomedical knowledge graphs remains challenging, particularly when integrating structured graph evidence with large language model (LLM)–based explanation in a reproducible and auditable manner. Existing approaches either rely on manual synthesis of graph-derived results or generate unconstrained narratives that [...] Read more.
Background: Mechanistic reporting from large-scale biomedical knowledge graphs remains challenging, particularly when integrating structured graph evidence with large language model (LLM)–based explanation in a reproducible and auditable manner. Existing approaches either rely on manual synthesis of graph-derived results or generate unconstrained narratives that lack traceability to underlying evidence. Methods: We present SemNet Explorer, an evidence-grounded knowledge graph–LLM unified framework for automated mechanistic reporting across biomedical domains using SemNet 2.0, a PubMed-scale heterogeneous knowledge graph. Given a set of target concepts and a selected semantic layer, the framework organizes graph-derived evidence into structured regions and generates two complementary report types: global reports for process-level mechanisms and anchor-centric reports for localized mediator-based explanations. A central methodological contribution is an ablation-derived adaptive grounding policy: we systematically compare alternative evidence-integration strategies across report types, semantic layers, and region structures, and use the resulting preferences to guide prompt selection in the deployed system. Results: SemNet Explorer produces stable region decompositions and interpretable report scaffolds across molecular (AAPP), disease-level (DSYN), and pharmacologic (PHSU) representations. For global reports, explicit evidence grounding improves expression quality more consistently than content accuracy, with benefits dependent on evidence density and semantic abstraction. In contrast, anchor-centric reports show consistent improvements in both content and expression under stronger, mediator-constrained prompting. These findings are supported by both pairwise ablation comparisons and absolute score analyses. Conclusions: SemNet Explorer establishes a generalizable unified framework and interactive platform for transforming knowledge graph evidence into reproducible mechanistic narratives across biomedical domains, including multimorbidity analysis, comparative pathophysiology, drug repurposing, and adverse event discovery. The results demonstrate that effective knowledge graph–LLM integration requires adaptive, context-dependent evidence grounding rather than fixed prompting strategies. Full article
Show Figures

Figure 1

23 pages, 677 KB  
Article
Large Language Models for Energy Market Analytics: An Exploratory Feasibility Study Across Geopolitical Monitoring, Commodity Summarisation, and Renewable Forecasting
by Alex Krempasky, Erik Kajati and Peter Papcun
Big Data Cogn. Comput. 2026, 10(6), 166; https://doi.org/10.3390/bdcc10060166 - 22 May 2026
Cited by 1 | Viewed by 1059
Abstract
Large Language Models (LLMs) offer opportunities for processing heterogeneous information streams relevant to energy-market decision-making, but their practical role in forecasting-oriented analytical workflows remains uncertain. This paper presents an exploratory feasibility study of LLM use across four energy-market tasks: geopolitical event monitoring for [...] Read more.
Large Language Models (LLMs) offer opportunities for processing heterogeneous information streams relevant to energy-market decision-making, but their practical role in forecasting-oriented analytical workflows remains uncertain. This paper presents an exploratory feasibility study of LLM use across four energy-market tasks: geopolitical event monitoring for Dutch Title Transfer Facility (TTF) market context using Global Database of Events, Language, and Tone (GDELT)-based data, structured summarisation of commodity-intelligence articles, prompt-engineered solar-power and grid-load forecasting for Austria, and a short-horizon exploratory TTF price-estimation case. The study is positioned as a pilot investigation and hybrid workflow blueprint rather than as a statistically conclusive forecasting benchmark. A four-layer reference architecture was devised, including structured market data, semi-structured news intelligence, web-scraping concepts, and implemented Twitter/X and GDELT monitoring layers. The empirical cases indicate that LLMs are most useful for text-heavy reasoning, event-context integration, source triage, and structured interpretation. In the 20-article summarisation corpus, Gemini 1.5 Pro achieved higher commodity-direction accuracy than GPT-4, while GPT-4 showed stronger output-format stability. In selected solar case checks, OpenAI models produced plausible generation curves close to the Fraunhofer ISE Energy Charts reference, while Energy Charts remained more accurate for aggregate load estimation in the available benchmark comparison. The two-day TTF experiment illustrated that LLMs can incorporate qualitative geopolitical context into short-horizon reasoning, but it did not establish reliable price-forecasting capability. The Twitter/X monitoring layer is retained as a documented negative pathway, showing the limitations of informal social-media scraping for reproducible market intelligence. Full article
(This article belongs to the Special Issue Large Language Models and Their Limitations)
Show Figures

Figure 1

23 pages, 4279 KB  
Article
Impact of Server-Side Aggregation on Federated Traffic Classification Under Heterogeneous Data Distributions
by Salam Allawi Hussein and Sándor R. Répás
Big Data Cogn. Comput. 2026, 10(6), 167; https://doi.org/10.3390/bdcc10060167 - 22 May 2026
Viewed by 2101
Abstract
The growing prevalence of encrypted network traffic has rendered traditional payload-based inspection ineffective, shifting attention toward flow-level statistical analysis combined with machine learning. At the same time, privacy regulations and distributed network architectures make centralised data collection increasingly impractical, motivating federated learning as [...] Read more.
The growing prevalence of encrypted network traffic has rendered traditional payload-based inspection ineffective, shifting attention toward flow-level statistical analysis combined with machine learning. At the same time, privacy regulations and distributed network architectures make centralised data collection increasingly impractical, motivating federated learning as a privacy-preserving alternative. Despite its promise, deploying federated learning for encrypted traffic classification in realistic environments remains challenging, particularly under heterogeneous client data distributions that arise when different network sites observe different subsets of services. This paper examines how server-side aggregation affects federated QUIC traffic classification under such heterogeneous conditions. We use a five-class Google QUIC dataset and represent each flow with eight statistical features derived from packet size and timing. We compare a centralised baseline with federated learning under three client partitions: mixed-label clients (C1), service-based single-class clients (C2), and hash-based semi-IID clients (C3). For each case, we evaluate four Flower aggregation strategies: FedAvg, FedAdam, FedAvgM, and FedYogi. Results show that client distribution has a greater impact on performance than the choice of aggregation strategy. Federated models match or closely approach centralised performance in C1 and C3, with accuracy up to 0.9969 and macro-AUC near 1.0. In C2, accuracy drops due to extreme label skew, but adaptive aggregation mitigates the effect. FedYogi achieves the best C2 accuracy of 0.9287, while FedAvgM attains the highest C2 macro-AUC of 0.9885. ROC curves and confusion matrices confirm that the choice of aggregation matters mainly under severe heterogeneity. Full article
Show Figures

Figure 1

26 pages, 396 KB  
Article
Blockchains for Data Management: The DIGI4ECO Use Case and Practical Lessons Beyond Theory
by Andreas Polyvios Delladetsimas, Elias Iosif, Stamatis Papangelou and George Giaglis
Big Data Cogn. Comput. 2026, 10(5), 162; https://doi.org/10.3390/bdcc10050162 - 18 May 2026
Cited by 1 | Viewed by 854
Abstract
This article examines blockchain as an enabling technological component for data management tasks that are independent of currency-related functionality, a less-discussed aspect of a technology commonly associated with cryptocurrencies and decentralized finance (DeFi). Drawing on empirical findings from the DIGI4ECO project as a [...] Read more.
This article examines blockchain as an enabling technological component for data management tasks that are independent of currency-related functionality, a less-discussed aspect of a technology commonly associated with cryptocurrencies and decentralized finance (DeFi). Drawing on empirical findings from the DIGI4ECO project as a case study, we present a structured literature review and cross-domain analysis of blockchain-based data management systems (BDMSs), examine a representative permissioned BDMS implementation, and synthesize practical design guidelines and implementation insights for BDMS development. This perspective is motivated by core blockchain properties such as immutability and transparency, as well as by the observation that existing resources for BDMS development, including methods, tools, and best practices, remain fragmented and less developed than those available for more mature technologies. Full article
(This article belongs to the Section Big Data)
Show Figures

Figure 1

20 pages, 3389 KB  
Article
Teaching AI to Decode Vaccine Hesitancy Narratives: A Few-Shot Learning and Topic Modeling Approach
by Md Enamul Kabir, Shakhawat H. Tanim, Deanna D. Sellnow, Geneva Lei P. Luteria and Lior Rennert
Big Data Cogn. Comput. 2026, 10(5), 159; https://doi.org/10.3390/bdcc10050159 - 16 May 2026
Cited by 1 | Viewed by 839
Abstract
Vaccine hesitancy—which can be defined as a delay in acceptance or the refusal to get vaccinated—has substantially increased over the past decade. This study introduces a computational and qualitative approach designed to efficiently classify stance and uncover narratives in social media discourse without [...] Read more.
Vaccine hesitancy—which can be defined as a delay in acceptance or the refusal to get vaccinated—has substantially increased over the past decade. This study introduces a computational and qualitative approach designed to efficiently classify stance and uncover narratives in social media discourse without relying on extensive manual annotation. Using 298,356 COVID-19 vaccine-related X posts geolocated to South Carolina (June 2021–May 2022), zero-shot and few-shot learning with instruction-tuned large language models (Mistral-7B, Meta-Llama-3.1, and DeepSeek-7B) was applied for stance detection while Latent Dirichlet Allocation (LDA) was used for topic modeling. The topic modeling identified five dominant themes in vaccine hesitant conversations: skepticism of vaccine efficacy, comparative framing, scientific justification, disapproval of regulations, and distrust. Temporal analysis revealed that skepticism peaked during late 2021, coinciding with booster campaigns and mandate debates. These findings suggest that vaccine hesitancy is influenced through complex rhetorical strategies rather than misinformation alone. These underlying narratives often frame skepticism as rational and evidence-based, using scientific language and statistical reasoning to challenge the effectiveness of vaccines. Full article
Show Figures

Figure 1

17 pages, 1159 KB  
Article
Performance Trade-Offs of Optimizing Small Language Models for E-Commerce
by Josip Tomo Licardo, Nikola Tanković, Ivan Osman, Ivan Lorencin and Sandi Baressi Šegota
Big Data Cogn. Comput. 2026, 10(5), 155; https://doi.org/10.3390/bdcc10050155 - 14 May 2026
Cited by 1 | Viewed by 1232
Abstract
Large Language Models (LLMs) offer state-of-the-art performance in natural language understanding and generation tasks. However, the deployment of leading commercial models for specialized tasks, such as e-commerce, is often hindered by high computational costs, latency, and operational expenses. This paper investigates the viability [...] Read more.
Large Language Models (LLMs) offer state-of-the-art performance in natural language understanding and generation tasks. However, the deployment of leading commercial models for specialized tasks, such as e-commerce, is often hindered by high computational costs, latency, and operational expenses. This paper investigates the viability of smaller, open-weight models as a resource-efficient alternative. We present a methodology for optimizing a one-billion-parameter Llama 3.2 model for multilingual e-commerce intent recognition. The model was fine-tuned using Quantized Low-Rank Adaptation (QLoRA) on a synthetically generated dataset designed to mimic real-world user queries. Subsequently, we applied post-training quantization techniques, creating GPU-optimized (GPTQ) and CPU-optimized (GGUF) versions. Our results demonstrate that the specialized 1B model achieves 98.8% accuracy, approaching the performance of the significantly larger GPT-4.1 model. A detailed performance analysis revealed critical, hardware-dependent trade-offs: while 4-bit GPTQ reduced VRAM usage by 41%, it paradoxically slowed inference by 82% on an older GPU architecture (NVIDIA T4) due to dequantization overhead. Conversely, GGUF formats on a CPU achieved a speedup of up to 4.3× in inference throughput and up to a 72% reduction in RAM consumption compared to the FP16 baseline. We conclude that small, properly optimized open-weight models are not just a viable but a more suitable alternative for domain-specific applications, offering state-of-the-art accuracy at a fraction of the computational cost. Full article
Show Figures

Figure 1

28 pages, 3206 KB  
Article
Consensus-Driven Framework for Data-Driven Optimization of Distributed Systems Through Blockchain Consensus Mechanism Selection
by Miljenko Švarcmajer, Mirko Kohler, Zdravko Krpić and Ivica Lukić
Big Data Cogn. Comput. 2026, 10(5), 154; https://doi.org/10.3390/bdcc10050154 - 13 May 2026
Cited by 1 | Viewed by 947
Abstract
Modern data-driven distributed systems increasingly rely on blockchain technologies to ensure trust, transparency, and decentralized coordination. However, the rapid proliferation of consensus mechanisms has created a complex design space, making the selection of an appropriate protocol a non-trivial architectural and decision-making challenge. Different [...] Read more.
Modern data-driven distributed systems increasingly rely on blockchain technologies to ensure trust, transparency, and decentralized coordination. However, the rapid proliferation of consensus mechanisms has created a complex design space, making the selection of an appropriate protocol a non-trivial architectural and decision-making challenge. Different consensus mechanisms rely on distinct security resources, validator admission models, and agreement architectures, leading to diverse trade-offs between scalability, decentralization, performance, and governance. Existing studies primarily focus on classification or performance comparison of consensus mechanisms, while the problem of systematic, requirement-driven selection remains insufficiently addressed. In particular, there is a lack of structured approaches that integrate multiple system requirements into a unified decision framework suitable for real-world environments. To address this gap, this paper proposes a consensus-driven, layered framework for blockchain consensus mechanism selection, formulated as a multi-criteria decision problem. The framework organizes the consensus design space across key architectural dimensions and analyzes 32 consensus mechanisms, enabling systematic comparison and supporting data-driven decision-making. The approach is further demonstrated through five representative use-case scenarios, showing its applicability in optimizing distributed system design. Full article
(This article belongs to the Special Issue AI and Blockchain for Trustworthy Social Computing)
Show Figures

Figure 1

24 pages, 1787 KB  
Article
Data-Driven Peak Demand Identification in Commercial Electricity Consumption for Load Curve Flattening
by Michał Gostkowski, Tomasz Ząbkowski and Krzysztof Gajowniczek
Big Data Cogn. Comput. 2026, 10(5), 152; https://doi.org/10.3390/bdcc10050152 - 12 May 2026
Viewed by 839
Abstract
Effective peak load management enables utilities to mitigate increased electricity demand and optimize the use of available resources during periods of maximum consumption. Accurate forecasting of the peak load is essential for ensuring the reliability, efficiency, and resilience of contemporary power systems. In [...] Read more.
Effective peak load management enables utilities to mitigate increased electricity demand and optimize the use of available resources during periods of maximum consumption. Accurate forecasting of the peak load is essential for ensuring the reliability, efficiency, and resilience of contemporary power systems. In this study, commercial customer-level data were employed to identify electricity peak demand within the Polish power system, drawing upon historical records of both energy consumption and meteorological variables. Departing from conventional time series forecasting approaches, the problem was intentionally reformulated as a pattern recognition task. Three classification techniques were systematically evaluated to identify individual customers’ peak load events, thereby offering a basis for demand-side management strategies and incentive mechanisms aimed at flattening load profiles and improving grid stability. The proposed approach demonstrates how data-driven analytics can support utilities in extracting actionable knowledge from large-scale energy datasets and enabling proactive demand response programs. Empirical results indicate that the proposed methods are capable of predicting up to 90% of electricity peak occurrences, with a forecasting horizon of 24 h leading to significant shifts in the load curve. Full article
(This article belongs to the Section Data Mining and Machine Learning)
Show Figures

Figure 1

16 pages, 284 KB  
Review
Talent Identification and AI-Driven Decision Tools in Sport: A Policy-Oriented Perspective on Algorithmic Bias, Data Privacy, and Digital Determinism in Player Evaluation
by Elia Morgulev and Ofer H. Azar
Big Data Cogn. Comput. 2026, 10(5), 146; https://doi.org/10.3390/bdcc10050146 - 7 May 2026
Cited by 1 | Viewed by 2795
Abstract
Big-data analytics are increasingly used in scouting and talent identification, with machine learning (ML) tools applied to evaluate and predict player performance based on match statistics, video tracking, physical and anthropometric tests, psychological assessments, social media data, and qualitative scouting reports. Advances in [...] Read more.
Big-data analytics are increasingly used in scouting and talent identification, with machine learning (ML) tools applied to evaluate and predict player performance based on match statistics, video tracking, physical and anthropometric tests, psychological assessments, social media data, and qualitative scouting reports. Advances in computer vision, together with the emergence of affordable automated broadcasting and data collection systems, have extended the deployment of ML-driven scouting from professional to youth sport. The use of algorithms in educational, employment, and healthcare settings has been shown to introduce biases and discrimination while wrongly assuming accuracy and objectivity because the decisions are made automatically and quantitatively. In this respect, we briefly describe the development of data-driven performance analysis and how ML-based technologies are currently applied for early screening and comparison of large player populations. Based on a narrative overview of the literature, we draw on evidence from education, employment, and healthcare to identify risks that may also emerge in ML-driven player evaluation, including algorithmic bias, non-representative training data, privacy concerns, and the persistence of model-based labels over time, especially in youth sport. Our main contribution is translating these threats into governance principles and operational safeguards for responsible use of AI in scouting and talent identification. Full article
(This article belongs to the Special Issue AI and Data Science in Sports Analytics)
22 pages, 5847 KB  
Article
BERT-Based Models for Normalization of Adverse Drug Event Expressions in Social Media to Standard Medical Terminology for Drug Safety Analysis
by Fan Dong, Wenjing Guo, Jie Liu, Ann Varghese, Weida Tong, Tucker A. Patterson and Huixiao Hong
Big Data Cogn. Comput. 2026, 10(5), 141; https://doi.org/10.3390/bdcc10050141 - 2 May 2026
Viewed by 811
Abstract
Social media platforms host abundant and timely descriptions of medication experiences that can complement traditional pharmacovigilance systems. Yet the linguistic informality of these data presents a major challenge for mapping adverse drug event (ADE) expressions to standardized medical terminology. In this study, we [...] Read more.
Social media platforms host abundant and timely descriptions of medication experiences that can complement traditional pharmacovigilance systems. Yet the linguistic informality of these data presents a major challenge for mapping adverse drug event (ADE) expressions to standardized medical terminology. In this study, we developed BERT-based language models to classify ADE mentions from social media into MedDRA System Organ Classes (SOCs). Using the SMM4H and CADEC corpora, as well as their combination, we performed 20 iterations of 20% holdout validation for 3-, 6-, 22-, and 25-SOC classification tasks with a selected fixed training configuration (learning rate, batch size, and training epochs) based on training-loss convergence. The models achieved accuracies ranging from 75% to 94%, demonstrating strong performance for SOC-level classification of noisy and informal ADE expressions under the evaluated settings. These results are based on a controlled mention-level evaluation using deduplicated adverse drug event strings and do not establish document-level or real-world deployment generalization. This work provides a systematic evaluation of BERT-based models for SOC-level classification of ADEs and demonstrates consistent performance within the evaluated datasets and label granularities. While direct comparison with prior studies is limited by differences in datasets and evaluation protocols, the results demonstrate that transformer-based models can effectively classify ADEs into SOCs. These findings support the use of transformer-based normalization for SOC-level aggregation of user-reported adverse events and their integration into large-scale social media pharmacovigilance pipelines as a downstream component under controlled conditions. Full article
(This article belongs to the Section Data Mining and Machine Learning)
Show Figures

Figure 1

20 pages, 17426 KB  
Article
Towards Improved Clinical Adoption of AI Segmentation Models: Benchmarking High-Performance Models for Resource-Constrained Settings
by Emmanuel Chibuikem Nnadozie, Susana Merino-Caviedes, Daniel A. de Luis-Román, Marcos Martín-Fernández and Carlos Alberola-López
Big Data Cogn. Comput. 2026, 10(5), 142; https://doi.org/10.3390/bdcc10050142 - 2 May 2026
Viewed by 845
Abstract
High-performance medical segmentation models are often benchmarked on high-end GPUs. Such benchmarks do not provide useful performance insights for point-of-care low-end devices. This work, firstly, posits that to achieve improved clinical adoption of AI-powered segmentation models, especially in reduced manpower settings like rural [...] Read more.
High-performance medical segmentation models are often benchmarked on high-end GPUs. Such benchmarks do not provide useful performance insights for point-of-care low-end devices. This work, firstly, posits that to achieve improved clinical adoption of AI-powered segmentation models, especially in reduced manpower settings like rural hospitals, we need benchmarks that provide actionable insights on the degree to which high-performance models address five deployment constraints viz: resource-effectiveness for low-end computing devices, clinically acceptable accuracy, clinically compatible execution times, localization of user data, and user-based finetuning. In this work, five state-of-the-art foundation segmentation models and one target-specific model were systematically evaluated on three multi-organ medical datasets. Furthermore, the best-ranking foundation model and target-specific model were benchmarked on three low-end devices. Our findings show that lightweight foundation models provided the best performance trade-off and are easily user-fine-tuned on custom datasets. Target-specific models provide high accuracy out-of-the-box, but may require significant optimisation to deliver comparably fast execution times and user-based finetuning on low-end devices. The methods and results from this research provide actionable insights on high-performance medical segmentation models for low-end computing devices, as a necessary step towards improved adoption in resource-limited clinical settings. Full article
Show Figures

Figure 1

19 pages, 496 KB  
Article
Evaluating Computational Approaches for Harmful Content Analysis: Promise, Pitfalls and Tools for Responsible Research
by Itai Himelboim and Mudit Baid
Big Data Cogn. Comput. 2026, 10(5), 143; https://doi.org/10.3390/bdcc10050143 - 2 May 2026
Viewed by 877
Abstract
This manuscript develops and demonstrates a practical framework for evaluating automated classifiers used in communication research, using harmful language detection as an illustrative case. We combine (a) a structured review of documentation practices for 27 publicly available classifiers and their associated annotation processes [...] Read more.
This manuscript develops and demonstrates a practical framework for evaluating automated classifiers used in communication research, using harmful language detection as an illustrative case. We combine (a) a structured review of documentation practices for 27 publicly available classifiers and their associated annotation processes with (b) a cross-dataset evaluation that re-tests each model beyond its original training context. Across 27 datasets, we extract and compare reporting on construct definitions, annotator instructions, and inter-annotator agreement, and we quantify generalization by applying each model to multiple out-of-domain test sets. We also benchmark a contemporary large language model (GPT-5) under a consistent prompting protocol to illustrate how LLM-based classification compares to fine-tuned classifiers. Results show that documentation is uneven and often insufficient for theory-driven measurement, inter-annotator agreement varies widely across datasets, and cross-dataset performance frequently drops substantially relative to within-dataset evaluations. Building on these findings and existing validation guidance, we provide a reusable checklist and decision flow to help researchers select, justify, and report classifier-based measures in ways that support transparency and cumulative science. Recommendations for researchers, reviewers, and journal editors stress aligning model selection with standards of validity, reliability, and transparency. Full article
(This article belongs to the Section Data Mining and Machine Learning)
Show Figures

Figure 1

29 pages, 1779 KB  
Article
BWT-Enhanced Compression for GIS Raster Data: A Hybrid AV1-Inspired Approach with Burrows–Wheeler Transform
by Yair Wiseman
Big Data Cogn. Comput. 2026, 10(5), 140; https://doi.org/10.3390/bdcc10050140 - 1 May 2026
Viewed by 959
Abstract
The AVIF (AV1 Image File Format) is a modern, royalty-free image format that leverages the AV1 video codec for superior compression efficiency, supporting both lossy and lossless modes. Its entropy encoding relies on a multi-symbol context-adaptive arithmetic coder (range coding with adaptive cumulative [...] Read more.
The AVIF (AV1 Image File Format) is a modern, royalty-free image format that leverages the AV1 video codec for superior compression efficiency, supporting both lossy and lossless modes. Its entropy encoding relies on a multi-symbol context-adaptive arithmetic coder (range coding with adaptive cumulative distribution functions (CDFs)), which is effective for general imagery but may not optimally exploit the repetitive structures common in Geographic Information System (GIS) maps/data. This paper proposes replacing AVIF’s entropy encoder with the Burrows–Wheeler Transform (BWT), a reversible preprocessing algorithm that rearranges data to create runs of similar symbols, enhancing subsequent compression. We detail the technical steps for modification, drawing from AV1’s open-source implementation, and explain why BWT is advantageous for GIS raster maps/data, which often feature large uniform areas, limited color palettes, and spatial redundancies. Empirical evidence from related studies on BWT-based image compression shows improvements in lossless scenarios, potentially considerably reducing file sizes over standard methods while preserving data integrity critical for geospatial analysis. This swap could improve storage, transmission, and processing efficiency in GIS applications, such as remote sensing and cartography. The discussion includes challenges like computational overhead and compatibility, with recommendations for implementations. The resulting BWT-AVIF hybrid produces a non-standard AV1 bit-stream that is not compliant with the AV1 or AVIF specifications and therefore requires custom decoders. It is presented here as a research prototype for GIS-specific compression rather than a compliant AVIF extension. Full article
Show Figures

Figure 1

30 pages, 1706 KB  
Article
Understanding the Global Trends of 2025 Through the Defly Compass Methodology
by Mabel López Bordao, Antonia Ferrer Sapena, Carlos A. Reyes Pérez and Enrique A. Sánchez Pérez
Big Data Cogn. Comput. 2026, 10(4), 124; https://doi.org/10.3390/bdcc10040124 - 17 Apr 2026
Viewed by 4503
Abstract
This study aims to identify and synthesize the major global trends that shaped 2025 by applying the DeflyCompass methodology to a curated corpus of strategic foresight reports. The study synthesizes insights from 23 strategic reports published by leading international organizations, including the World [...] Read more.
This study aims to identify and synthesize the major global trends that shaped 2025 by applying the DeflyCompass methodology to a curated corpus of strategic foresight reports. The study synthesizes insights from 23 strategic reports published by leading international organizations, including the World Economic Forum, Accenture, Euromonitor, and major technology firms. Methodologically, DeflyCompass operationalizes a structured hybrid human–AI pipeline comprising the deployment of multi-agent AI systems, automated knowledge graph construction, semantic clustering, and hybrid human–AI validation processes, reducing an initial set of 816 preliminary signals to a validated catalog of 50 high-priority trends across six PESTEL domains: Political, Economic, Social, Technological, Environmental, and Legal/Governance. Key findings indicate that artificial intelligence functions as a systemic enabling technology across all domains, climate and sustainability imperatives permeate multiple domains, geopolitical fragmentation introduces systemic tension, and trust deficits emerge as a critical vulnerability. The study contributes a replicable and scalable framework for global-level strategic foresight that operationalizes human–AI integration within a rigorous expert-driven validation process, complementing existing hybrid analytical approaches in the literature. Implications extend to decision-making in technology governance, sustainability strategy, social adaptation, and scenario planning, highlighting the necessity of integrating AI augmentation with human expertise for effective future-oriented planning. Full article
Show Figures

Graphical abstract

33 pages, 3137 KB  
Article
Distilling the Complexity of Agent-Based Simulations into Textual Explanations via Large Language Models
by Noé Y. Flandre and Philippe J. Giabbanelli
Big Data Cogn. Comput. 2026, 10(4), 121; https://doi.org/10.3390/bdcc10040121 - 15 Apr 2026
Viewed by 1490
Abstract
Communicating the design and results of agent-based models (ABMs) to subject matter experts is challenging, which hinders participation and limits trust in simulation-based decision support. Large language models (LLMs) can communicate ABMs as textual summaries, thus complementing traditional disclosure through statistical and visualization [...] Read more.
Communicating the design and results of agent-based models (ABMs) to subject matter experts is challenging, which hinders participation and limits trust in simulation-based decision support. Large language models (LLMs) can communicate ABMs as textual summaries, thus complementing traditional disclosure through statistical and visualization techniques. While prior work translated the structure of conceptual models into narratives via LLMs, our extension covers the dynamics of simulation models via an automated simulation-to-text method that extracts contextual information from NetLogo ABMs, performs repeated simulations, and generates narrative descriptions (including the model’s purpose, parameters, and simulation dynamics) using mutimodal LLMs. Furthermore, four summarization algorithms spanning abstractive and extractive methods provide shorter reports. Using Design-of-Experiments methods over three peer-reviewed ABMs, state-of-the-art multimodal LLMs from 2026 (Gemini 3.1 Pro, Qwen 3.5, Kimi K2.5, Claude Opus 4.6) and different prompt elements (e.g., roles, examples, generating insights, statistical analyses), we compare our results with several reference reports (e.g., from associate professors). We find that report quality is determined mainly (i.e., up to 34% of the variance) by the summarization algorithm and its interaction with the LLM, with abstractive summarizers (BART, T5) producing more coherent and readable reports, while Claude Opus 4.6 is the most robust LLM. Full article
(This article belongs to the Section Large Language Models and Embodied Intelligence)
Show Figures

Figure 1

18 pages, 555 KB  
Article
Enhancing Retrieval-Augmented Generation with Entity Linking for Educational Platforms
by Francesco Granata, Francesco Poggi and Misael Mongiovì
Big Data Cogn. Comput. 2026, 10(4), 120; https://doi.org/10.3390/bdcc10040120 - 13 Apr 2026
Cited by 1 | Viewed by 1445
Abstract
In the era of Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) architectures are gaining significant attention for their ability to ground language generation in reliable knowledge sources. Despite their effectiveness, RAG systems based solely on semantic similarity often fail to ensure factual accuracy [...] Read more.
In the era of Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) architectures are gaining significant attention for their ability to ground language generation in reliable knowledge sources. Despite their effectiveness, RAG systems based solely on semantic similarity often fail to ensure factual accuracy in specialized domains, where terminological ambiguity can affect retrieval relevance. This study proposes Entity Linking Enhanced RAG (ELERAG), an enhanced RAG architecture that integrates a factual signal derived from Entity Linking to improve the accuracy of educational question-answering systems in Italian. The system includes a Wikidata-based Entity Linking module and implements a hybrid re-ranking strategy based on Reciprocal Rank Fusion (RRF). To validate our approach, we compared it against standard baselines and state-of-the-art methods, including a Weighted-Score Re-ranking, a standalone Cross-Encoder and a combined RRF + Cross-Encoder pipeline. Experiments were conducted on two benchmarks: a custom academic dataset and the standard SQuAD-it dataset. Results show that, in domain-specific contexts, ELERAG significantly outperforms both the baseline and the Cross-Encoder configurations. Conversely, the Cross-Encoder approaches achieve the best results on the general-domain dataset. These findings provide strong experimental evidence of the domain mismatch effect, highlighting the importance of domain-adapted hybrid strategies to enhance factual precision in educational RAG systems without relying on computationally expensive models trained on disparate data distributions. They also demonstrate the potential of entity-aware RAG systems in educational environments, fostering adaptive and reliable AI-based tutoring tools. Full article
(This article belongs to the Section Large Language Models and Embodied Intelligence)
Show Figures

Figure 1

29 pages, 6592 KB  
Article
Non-Invasive Sleep Stage Classification with Imbalance-Aware Machine Learning for Healthcare Monitoring
by Luisiana Sabbatini, Alberto Belli, Sara Bruschi, Marco Esposito, Sara Raggiunto and Paola Pierleoni
Big Data Cogn. Comput. 2026, 10(4), 116; https://doi.org/10.3390/bdcc10040116 - 10 Apr 2026
Cited by 4 | Viewed by 1166
Abstract
Sleep plays a fundamental role in human health and cognitive functioning, motivating the development of reliable and scalable methodologies for sleep stage classification (SSC). Recent advances in non-invasive and economically sustainable sensing technologies enable continuous sleep monitoring beyond laboratory settings. However, SSC remains [...] Read more.
Sleep plays a fundamental role in human health and cognitive functioning, motivating the development of reliable and scalable methodologies for sleep stage classification (SSC). Recent advances in non-invasive and economically sustainable sensing technologies enable continuous sleep monitoring beyond laboratory settings. However, SSC remains a challenging data analytics task due to the intrinsic class imbalance among sleep stages. This study investigates the effectiveness of different imbalanced data management strategies within a machine learning framework for non-invasive SSC. The proposed approach relies exclusively on heart rate and motion signals, which can be acquired through wearable devices or contactless under-mattress sensors, making it suitable for longitudinal monitoring scenarios. Using the PhysioNet DREAMT dataset, 32 experimental scenarios are defined by combining data-level techniques (ADASYN oversampling with different balancing weights), algorithm-level strategies (cost-sensitive learning), and hybrid solutions. Four model families are evaluated—Decision Tree, k-Nearest Neighbors, Ensemble Classifiers, and Artificial Neural Networks—across classification tasks involving 2, 3, 4, and 5 sleep stages. The experimental results show that ensemble-based models provide robust and consistent performance under severe class imbalance, achieving macro accuracies of 82% for sleep–wake detection, 73% for 3-stage classification, 72% for 4-stage classification, and 64% for 5-stage classification. These findings confirm the relevance of imbalance-aware analytics and demonstrate the feasibility of accurate, minimally invasive SSC within big data and cognitive computing paradigms. Full article
Show Figures

Figure 1

37 pages, 1919 KB  
Article
LLMs for Integrated Business Intelligence: A Big Data-Driven Framework Integrating Marketing Optimization, Financial Performance, and Audit Quality
by Leonidas Theodorakopoulos, Aristeidis Karras, Alexandra Theodoropoulou and Christos Klavdianos
Big Data Cogn. Comput. 2026, 10(4), 110; https://doi.org/10.3390/bdcc10040110 - 5 Apr 2026
Cited by 3 | Viewed by 2565
Abstract
Enterprise decision making in marketing, finance, and audit remains fragmented, leading to inefficient budget allocation and incomplete risk assessment. This study proposes an integrated, Big Data-driven decision-support framework that unifies Large Language Models (LLMs), attention-based marketing mix modeling, and multi-agent, game-theoretic optimization to [...] Read more.
Enterprise decision making in marketing, finance, and audit remains fragmented, leading to inefficient budget allocation and incomplete risk assessment. This study proposes an integrated, Big Data-driven decision-support framework that unifies Large Language Models (LLMs), attention-based marketing mix modeling, and multi-agent, game-theoretic optimization to coordinate cross-functional decisions. The architecture combines five modules: LLM-enhanced customer segmentation and customer lifetime value prediction, attention-weighted marketing mix modeling, multi-agent LLM systems for hierarchical budget optimization, attention-informed Markov multi-touch attribution, and LLM-augmented audit quality assessment. Empirical validation on a large-scale e-commerce dataset with 2.8 million customers and USD 156 million in marketing expenditure shows that marketing return on investment increases from 4.2 to 6.78 (61.4% relative improvement), financial forecasting error (MAPE) decreases from 12.8% to 4.7% (63.3% reduction), fraud detection accuracy improves by 29.8%, the Audit Quality Index reaches 0.951, and customer lifetime value prediction accuracy improves from 76.4% to 91.3%. By operationalizing the convergence of LLMs, attention mechanisms, and game-theoretic reasoning within a unified and empirically validated framework, the study delivers both theoretical advances and practically deployable tools for integrated business intelligence in digital economies. Full article
(This article belongs to the Section Large Language Models and Embodied Intelligence)
Show Figures

Figure 1

16 pages, 589 KB  
Article
Exploring the Mechanisms Influencing Graduate Students’ Adoption of Generative AI: Insights from the Technology Acceptance Model
by Qing Chen, Yujie Xue, Jie Lin and Chang Zhu
Big Data Cogn. Comput. 2026, 10(4), 108; https://doi.org/10.3390/bdcc10040108 - 3 Apr 2026
Viewed by 1920
Abstract
The rapid development of Generative Artificial Intelligence (GenAI) in graduate education has changed human–AI interaction within knowledge-intensive environments, leading to important questions about user-side cognitive adaptation in probabilistic AI systems. While many studies focus on ethical implications, limited attention has been paid to [...] Read more.
The rapid development of Generative Artificial Intelligence (GenAI) in graduate education has changed human–AI interaction within knowledge-intensive environments, leading to important questions about user-side cognitive adaptation in probabilistic AI systems. While many studies focus on ethical implications, limited attention has been paid to the cognitive mechanisms underlying graduate students’ adoption of GenAI. Drawing on the Technology Acceptance Model (TAM), this study explores the cognitive and interactional mechanisms shaping graduate students’ adoption and usage of GenAI. Using thematic analysis of in-depth interviews with 20 graduate students from diverse academic backgrounds, the study identifies seven interrelated constructs: perceived usefulness, perceived ease of use, external environment, risk perception, attitude, behavioral intention, and interaction subjectivity. This study demonstrates that the adoption of GenAI is not merely a result of perceived efficiency but is shaped by cognitive calibration between trust and risk evaluation. Moreover, interaction subjectivity emerges as a metacognitive factor that determines whether engagement results in human–AI collaboration or passive automation. By integrating external environment, risk perception, and interaction subjectivity, this study provides a cognitively grounded framework for understanding human–AI adoption and interaction dynamics. Practically, the findings provide design-relevant insights for developing GenAI systems that support calibrated trust, uncertainty awareness, and adaptive cognitive participation. Full article
Show Figures

Figure 1

35 pages, 4221 KB  
Article
Semantic Agent-Based Intelligent Digital Twins Integrating Demand, Production and Product Through Asset Administration Shells
by Joel Lehmann, Tim Markus Häußermann and Julian Reichwald
Big Data Cogn. Comput. 2026, 10(4), 103; https://doi.org/10.3390/bdcc10040103 - 26 Mar 2026
Cited by 1 | Viewed by 1676
Abstract
Complex products and production processes are intertwined and demand expressive, lifecycle-wide digital representations. The Asset Administration Shell emerged as a standard for Digital Twins (DTs), structuring heterogeneous data across cloud-based Industrial Internet of Things (IIoT) infrastructures. However, today’s deployments predominantly realize passive or [...] Read more.
Complex products and production processes are intertwined and demand expressive, lifecycle-wide digital representations. The Asset Administration Shell emerged as a standard for Digital Twins (DTs), structuring heterogeneous data across cloud-based Industrial Internet of Things (IIoT) infrastructures. However, today’s deployments predominantly realize passive or reactive DTs, while intelligent behavior remains underexploited. This paper addresses this gap, proposing an end-to-end architecture operationalizing the DT Reference Model through the integration of machine-interpretable granulated industrial skills, which are semantically accumulated into a knowledge graph enabling discovery and reasoning, while a multi-agent system provides autonomous, utility-based negotiation via machine-to-machine interactions within a federated marketplace. The approach is applied in a real smart manufacturing demonstrator, combining order processes, production orchestration, and lifecycle documentation into a unified execution pipeline spanning IIoT-connected shopfloor assets and cloud-based services. Quantitative experiments evaluating negotiation latency, renegotiation robustness, and utility variation demonstrate stable, predictable behavior even under concurrent demand and failure scenarios. The architecture lays a foundation for interoperable, sovereign collaboration across value chains to realize shared production. The results underline the effectiveness of the tightly coupled enabler technologies realizing proactive, reconfigurable, and semantically enriched intelligent DTs. Full article
Show Figures

Figure 1

28 pages, 502 KB  
Article
Emotional Framing in Prompts Modulates Large Language Model Performance
by Manuel Gozzi and Francesca Fallucchi
Big Data Cogn. Comput. 2026, 10(4), 102; https://doi.org/10.3390/bdcc10040102 - 24 Mar 2026
Viewed by 4126
Abstract
Large Language Models (LLMs) demonstrate remarkable performance across a variety of natural language understanding tasks, yet their sensitivity to emotional framing in user prompts remains underexplored. This paper presents an empirical study investigating how four emotional tones—joy, apathy, anger, and fear—affect LLM performance [...] Read more.
Large Language Models (LLMs) demonstrate remarkable performance across a variety of natural language understanding tasks, yet their sensitivity to emotional framing in user prompts remains underexplored. This paper presents an empirical study investigating how four emotional tones—joy, apathy, anger, and fear—affect LLM performance on the SuperGLUE benchmark. We evaluate five instruction-tuned, open-weight models across eight diverse tasks, systematically modulating input prompts with affective cues while keeping semantic content constant. Results reveal that prompts framed with joy and apathy lead to consistently higher accuracy, with gains of up to 4.5 percentage points compared to fear-framed inputs, which yield the lowest performance. These findings demonstrate that affective modulation in user prompts measurably impacts LLM reasoning and task outcomes, suggesting that emotional framing is not merely stylistic but functionally relevant to model behavior. Our study provides a reproducible experimental framework and an open-source prompt set, offering a foundation for future research on affect-aware prompting strategies and their implications in human–AI interaction. Full article
Show Figures

Figure 1

46 pages, 2822 KB  
Review
Generative AI and the Foundation Model Era: A Comprehensive Review
by Abdussalam Elhanashi, Siham Essahraui, Pierpaolo Dini, Davide Paolini, Qinghe Zheng and Sergio Saponara
Big Data Cogn. Comput. 2026, 10(3), 94; https://doi.org/10.3390/bdcc10030094 - 20 Mar 2026
Cited by 5 | Viewed by 9332
Abstract
Generative artificial intelligence and foundation models have changed machine learning by allowing systems to produce readable text, realistic images, and other multimodal content with little direct input from a user. Foundation models are large neural networks trained on very large and varied datasets, [...] Read more.
Generative artificial intelligence and foundation models have changed machine learning by allowing systems to produce readable text, realistic images, and other multimodal content with little direct input from a user. Foundation models are large neural networks trained on very large and varied datasets, and they form the core of many current generative AI (GenAI) systems. Their rapid development has led to major advances in areas like natural language processing, computer vision, multimodal learning, and robotics. Examples include GPT, LLaMA, and diffusion-based architectures, such as models often used for image generation. Systems such as Stable Diffusion show this shift by illustrating how AI can interpret information, draw basic inferences, and produce new outputs using more than one type of data. This review surveys common foundation model architectures and examines what they can do in generative tasks. It reviews Transformer, diffusion, and multimodal architectures, focusing on methods that support scaling and transfer across domains. The paper also reviews key approaches to pretraining and fine-tuning, including self-supervised learning, instruction tuning, and parameter-efficient adaptation, which support these systems’ ability to generalize across tasks. In addition to the technical details, this review discusses how GenAI is being used for text generation, image synthesis, robotics, and biomedical research. The study also notes continuing challenges, such as the high computing and energy demands of large models, ethical concerns about data bias and misinformation, and worries about privacy, reliability, and responsible use of AI in real settings. This review brings together ideas about model design, training methods, and social implications to point future research toward GenAI systems that are efficient, easy to interpret, and reliable, while supporting scientific progress and ethical responsibility. Full article
(This article belongs to the Special Issue Multimodal Deep Learning and Its Applications)
Show Figures

Figure 1

23 pages, 5079 KB  
Article
Dual-Stream Transformer with Kalman-Based Sensor Fusion for Wearable Fall Detection
by Abheek Pradhan, Sana Alamgeer, Rakesh Suvvari, Syed Tousiful Haque and Anne H. H. Ngu
Big Data Cogn. Comput. 2026, 10(3), 90; https://doi.org/10.3390/bdcc10030090 - 17 Mar 2026
Viewed by 2216
Abstract
Wearable fall detection systems face a fundamental challenge: while gyroscope data provide valuable orientation cues, naively combining raw gyroscope and accelerometer signals can degrade performance due to noise contamination. To overcome this challenge, we present a dual-stream transformer architecture that incorporates (i) Kalman-based [...] Read more.
Wearable fall detection systems face a fundamental challenge: while gyroscope data provide valuable orientation cues, naively combining raw gyroscope and accelerometer signals can degrade performance due to noise contamination. To overcome this challenge, we present a dual-stream transformer architecture that incorporates (i) Kalman-based sensor fusion to convert noisy gyroscope angular velocities into stable orientation estimates (roll, pitch, yaw), maintaining an internal state of body pose, and (ii) processing accelerometer and orientation streams in separate encoder pathways before fusion to prevent cross-modal interference. Our architecture further integrates Squeeze-and-Excitation channel attention and Temporal Attention Pooling to focus on fall-critical temporal patterns. Evaluated on the SmartFallMM dataset using 21-fold leave-one-subject-out cross-validation, the dual-stream Kalman transformer achieves 91.10% F1, outperforming single-stream Kalman transformers (89.80% F1) by 1.30% and single-stream baseline transformers (88.96% F1) by 2.14%. We further evaluate the model in real time using a watch-based SmartFall App on five participants, maintaining an average F1 score of 83% and an accuracy of 90%. These results indicate robust performance in both offline and real-world deployment settings, establishing a new state-of-the-art for inertial-measurement-unit-based fall detection on commodity smartwatch devices. Full article
Show Figures

Figure 1

34 pages, 5422 KB  
Article
Home-Based Telerehabilitation Through a Modular, Sensor-Integrated Virtual Monitoring System
by Zoltán Mészáros, M. A. Hannan Bin Azhar, Tasmina Islam and Soumya Kanti Manna
Big Data Cogn. Comput. 2026, 10(3), 84; https://doi.org/10.3390/bdcc10030084 - 8 Mar 2026
Viewed by 1522
Abstract
Home based telerehabilitation has expanded after COVID-19, but delivering timely guidance and monitoring exercise performance outside the clinic remains difficult. Traditional physiotherapy often relies on repeated execution of simple routines, yet clinicians have limited visibility into adherence and movement quality during unsupervised sessions. [...] Read more.
Home based telerehabilitation has expanded after COVID-19, but delivering timely guidance and monitoring exercise performance outside the clinic remains difficult. Traditional physiotherapy often relies on repeated execution of simple routines, yet clinicians have limited visibility into adherence and movement quality during unsupervised sessions. From a systems perspective, many telerehabilitation approaches also face constraints in accessibility, bandwidth, and computational cost that can limit practical deployment. This paper presents a modular telerehabilitation framework and prototype that captures and records rehabilitation exercise sessions for asynchronous clinician review in a 3D visualisation environment. The system integrates skeletal motion capture with plantar pressure sensing, and stores sessions as portable artefacts to support replay, inspection, and downstream analysis. A connector-based architecture enables extension to additional sensors without redesigning the core application, and the design aims to support deployment under constrained home computing and networking conditions. The manuscript contributes an implementation blueprint and reference architecture for multimodal capture and replay. Clinical effectiveness, usability outcomes, and quantitative sensor accuracy benchmarking are outside the scope of this work and are identified as necessary future evaluation. Full article
Show Figures

Figure 1

44 pages, 1099 KB  
Systematic Review
Sound Event Detection in Smart Cities: A Systematic Review of Methods, Datasets, and Applications
by Giuseppe Ciaburro and Virginia Puyana-Romero
Big Data Cogn. Comput. 2026, 10(3), 83; https://doi.org/10.3390/bdcc10030083 - 8 Mar 2026
Cited by 4 | Viewed by 2798
Abstract
Sound Event Detection (SED) is a growing area with vast prospects for understanding and designing the sonic fabric of smart cities. In this paper, the latest advances in SED are summarized, focusing on models, datasets, and applications from scientific papers listed on Scopus [...] Read more.
Sound Event Detection (SED) is a growing area with vast prospects for understanding and designing the sonic fabric of smart cities. In this paper, the latest advances in SED are summarized, focusing on models, datasets, and applications from scientific papers listed on Scopus and Web of Science. The paper provides a clear view of how SED is being used in smart cities, public safety, environment monitoring, and home security. The paper also addresses the challenges of SED, including dataset representativeness, model robustness under noisy or complex acoustic scenes, event rarity detection, as well as the ethics of using automatic listening. The paper also provides a view of future work to be undertaken in SED. The focus of the paper is on self-supervised learning, multi-modal fusion, neuro-inspired approaches, as well as privacy-preserving analytics. The paper provides a view of SED as a key technology to make smart cities safe, secure, and sustainable. SED has vast prospects as a key technology to enable artificial perception of smart cities. Full article
Show Figures

Figure 1

24 pages, 504 KB  
Article
Feasibility Study of CUDA-Accelerated Homomorphic Encryption and Benchmarking on Consumer-Grade and Embedded GPUs
by Volodymyr Dubetskyy and Maria-Dolores Cano
Big Data Cogn. Comput. 2026, 10(3), 79; https://doi.org/10.3390/bdcc10030079 - 6 Mar 2026
Cited by 1 | Viewed by 2265
Abstract
Fully Homomorphic Encryption (FHE) provides strong data confidentiality during computation but often suffers from high latency on Central Processing Units (CPUs). This study evaluates Graphics Processing Unit (GPU) acceleration for modern FHE libraries across a laptop (NVIDIA GTX 1650 Ti), a server (NVIDIA [...] Read more.
Fully Homomorphic Encryption (FHE) provides strong data confidentiality during computation but often suffers from high latency on Central Processing Units (CPUs). This study evaluates Graphics Processing Unit (GPU) acceleration for modern FHE libraries across a laptop (NVIDIA GTX 1650 Ti), a server (NVIDIA RTX 4060), and a Jetson Nano 2 GB embedded GPU. We benchmark key generation, arithmetic operations, Boolean-gate evaluation and scheme-specific tasks such as relinearization and key switching, using library-provided benchmarks with an explicit baseline (operation scope, timing boundaries, and parameter tuples). Moreover, we compare GPU-native libraries (NuFHE, Phantom-FHE, and Troy-Nova) with CPU-oriented ones (Microsoft SEAL, HElib, OpenFHE, Cupcake, and TFHE-rs). Results show GPUs deliver significant speedups for targeted operations. For example, NuFHE’s NVIDIA CUDA (Compute Unified Device Architecture) backend achieves about 1.4× faster Boolean-gate evaluation on the laptop and 3.4× faster on the server compared to its OpenCL backend. Likewise, RLWE (Ring Learning With Errors)-based schemes (BFV, CKKS, and BGV) see marked gains for polynomial arithmetic such as Number Theoretic Transform (NTT) when executed via Phantom-FHE. However, attempts to add CUDA support to Microsoft SEAL reveal four main challenges: high-precision modular arithmetic on GPUs, sequential dependencies in SEAL’s design, limited GPU memory and complex build-system changes. In light of these findings, we propose revised guidelines for GPU-first FHE libraries and practical recommendations for deploying high-throughput, privacy-preserving solutions on modern GPUs. Full article
(This article belongs to the Section Big Data)
Show Figures

Figure 1

24 pages, 3755 KB  
Article
Automating Data Product Discovery with Large Language Models and Metadata Reasoning
by Michalis Pingos, Artemis Photiou and Andreas S. Andreou
Big Data Cogn. Comput. 2026, 10(3), 72; https://doi.org/10.3390/bdcc10030072 - 28 Feb 2026
Viewed by 2855
Abstract
The exponential growth of data over the past decade has created new challenges in transforming raw information into actionable knowledge, particularly through the development of data products. The latter is essentially the result of querying and retrieving specific portions of data from a [...] Read more.
The exponential growth of data over the past decade has created new challenges in transforming raw information into actionable knowledge, particularly through the development of data products. The latter is essentially the result of querying and retrieving specific portions of data from a data storage architecture at various levels of granularity. Traditionally, this transformation depends on domain experts manually analyzing datasets and providing feedback to effectively describe or annotate data that facilitates data retrieval. Nevertheless, this is a very time-consuming process that highlights the need for its potential automation. To address this challenge, the present paper proposes a framework which utilizes Large Language Models to support data product discovery through semantic metadata reasoning and executable query prototyping. The framework is evaluated across two domains and three levels of concept complexity to assess the LLM’s ability to identify relevant datasets and generate executable data product queries under varying analytical demands. The findings indicate that LLMs perform effectively in simpler scenarios, but their performance declines as conceptual complexity and dataset volume increase. Full article
Show Figures

Figure 1

26 pages, 530 KB  
Review
Generative AI as a General-Purpose Technology: Foundations, Applications, and Labor Market Implications Through 2030
by Maikel Leon
Big Data Cogn. Comput. 2026, 10(3), 69; https://doi.org/10.3390/bdcc10030069 - 27 Feb 2026
Cited by 4 | Viewed by 5492
Abstract
Generative Artificial Intelligence (AI) has transitioned from a research milestone to a general-purpose technology with wide-ranging implications for organizations, labor markets, and information systems. Thanks to improvements in deep learning, generative adversarial networks (GANs), variational autoencoders (VAEs), diffusion models, transformer-based language models, and [...] Read more.
Generative Artificial Intelligence (AI) has transitioned from a research milestone to a general-purpose technology with wide-ranging implications for organizations, labor markets, and information systems. Thanks to improvements in deep learning, generative adversarial networks (GANs), variational autoencoders (VAEs), diffusion models, transformer-based language models, and reinforcement learning from human feedback (RLHF), generative AI can now create high-quality text, images, audio, code, and other types of content. This review synthesizes the core technical foundations and best practices for training, evaluation, and governance, with an emphasis on scalability and human oversight. The paper examines applications across customer service, marketing, software development, healthcare, finance, law, logistics, and the creative industries, and assesses the labor implications of generative AI using a sociotechnical lens. This study also develops a disruption index that integrates task exposure, adoption rates, time savings, and skill complementarity. The paper concludes with actionable recommendations for policymakers, organizations, and workers, emphasizing the importance of reskilling, algorithmic transparency, and inclusive innovation. Taken together, these contributions situate generative AI within broader debates about automation, augmentation, and the future of work. Full article
(This article belongs to the Section Large Language Models and Embodied Intelligence)
26 pages, 1059 KB  
Article
Validating the Effectiveness of Fine-Tuning for Semantic Classification of Japanese Katakana Words: An Analysis of Frequency and Polysemy Effects on Accuracy
by Kazuki Kodaki and Minoru Sasaki
Big Data Cogn. Comput. 2026, 10(3), 67; https://doi.org/10.3390/bdcc10030067 - 26 Feb 2026
Viewed by 1318
Abstract
In semantic classification of katakana words using large language models and pre-trained language models, semantic divergences from original English meanings, such as those found in Wasei-Eigo which is Japanese-made English, and the inherent sense ambiguity in katakana words may affect model accuracy. To [...] Read more.
In semantic classification of katakana words using large language models and pre-trained language models, semantic divergences from original English meanings, such as those found in Wasei-Eigo which is Japanese-made English, and the inherent sense ambiguity in katakana words may affect model accuracy. To analyze the impact of these loanword semantic characteristics on classification accuracy, we created a large-scale dataset from the Balanced Corpus of Contemporary Written Japanese. We extracted 403,819 sentences covering 230 katakana words defined in dictionaries and suitable for word sense disambiguation tasks, and used the gpt-4.1-mini model to predict the meaning of the target words based on their context, to create annotation data. We then fine-tuned the pre-trained language model DeBERTa V3 with this data. We compared baseline and fine-tuned model accuracy, dividing data into four quadrants based on frequency and polysemy to conduct statistical analysis and explore strategies for improving accuracy. We also tested the hypothesis that high-frequency, low-polysemy words would achieve the highest accuracy, while low-frequency, high-polysemy words would achieve the lowest. As a result, the fine-tuned model showed an average accuracy improvement of approximately 53% compared to the baseline model. As hypothesized, high-frequency, low-polysemy words achieved the highest accuracy (93.93%), while low-frequency, high-polysemy words achieved the lowest (81.14%). Our analysis quantitatively revealed that both frequency and polysemy contributed to accuracy improvement, but polysemy had a greater impact on accuracy than frequency. Full article
(This article belongs to the Special Issue Artificial Intelligence (AI) and Natural Language Processing (NLP))
Show Figures

Figure 1

23 pages, 1201 KB  
Article
Comparative Read Performance Analysis of PostgreSQL and MongoDB in E-Commerce: An Empirical Study of Filtering and Analytical Queries
by Jovita Urnikienė, Vaida Steponavičienė and Svetoslav Atanasov
Big Data Cogn. Comput. 2026, 10(2), 66; https://doi.org/10.3390/bdcc10020066 - 19 Feb 2026
Viewed by 3995
Abstract
This paper presents a comparative analysis of read performance for PostgreSQL and MongoDB in e-commerce scenarios, using identical datasets in a resource-constrained single-host environment. The results demonstrate that PostgreSQL executes complex analytical queries 1.6–15.1 times faster, depending on the query type and data [...] Read more.
This paper presents a comparative analysis of read performance for PostgreSQL and MongoDB in e-commerce scenarios, using identical datasets in a resource-constrained single-host environment. The results demonstrate that PostgreSQL executes complex analytical queries 1.6–15.1 times faster, depending on the query type and data volume. The study employed synthetic data generation with the Faker library across three stages, processing up to 300,000 products and executing each of 6 query types 15 times. Both filtering and analytical queries were tested on non-indexed data in a controlled localhost environment with PostgreSQL 17.5 and MongoDB 7.0.14, using default configurations. PostgreSQL showed 65–80% shorter execution times for multi-criteria queries, while MongoDB required approximately 33% less disk space. These findings suggest that normalized relational schemas are advantageous for transactional e-commerce systems where analytical queries dominate the workload. The results are directly applicable to small and medium e-commerce developers operating in budget-constrained, single-host deployment environments when choosing between relational and document-oriented databases for structured transactional data with read-heavy analytical workloads. A minimal indexed validation confirms that the baseline trends remain consistent under a simple indexing configuration. Future work will examine broader indexing strategies, write-intensive workloads, and distributed deployment scenarios. Full article
Show Figures

Figure 1

74 pages, 45992 KB  
Perspective
Integration of Lean Analytics and Industry 6.0: A Novel Meta-Theoretical Framework for Antifragile, Generative AI-Orchestrated, Circular–Regenerative, and Hyper-Connected Manufacturing Ecosystems
by Mohammad Shahin, Mazdak Maghanaki and F. Frank Chen
Big Data Cogn. Comput. 2026, 10(2), 65; https://doi.org/10.3390/bdcc10020065 - 17 Feb 2026
Cited by 9 | Viewed by 2751
Abstract
The convergence of Lean manufacturing principles with Industry 4.0 has yielded significant operational improvements, yet the emerging paradigm of Industry 6.0—characterized by antifragile, autonomous, and sustainable systems—demands a fundamental rethinking of existing analytical frameworks. This paper introduces the Industry 6.0 Lean Analytics (I6LA) [...] Read more.
The convergence of Lean manufacturing principles with Industry 4.0 has yielded significant operational improvements, yet the emerging paradigm of Industry 6.0—characterized by antifragile, autonomous, and sustainable systems—demands a fundamental rethinking of existing analytical frameworks. This paper introduces the Industry 6.0 Lean Analytics (I6LA) Framework, a novel meta-theoretical approach that integrates Lean principles with the core concepts of Industry 6.0. By systematically analyzing the limitations of current Lean analytics in the context of Industry 6.0 requirements, we identify critical gaps in areas such as system resilience, AI-driven autonomy, and circular economy integration. The I6LA Framework addresses these gaps through four new theoretical pillars: Antifragile Lean Systems Theory, generative AI-Orchestrated Value Streams, Circular–Regenerative Analytics, and Hyper-Connected Ecosystem Integration. This research provides a new set of mathematical models for measuring antifragility, generative orchestration efficiency, and circularity, offering a comprehensive analytical toolkit for the next generation of manufacturing. The framework’s primary contribution is a paradigm shift from optimizing stable, human-in-the-loop systems to managing dynamic, autonomous ecosystems that thrive on volatility and are regenerative by design. This paper provides both a robust theoretical foundation and practical implementation guidance for organizations navigating the transition to Industry 6.0. Full article
(This article belongs to the Section Cognitive System)
Show Figures

Figure 1

23 pages, 291 KB  
Review
Cognitive Assemblages: Living with Algorithms
by Stéphane Grumbach
Big Data Cogn. Comput. 2026, 10(2), 63; https://doi.org/10.3390/bdcc10020063 - 16 Feb 2026
Cited by 2 | Viewed by 2461
Abstract
The rapid expansion of algorithmic systems has transformed cognition into an increasingly distributed and collective enterprise, giving rise to what can be described as cognitive assemblages, dynamic constellations of humans, institutions, data infrastructures, and artificial agents. This paper traces the historical and conceptual [...] Read more.
The rapid expansion of algorithmic systems has transformed cognition into an increasingly distributed and collective enterprise, giving rise to what can be described as cognitive assemblages, dynamic constellations of humans, institutions, data infrastructures, and artificial agents. This paper traces the historical and conceptual evolution that has led to this shift. First, we show how cognition, once conceived as the property of autonomous individuals, has progressively become embedded in socio-technical networks in which algorithmic processes participate as co-agents. Second, we revisit the progressive awareness of human cognitive limits, from bounded rationality to contemporary theories of extended mind. These frameworks anticipate and help explain today’s hybrid cognitive ecologies. Third, we assess the philosophical implications for Enlightenment ideals of autonomy, rationality, and self-governance, showing how these concepts must be reinterpreted in light of pervasive algorithmic intermediation. Finally, we examine global initiatives that seek to integrate augmented cognitive capacities into large-scale cybernetic forms of societal coordination, ranging from digital platforms and data spaces to AI-driven governance systems. These developments offer new opportunities for steering complex societies under conditions of globalization, environmental disruption, and the rise of autonomous intelligent systems, yet they also raise profound questions regarding control, accountability, and democratic legitimacy. We argue that understanding cognitive assemblages is essential to designing socio-technical systems capable of supporting collective intelligence while preserving human values in an era of accelerating complexity. Full article
Show Figures

Figure 1

17 pages, 1902 KB  
Article
Skill Classification of Youth Table Tennis Players Using Sensor Fusion and the Random Forest Algorithm
by Yung-Hoh Sheu, Cheng-Yu Huang, Li-Wei Tai, Tzu-Hsuan Tai and Sheng K. Wu
Big Data Cogn. Comput. 2026, 10(2), 62; https://doi.org/10.3390/bdcc10020062 - 15 Feb 2026
Viewed by 1402
Abstract
This study addresses the issue of inaccurate results in traditional table tennis player classification, which is often influenced by subjective judgment and environmental factors, by proposing a youth table tennis player classification system based on sensor fusion and the random forest algorithm. The [...] Read more.
This study addresses the issue of inaccurate results in traditional table tennis player classification, which is often influenced by subjective judgment and environmental factors, by proposing a youth table tennis player classification system based on sensor fusion and the random forest algorithm. The system utilizes an embedded intelligent table tennis racket equipped with an ICM20948 nine-axis sensor and a wireless transmission module to capture real-time acceleration and angular velocity data during players’ strokes while synchronously employing a camera with OpenPose to extract joint angle variations. A total of 40 players’ stroke data were collected. Due to the limited sample size of top-tier players, the Synthetic Minority Over-sampling Technique (SMOTE) was applied, resulting in a final dataset of 360 records. Multiple key motion indicators were then computed and stored in a dedicated database. Experimental results showed that the proposed system, powered by the random forest algorithm, achieved a classification accuracy of 91.3% under conventional cross-validation, while subject-independent LOSO validation yielded a more conservative accuracy of 70.89%, making it a valuable reference for coaches and referees in conducting objective player classification. Future work will focus on expanding the dataset of domestic high-performance athletes and integrating precise sports science resources to further enhance the system’s performance and algorithmic models, thereby promoting the scientific selection of national team players and advancing the intelligent development of table tennis. Full article
(This article belongs to the Section Artificial Intelligence and Multi-Agent Systems)
Show Figures

Figure 1

17 pages, 698 KB  
Review
What Distinguishes AI-Generated from Human Writing? A Rapid Review of the Literature
by Georgios P. Georgiou
Big Data Cogn. Comput. 2026, 10(2), 55; https://doi.org/10.3390/bdcc10020055 - 8 Feb 2026
Cited by 3 | Viewed by 9427
Abstract
Large language models (LLMs) are now routine writing tools across various domains, intensifying questions about when text should be treated as human-authored, artificial intelligence (AI)-generated, or collaboratively produced. This rapid review aims to identify cue families reported in empirical studies as distinguishing AI [...] Read more.
Large language models (LLMs) are now routine writing tools across various domains, intensifying questions about when text should be treated as human-authored, artificial intelligence (AI)-generated, or collaboratively produced. This rapid review aims to identify cue families reported in empirical studies as distinguishing AI from human-authored text and to assess how stable these cues are across genres/tasks, text lengths, and revision conditions. Following the Preferred Reporting Items for Systematic Reviews and Meta-Analysis (PRISMA) guidelines, we searched four online databases for peer-reviewed empirical articles (1 January 2022–1 January 2026). After deduplication and screening, 40 studies were included. Evidence converged on five cue families: surface, discourse/pragmatic, epistemic/content, predictability/probabilistic, and provenance. Surface cues dominated the literature and were the most consistently operationalized. Discourse/pragmatic cues followed, particularly in discipline-bound academic genres where stance and metadiscourse differentiated AI from human writing. Predictability/probabilistic cues were central in detector-focused studies, while epistemic/content cues emerged primarily in tasks where grounding and authenticity were salient. Provenance cues were concentrated in watermarking research. Across studies, cue stability was consistently conditional rather than universal. Specifically, surface and discourse cues often remained discriminative within constrained genres, but shifted with register and discipline; probabilistic cues were powerful yet fragile under paraphrasing, post-editing, and evasion; and provenance signals required robustness to editing, mixing, and span localization. Overall, the literature indicates that AI–human distinction emerges from layered and context-dependent cue profiles rather than from any single reliable marker. High-stakes decisions, therefore, require condition-aware interpretation, triangulation across multiple cue families, and human oversight rather than automated classification in isolation. Full article
(This article belongs to the Special Issue Machine Learning Applications in Natural Language Processing)
Show Figures

Figure 1

25 pages, 2294 KB  
Article
SiAraSent: From Features to Deep Transformers for Large-Scale Arabic Sentiment Analysis
by Omar Almousa, Yahya Tashtoush, Anas AlSobeh, Plamen Zahariev and Omar Darwish
Big Data Cogn. Comput. 2026, 10(2), 49; https://doi.org/10.3390/bdcc10020049 - 3 Feb 2026
Cited by 4 | Viewed by 1557
Abstract
Sentiment analysis of Arabic text, particularly on social media platforms, presents a formidable set of unique challenges that stem from the language’s complex morphology, its numerous dialectal variations, and the frequent and nuanced use of emojis to convey emotional context. This paper presents [...] Read more.
Sentiment analysis of Arabic text, particularly on social media platforms, presents a formidable set of unique challenges that stem from the language’s complex morphology, its numerous dialectal variations, and the frequent and nuanced use of emojis to convey emotional context. This paper presents SiAraSent, a hybrid framework that integrates traditional text representations, emoji-aware features, and deep contextual embeddings based on Arabic transformers. Starting from a strong and fully interpretable baseline built on Term Frequency–Inverse Definition Frequency (TF–IDF)-weighted character and word N-grams combined with emoji embeddings, we progressively incorporate SinaTools for linguistically informed preprocessing and AraBERT for contextualized encodings. The framework is evaluated on a large-scale dataset of 58,751 Arabic tweets labeled for sentiment polarity. Our design works within four experimental configurations: (1) a baseline traditional machine learning architecture that employs TF-IDF, N-grams, and emoji features with an Support Vector Machine (SVM) classifier; (2) an Large-language Model (LLM) feature extraction approach that leverages deep contextual embeddings from the pre-trained AraBERT model; (3) a novel hybrid fusion model that concatenates traditional morphological features, AraBERT embeddings, and emoji-based features into a high-dimensional vector; and (4) a fully fine-tuned AraBERT model specifically adapted for the sentiment classification task. Our experiments demonstrate the remarkable efficacy of our proposed framework, with the fine-tuned AraBERT architecture achieving an accuracy of 93.45%, a significant 10.89% improvement over the best traditional baseline. Full article
(This article belongs to the Special Issue Advances in Natural Language Processing and Text Mining: 2nd Edition)
Show Figures

Figure 1

18 pages, 800 KB  
Article
Free Access to World News: Reconstructing Full-Text Articles from GDELT
by Andrea Fronzetti Colladon and Roberto Vestrelli
Big Data Cogn. Comput. 2026, 10(2), 45; https://doi.org/10.3390/bdcc10020045 - 2 Feb 2026
Cited by 1 | Viewed by 3546
Abstract
News data have become essential resources across various disciplines. Still, access to full-text news corpora remains challenging due to high costs and the limited availability of free alternatives. This paper presents a novel Python package (gdeltnews) that reconstructs full-text newspaper articles at near-zero [...] Read more.
News data have become essential resources across various disciplines. Still, access to full-text news corpora remains challenging due to high costs and the limited availability of free alternatives. This paper presents a novel Python package (gdeltnews) that reconstructs full-text newspaper articles at near-zero cost by leveraging the Global Database of Events, Language, and Tone (GDELT) Web News NGrams 3.0 dataset. Our method merges overlapping n-grams extracted from global online news to rebuild complete articles. We validate the approach on a benchmark set of 2211 articles from major U.S. news outlets, achieving up to 95% text similarity against original articles based on Levenshtein and SequenceMatcher metrics. Our tool facilitates economic forecasting, computational social science, information science, and natural language processing applications by enabling free and large-scale access to full-text news data. Full article
(This article belongs to the Section Big Data)
Show Figures

Figure 1

34 pages, 2216 KB  
Review
Big Data Analytics and AI for Consumer Behavior in Digital Marketing: Applications, Synthetic and Dark Data, and Future Directions
by Leonidas Theodorakopoulos, Alexandra Theodoropoulou and Christos Klavdianos
Big Data Cogn. Comput. 2026, 10(2), 46; https://doi.org/10.3390/bdcc10020046 - 2 Feb 2026
Cited by 8 | Viewed by 15259
Abstract
In the big data era, understanding and influencing consumer behavior in digital marketing increasingly relies on large-scale data and AI-driven analytics. This narrative, concept-driven review examines how big data technologies and machine learning reshape consumer behavior analysis across key decision-making areas. After outlining [...] Read more.
In the big data era, understanding and influencing consumer behavior in digital marketing increasingly relies on large-scale data and AI-driven analytics. This narrative, concept-driven review examines how big data technologies and machine learning reshape consumer behavior analysis across key decision-making areas. After outlining the theoretical foundations of consumer behavior in digital settings and the main data and AI capabilities available to marketers, this paper discusses five application domains: personalized marketing and recommender systems, dynamic pricing, customer relationship management, data-driven product development and fraud detection. For each domain, it highlights how algorithmic models affect targeting, prediction, consumer experience and perceived fairness. This review then turns to synthetic data as a privacy-oriented way to support model development, experimentation and scenario analysis, and to dark data as a largely underused source of behavioral insight in the form of logs, service interactions and other unstructured records. A discussion section integrates these strands, outlines implications for digital marketing practice and identifies research needs related to validation, governance and consumer trust. Finally, this paper sketches future directions, including deeper integration of AI in real-time decision systems, increased use of edge computing, stronger consumer participation in data use, clearer ethical frameworks and exploratory work on quantum methods. Full article
(This article belongs to the Section Big Data)
Show Figures

Figure 1

36 pages, 1519 KB  
Review
Thinking Machines: Mathematical Reasoning in the Age of LLMs
by Andrea Asperti, Alberto Naibo and Claudio Sacerdoti Coen
Big Data Cogn. Comput. 2026, 10(1), 38; https://doi.org/10.3390/bdcc10010038 - 22 Jan 2026
Cited by 3 | Viewed by 6327
Abstract
Large Language Models (LLMs) have demonstrated impressive capabilities in structured reasoning and symbolic tasks, with coding emerging as a particularly successful application. This progress has naturally motivated efforts to extend these models to mathematics, both in its traditional form, expressed through natural-style mathematical [...] Read more.
Large Language Models (LLMs) have demonstrated impressive capabilities in structured reasoning and symbolic tasks, with coding emerging as a particularly successful application. This progress has naturally motivated efforts to extend these models to mathematics, both in its traditional form, expressed through natural-style mathematical language, and in its formalized counterpart, expressed in a symbolic syntax suitable for automatic verification. Yet, despite apparent parallels between programming and proof construction, advances in formalized mathematics have proven significantly more challenging. This gap raises fundamental questions about the nature of reasoning in current LLM architectures, the role of supervision and feedback, and the extent to which such models maintain an internal notion of computational or deductive state. In this article, we review the current state-of-the-art in mathematical reasoning with LLMs, focusing on recent models and benchmarks. We explore three central issues at the intersection of machine learning and mathematical cognition: (i) the trade-offs between traditional and formalized mathematics as training and evaluation domains; (ii) the structural and methodological reasons why proof synthesis remains more brittle than code generation; and (iii) whether LLMs genuinely represent or merely emulate a notion of evolving logical state. Our goal is not to draw rigid distinctions but to clarify the present boundaries of these systems and outline promising directions for their extension. Full article
Show Figures

Figure 1

30 pages, 1372 KB  
Systematic Review
A Systematic Review and Bibliometric Analysis of Automated Multiple-Choice Question Generation
by Dimitris Mitroulias and Spyros Sioutas
Big Data Cogn. Comput. 2026, 10(1), 35; https://doi.org/10.3390/bdcc10010035 - 18 Jan 2026
Cited by 3 | Viewed by 2965
Abstract
The aim of this study is to systematically capture, synthesize, and evaluate current research trends related to Automated Multiple-Choice Question Generation as they emerge within the broader landscape of natural language processing (NLP) and large language model (LLM)-based educational and assessment research. A [...] Read more.
The aim of this study is to systematically capture, synthesize, and evaluate current research trends related to Automated Multiple-Choice Question Generation as they emerge within the broader landscape of natural language processing (NLP) and large language model (LLM)-based educational and assessment research. A systematic search and selection process was conducted following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, using predefined inclusion and exclusion criteria. A total of 240 eligible publications indexed in the Scopus database were identified and analyzed. To provide a comprehensive overview of this evolving research landscape, a bibliometric analysis was performed utilizing performance analysis and scientific mapping methods, supported by the Bibliometrix (version 4.2.2) R package and VOSviewer (version 1.6.19) software. The findings of the performance analysis indicate a steady upward trend in publications and citations, with significant contributions from leading academic institutions—primarily from the United States—and a strong presence in high quality academic journals. Scientific mapping through co-authorship analysis reveals that, despite the increasing research activity, there remains a need for enhanced collaborative efforts. Bibliographic coupling organizes the analyzed literature into seven thematic clusters, highlighting the main research axes and their diachronic evolution. Furthermore, co-word analysis identifies emerging research trends and underexplored directions, indicating substantial opportunities for future investigation. To the best of our knowledge, this study represents the first systematic bibliometric analysis that examines Automated Multiple-Choice Question Generation research within the context of the broader LLM-driven educational assessment literature. By mapping the relevant scientific production and identifying research gaps and future directions, this work contributes to a more coherent understanding of the field and supports the ongoing development of research at the intersection of generative AI and educational assessment. Full article
(This article belongs to the Special Issue Generative AI and Large Language Models)
Show Figures

Figure 1

22 pages, 6241 KB  
Article
Using Large Language Models to Detect and Debunk Climate Change Misinformation
by Zeinab Shahbazi and Sara Behnamian
Big Data Cogn. Comput. 2026, 10(1), 34; https://doi.org/10.3390/bdcc10010034 - 17 Jan 2026
Cited by 7 | Viewed by 2895
Abstract
The rapid spread of climate change misinformation across digital platforms undermines scientific literacy, public trust, and evidence-based policy action. Advances in Natural Language Processing (NLP) and Large Language Models (LLMs) create new opportunities for automating the detection and correction of misleading climate-related narratives. [...] Read more.
The rapid spread of climate change misinformation across digital platforms undermines scientific literacy, public trust, and evidence-based policy action. Advances in Natural Language Processing (NLP) and Large Language Models (LLMs) create new opportunities for automating the detection and correction of misleading climate-related narratives. This study presents a multi-stage system that employs state-of-the-art large language models such as Generative Pre-trained Transformer 4 (GPT-4), Large Language Model Meta AI (LLaMA) version 3 (LLaMA-3), and RoBERTa-large (Robustly optimized BERT pretraining approach large) to identify, classify, and generate scientifically grounded corrections for climate misinformation. The system integrates several complementary techniques, including transformer-based text classification, semantic similarity scoring using Sentence-BERT, stance detection, and retrieval-augmented generation (RAG) for evidence-grounded debunking. Misinformation instances are detected through a fine-tuned RoBERTa–Multi-Genre Natural Language Inference (MNLI) classifier (RoBERTa-MNLI), grouped using BERTopic, and verified against curated climate-science knowledge sources using BM25 and dense retrieval via FAISS (Facebook AI Similarity Search). The debunking component employs RAG-enhanced GPT-4 to produce accurate and persuasive counter-messages aligned with authoritative scientific reports such as those from the Intergovernmental Panel on Climate Change (IPCC). A diverse dataset of climate misinformation categories covering denialism, cherry-picking of data, false causation narratives, and misleading comparisons is compiled for evaluation. Benchmarking experiments demonstrate that LLM-based models substantially outperform traditional machine-learning baselines such as Support Vector Machines, Logistic Regression, and Random Forests in precision, contextual understanding, and robustness to linguistic variation. Expert assessment further shows that generated debunking messages exhibit higher clarity, scientific accuracy, and persuasive effectiveness compared to conventional fact-checking text. These results highlight the potential of advanced LLM-driven pipelines to provide scalable, real-time mitigation of climate misinformation while offering guidelines for responsible deployment of AI-assisted debunking systems. Full article
(This article belongs to the Special Issue Natural Language Processing Applications in Big Data)
Show Figures

Figure 1

25 pages, 5725 KB  
Article
Data-Driven Life-Cycle Assessment of Household Air Conditioners: Identifying Low-Carbon Operation Patterns Based on Big Data Analysis
by Genta Sugiyama, Tomonori Honda and Norihiro Itsubo
Big Data Cogn. Comput. 2026, 10(1), 32; https://doi.org/10.3390/bdcc10010032 - 15 Jan 2026
Cited by 1 | Viewed by 3714
Abstract
Air conditioners are a critical adaptation measure against heat- and cold-related risks under climate change. However, their electricity use and refrigerant leakage increase greenhouse gas (GHG) emissions. This study developed a data-driven life-cycle assessment (LCA) framework for residential room air conditioners in Japan [...] Read more.
Air conditioners are a critical adaptation measure against heat- and cold-related risks under climate change. However, their electricity use and refrigerant leakage increase greenhouse gas (GHG) emissions. This study developed a data-driven life-cycle assessment (LCA) framework for residential room air conditioners in Japan by integrating large-scale field operation data with life-cycle climate performance (LCCP) modeling. We aggregated 1 min records for approximately 4100 wall-mounted split units and evaluated the 10-year LCCP across nine climate regions. Using the annual operating hours and electricity consumption, we classified the units into four behavioral quadrants and quantified the life-cycle GHG emissions and parameter sensitivities for each. The results show that the use-phase electricity dominated the total emissions, and that even under the same climate and capacity class, the 10-year per-unit emissions differed by roughly a factor of two between the high- and low-load quadrants. The sensitivity analysis identified the heating hours and the setpoint–indoor temperature difference as the most influential drivers, whereas the grid CO2 intensity, equipment lifetime, and refrigerant assumptions were of secondary importance. By replacing a single assumed use scenario with empirical profiles and behavior-based clusters, the proposed framework improves the representativeness of the LCA for air conditioners. This enabled the design of cluster-specific mitigation strategies. Full article
(This article belongs to the Special Issue Energy Conservation Towards a Low-Carbon and Sustainability Future)
Show Figures

Figure 1

26 pages, 3399 KB  
Article
Adaptive Data Prefetching for File Storage Systems Using Online Machine Learning
by George Savva and Herodotos Herodotou
Big Data Cogn. Comput. 2026, 10(1), 28; https://doi.org/10.3390/bdcc10010028 - 10 Jan 2026
Viewed by 2200
Abstract
Data prefetching is essential for modern file storage systems operating in large-scale cloud and data-intensive environments, where high performance increasingly depends on intelligent, adaptive mechanisms. Traditional rule-based methods and recently proposed machine learning-based techniques often struggle to cope with the complex and rapidly [...] Read more.
Data prefetching is essential for modern file storage systems operating in large-scale cloud and data-intensive environments, where high performance increasingly depends on intelligent, adaptive mechanisms. Traditional rule-based methods and recently proposed machine learning-based techniques often struggle to cope with the complex and rapidly evolving data access patterns characteristic of big-data workloads. In this paper, we introduce an online, streaming machine learning (SML) approach for predictive data prefetching that retrieves useful data into the cache ahead of time. We present a novel online training framework that extracts features in real time and continuously updates streaming ML models to learn and adapt from large and dynamic access streams. Building on this framework, we design new SML-driven prefetching algorithms that decide when, how, and what data to prefetch into the cache with minimal overhead. Extensive experiments using production traces from Huawei Technologies Inc. and Google workloads from the SNIA IOTTA repository demonstrate that our intelligent policies consistently deliver the highest byte hits among competing approaches, achieving 97% prefetch byte precision and reducing data access latency by up to 2.8 times. These results show that streaming ML can deliver immediate performance gains and offers a scalable foundation for future adaptive storage systems. Full article
Show Figures

Figure 1

Back to TopTop