Journal Description
Big Data and Cognitive Computing
Big Data and Cognitive Computing
is an international, peer-reviewed, open access journal on big data and cognitive computing published monthly online by MDPI.
- Open Access— free for readers, with article processing charges (APC) paid by authors or their institutions.
- High Visibility: indexed within Scopus, ESCI (Web of Science), dblp, Inspec, Ei Compendex, and other databases.
- Journal Rank: JCR - Q1 (Computer Science, Theory and Methods) / CiteScore - Q1 (Computer Science Applications)
- Rapid Publication: manuscripts are peer-reviewed and a first decision is provided to authors approximately 23.3 days after submission; acceptance to publication is undertaken in 4.8 days (median values for papers published in this journal in the first half of 2026).
- Recognition of Reviewers: reviewers who provide timely, thorough peer-review reports receive vouchers entitling them to a discount on the APC of their next publication in any MDPI journal, in appreciation of the work done.
- Journal Cluster of Artificial Intelligence: AI, AI in Medicine, Algorithms, BDCC, MAKE, MTI, Stats, Virtual Worlds, Computers and Journal of Superintelligence.
Impact Factor:
5.3 (2025);
5-Year Impact Factor:
4.9 (2025)
Latest Articles
Enhanced TabNet with Entmax-Based Sparse Attention and Modified GLU for Interpretable Cardiovascular Risk Prediction Using an Edge-IoT Framework
Big Data Cogn. Comput. 2026, 10(8), 271; https://doi.org/10.3390/bdcc10080271 - 11 Aug 2026
Abstract
►
Show Figures
Cardiovascular diseases remain the leading global cause of mortality, necessitating continuous monitoring solutions that extend beyond clinical settings. This paper proposes a real-time, end-to-end Edge-IoT framework for cardiovascular risk assessment that integrates biomedical signal acquisition, edge processing, and interpretable deep learning. The system
[...] Read more.
Cardiovascular diseases remain the leading global cause of mortality, necessitating continuous monitoring solutions that extend beyond clinical settings. This paper proposes a real-time, end-to-end Edge-IoT framework for cardiovascular risk assessment that integrates biomedical signal acquisition, edge processing, and interpretable deep learning. The system includes a three-tier architecture: (i) physiological signal acquisition using AD8232 ECG, MAX30102 photoplethysmography, DS18B20 temperature, and NEO-6M GPS sensors interfaced with an ESP32 microcontroller; (ii) real-time signal preprocessing, including digital filtering, normalisation, and PQRST feature extraction performed at the edge; and (iii) cloud-based analytics using an Enhanced TabNet classifier with modified attention mechanisms for cardiovascular risk prediction. The Enhanced TabNet architecture incorporates Entmax-based sparse attention and modified Gated Linear Units to improve predictive performance and clinical interpretability. Signal quality enhancement using Kalman filtering and class imbalance correction using SMOTE further support robust model performance. The Enhanced TabNet model achieves 97.43% accuracy, 96.18% precision, and 97.24% recall on the combined Cleveland, Hungarian, Switzerland, Long Beach VA, and Statlog heart disease datasets ( ). The developed Edge-IoT prototype maintains an end-to-end communication and processing latency below 200 ms. The framework also includes automated risk alert generation via SMS when the predicted cardiovascular risk probability exceeds a predefined threshold (e.g., 0.85), including the patient’s vital information and geolocation to support emergency response. The integrated edge-cloud architecture with attention-based feature selection provides interpretable cardiovascular risk predictions while maintaining the computational efficiency required for potential continuous patient monitoring outside hospital settings.
Full article
Open AccessArticle
Evaluating Machine Learning and Deep Learning Models for Early Detection of Alzheimer’s and Parkinson’s Disease: An Explainable Dual-Dataset Study on Clinical Generalizability
by
Muneera Mohammed Al-Dossary and Atta Rahman
Big Data Cogn. Comput. 2026, 10(8), 270; https://doi.org/10.3390/bdcc10080270 - 11 Aug 2026
Abstract
Neurodegenerative diseases such as Alzheimer’s disease (AD) and Parkinson’s disease (PD) pose a significant global healthcare burden due to challenges in early diagnosis. This study investigates the clinical generalizability of machine learning (ML) and deep learning (DL) models for early AD and PD
[...] Read more.
Neurodegenerative diseases such as Alzheimer’s disease (AD) and Parkinson’s disease (PD) pose a significant global healthcare burden due to challenges in early diagnosis. This study investigates the clinical generalizability of machine learning (ML) and deep learning (DL) models for early AD and PD classification across multiple data modalities. A dual-dataset framework was employed, combining global benchmarks (ADNI, PPMI, OASIS, and UCI voice) with local clinical data from King Fahd Hospital of the University (KFHU) in Saudi Arabia. We evaluated ensemble methods, SVMs, neural networks, CNNs, and LSTMs. On structured global data, tree-based ensembles achieved the best performance, with Random Forest reaching 88.24% accuracy for PD and Gradient Boosting achieving 94.44% for AD. For neuroimaging, an LSTM on CNN features attained 98.68% accuracy on a curated MRI dataset. A critical finding was a substantial generalization gap: models excelling on global data showed markedly reduced performance on local KFHU data, with AUC values between 0.50 and 0.77. This degradation is attributed to real-world clinical challenges including severe class imbalance, diagnostic uncertainty in EHRs, and heterogeneous feature representations. The results underscore that data quality and modality are often more consequential than algorithmic complexity. This study provides a reproducible validation framework, highlights the necessity of institution-specific evaluation, establishes performance benchmarks for Saudi healthcare (with the caveat that local sample sizes remain small and the results should be interpreted as exploratory), and demonstrates the use of explainable AI to validate clinical relevance.
Full article
(This article belongs to the Special Issue Artificial Intelligence and Big Data Analytics for Sustainable Healthcare Systems)
►▼
Show Figures

Figure 1
Open AccessArticle
From Unstructured Reports to Exploratory Causal Modeling: A Modality-Aware AI Pipeline for Infrastructure Delay Analysis
by
Florence Gundidza, Masato Kikuchi and Tadachika Ozono
Big Data Cogn. Comput. 2026, 10(8), 269; https://doi.org/10.3390/bdcc10080269 - 11 Aug 2026
Abstract
Infrastructure project reports contain rich narrative evidence on delay causes, yet transforming such unstructured text into reliable causal knowledge remains challenging because reports mix confirmed events with hypothetical, conditional, or localized statements. This study proposes an eight-stage computational pipeline that converts infrastructure project
[...] Read more.
Infrastructure project reports contain rich narrative evidence on delay causes, yet transforming such unstructured text into reliable causal knowledge remains challenging because reports mix confirmed events with hypothetical, conditional, or localized statements. This study proposes an eight-stage computational pipeline that converts infrastructure project evaluation reports into a Bayesian-network model for exploratory structure learning and probabilistic dependency modeling. The central methodological contribution is a modality-aware extraction layer that distinguishes confirmed, project-wide delay evidence from conditional, hypothetical, or component-level statements before causal analysis. The pipeline was evaluated on 55 road infrastructure project reports financed by the Asian Development Bank, the African Development Bank, and JICA, from which delay events across 15 cause categories were extracted and stratified by epistemic modality and scope. Ablation analysis shows that the principal dependency structure recovered by the Bayesian network is not recoverable without modality-aware filtering, indicating that evidence-quality stratification materially shapes downstream causal-structure exploration. Among the recovered dependencies, a financial-to-project-management pathway was the most consistent signal: its undirected skeleton edge was the only relationship recovered by all four causal-discovery algorithms tested (with the orientation determined only by the score-based search), its association was nominally positive—though weak and not uniformly discernible—across nine extraction models spanning three commercial vendors and open-weight families, and it is consistent with prior delay-factor literature. Its model-based scenario contrast ( , 95% CI ) is reported as hypothesis-generating rather than as a validated policy effect: under structure-learning uncertainty, the interval extends to zero, and the effect magnitude and the specific learned edge depend on the extraction model and the small effective sample. These findings suggest that incorporating modality awareness into narrative-evidence extraction improves the reliability of exploratory causal-structure analysis from infrastructure project reports.
Full article
(This article belongs to the Special Issue Text Mining and Big Data Analysis)
►▼
Show Figures

Figure 1
Open AccessReview
A Review of Human-AI Complementarities Across Multiple Dimensions of Organisational Complexity
by
Ganesh Sankaran, Marco A. Palomino and Guido Siestrup
Big Data Cogn. Comput. 2026, 10(8), 268; https://doi.org/10.3390/bdcc10080268 - 11 Aug 2026
Abstract
The growing capabilities of artificial intelligence (AI) have not translated straightforwardly into organisational value. A persistent disconnect—the “last-mile problem”—arises from structural gaps between idealised AI tasks and real-world organisational contexts. Synthesising insights from organisational theory, cognitive science, and computer science, we have developed
[...] Read more.
The growing capabilities of artificial intelligence (AI) have not translated straightforwardly into organisational value. A persistent disconnect—the “last-mile problem”—arises from structural gaps between idealised AI tasks and real-world organisational contexts. Synthesising insights from organisational theory, cognitive science, and computer science, we have developed a five-dimensional diagnostic framework that maps the challenges of human-AI collaboration across Integration, Representation, Scale, Temporality, and Adequacy gaps. These gaps illuminate how socio-technical complexity, contextualised problem representations, interdependencies among agents, dynamic environments, and limitations in current AI reasoning collectively constrain full automation and demand human judgement. By reviewing the historical evolution of AI—from symbolic systems to machine learning, generative models, and emerging agentic approaches—we show that augmentation remains the dominant and most viable mode of use in complex environments. An illustrative system-dynamics example demonstrates how improvements in algorithmic performance do not automatically yield proportional system-level gains. Overall, our framework provides researchers with a conceptual lens and practitioners with a diagnostic tool for assessing complementarities and informing the design of human-AI collaborations. The framework is offered as a conceptual synthesis and diagnostic instrument rather than an empirically validated model.
Full article
(This article belongs to the Special Issue Big Data and Cognitive Computing in 2026)
►▼
Show Figures

Figure 1
Open AccessArticle
A Self-Adaptive Agentic Mixture-of-Agent Families with Agent-to-Agent Communication for Automatic Sentiment Analysis
by
Wiam Saidi, Boutaina Satouri, Abdellatif El Abderrahmani and Khalid Satori
Big Data Cogn. Comput. 2026, 10(8), 267; https://doi.org/10.3390/bdcc10080267 - 10 Aug 2026
Abstract
Automatic sentiment analysis requires models that can effectively handle texts of varying levels of complexity. Monolithic methods use the same algorithm uniformly, without taking into account the intrinsic complexity of the input text. In response to this need, we developed a Dynamic Agentic
[...] Read more.
Automatic sentiment analysis requires models that can effectively handle texts of varying levels of complexity. Monolithic methods use the same algorithm uniformly, without taking into account the intrinsic complexity of the input text. In response to this need, we developed a Dynamic Agentic Mixture-of-Agents with Inter-Agent Communication and Adaptive Routing for Robust Sentiment Analysis (DAMA-Sent). This approach merges three distinct algorithmic paradigms: statistical learning, deep learning, and attention models. The decomposition process is carried out in a sophisticated system featuring a hierarchical routing system and an inter-agent communication system based on differentiable attention. Furthermore, each agent has a self-reflection module, a weighting mechanism that takes uncertainty into account and allows the agent to assess its own reliability. Finally, an adaptive early exit system halts processing as soon as an appropriate confidence threshold is reached or the computation budget is exhausted. In-depth analyses conducted on a corpus of tweets from American airlines reveal that the suggested approach can adjust to the intrinsic variability of textual complexity and surpasses static ensemble methods in terms of accuracy and computational cost, achieving an accuracy of 95.96%. Additional validations corroborate these trends, demonstrating both the structural relevance and the validity of our proposed framework.
Full article
(This article belongs to the Section Artificial Intelligence and Multi-Agent Systems)
►▼
Show Figures

Figure 1
Open AccessArticle
Descriptive Process Mining of Pulmonary Clinical Pathways Before and During COVID-19
by
Luca Murazzano, Paolo Landa, Jean-Baptiste Gartner and André Côté
Big Data Cogn. Comput. 2026, 10(8), 266; https://doi.org/10.3390/bdcc10080266 - 10 Aug 2026
Abstract
Understanding how clinical pathways evolve over time is essential for characterizing care processes. It also helps identify potential shifts in diagnostic and organizational practices. This study provides a descriptive analysis of patient trajectories for four major respiratory conditions: lung cancer, interstitial fibrosis, chronic
[...] Read more.
Understanding how clinical pathways evolve over time is essential for characterizing care processes. It also helps identify potential shifts in diagnostic and organizational practices. This study provides a descriptive analysis of patient trajectories for four major respiratory conditions: lung cancer, interstitial fibrosis, chronic obstructive pulmonary disease (COPD), and pneumonia. Trajectories were compared between a pre-COVID-19 period (2018–2019) and a COVID-19 period (2020–2022) in a specialized hospital. Using process mining applied to administrative event logs, we examined three aspects of care: the structure and sequencing of activities, the timing of transitions between care encounters, and imaging timeliness. The analysis spanned inpatient, emergency department, and outpatient settings. Indicators of care duration and transition timing revealed heterogeneous temporal patterns. Several conditions showed shorter intervals in the COVID-19 period, whereas others varied little. Activity-level analyses complemented these findings. Process maps indicated stable structural components in many pathways, together with differences in timing and execution. In the emergency department, care shifted toward bedside radiography, whereas CT chest volumes remained relatively stable across periods and settings. Imaging timeliness stayed consistently high in the emergency department and relatively stable for most inpatient conditions. Outcome-related indicators, including 30-day readmission and prolonged care trajectories, showed only modest differences between periods. Overall, the study demonstrates the value of process mining for describing real-world clinical pathways and identifying temporal variations in care. These results provide a foundation for future work that integrates richer clinical information and analytical approaches capable of assessing causal relationships.
Full article
(This article belongs to the Topic Data Intelligence and Computational Analytics)
►▼
Show Figures

Figure 1
Open AccessArticle
Joint MLP and Token Pruning for Personalizing Vision Transformers
by
Zhiyue Li, Tong Liu, Feng Huang, Xinzhi Huang and Zhihao Zou
Big Data Cogn. Comput. 2026, 10(8), 265; https://doi.org/10.3390/bdcc10080265 - 9 Aug 2026
Abstract
ViTs have achieved excellent performance in image recognition tasks, but their large parameter counts and high computational complexity limit their deployment on resource-constrained devices. Most existing ViT pruning methods adopt class-agnostic pruning strategies, which fail to distinguish the diverse structural requirements of different
[...] Read more.
ViTs have achieved excellent performance in image recognition tasks, but their large parameter counts and high computational complexity limit their deployment on resource-constrained devices. Most existing ViT pruning methods adopt class-agnostic pruning strategies, which fail to distinguish the diverse structural requirements of different target classes. As a result, they are prone to removing critical features, leading to class-wise accuracy imbalance in practical deployment. To address this issue, this paper proposes a class-aware joint pruning framework for ViTs, which collaboratively compresses the model from two orthogonal dimensions: MLP neurons and visual tokens. Specifically, (1) based on first-order Taylor expansion, we quantify the contribution of each MLP neuron to the target classes and adaptively prune redundant neurons to achieve structured compression, followed by lightweight fine-tuning on the target class subset; (2) we propose a Class-Guided Token Selection (CGTS) method, which constructs class prototype vectors using a few support samples of the target classes and then dynamically selects patch tokens that are semantically highly relevant to the target classes during inference in a zero-shot manner, requiring no additional training or fine-tuning. The two modules complement each other, achieving dual compression from the parameter dimension and the inference data dimension. Experiments on CIFAR-100 and TinyImageNet datasets using DeiT-Tiny/Small models demonstrate that, compared with state-of-the-art pruning methods, our method reduces GMACs on target class subsets by up to 48%, improves inference speed by nearly 50%, and requires only 0.8 KB of additional storage overhead per subset, ultimately achieving a superior trade-off among accuracy, computational efficiency, and storage overhead.
Full article
(This article belongs to the Special Issue Deep Learning in Sensor Networks and Real-Time and Embedded Applications)
►▼
Show Figures

Figure 1
Open AccessArticle
A Three-Stage Cross-Lingual Knowledge Transfer Approach Based on the XLM-RoBERTa Model for Detecting Fake News in Ukrainian
by
Volodymyr Smahliuk, Yaroslav Kovivchak and Yurii Kynash
Big Data Cogn. Comput. 2026, 10(8), 264; https://doi.org/10.3390/bdcc10080264 - 8 Aug 2026
Abstract
In recent years, there has been an increase in the amount of fake news in the media, which is why fact-checking systems are gaining popularity, particularly those that use natural language processing (NLP) to quickly identify and flag fake news. One of the
[...] Read more.
In recent years, there has been an increase in the amount of fake news in the media, which is why fact-checking systems are gaining popularity, particularly those that use natural language processing (NLP) to quickly identify and flag fake news. One of the main limitations in the development of such systems is the limited number of datasets containing verified information, which are necessary for the effective training of models. The situation is particularly critical for non-English datasets, specifically those in the Ukrainian language. This article proposes a three-stage algorithm for training a model to recognize fake news in the Ukrainian language. At the core of the proposed approach lies the multilingual transformer model XLM-RoBERTa, which solves this problem by utilizing cross-lingual knowledge transfer from English to Ukrainian. This approach means there is no need to search for a large, high-quality dataset in Ukrainian; instead, a significantly smaller dataset in Ukrainian can be used for the final calibration of the model. The model developed as a result of the experiment proved effective in extreme low-resource scenarios, achieving 90.7% accuracy on just 500 training records and outperforming the baseline model by 9.7%.
Full article
(This article belongs to the Special Issue Artificial Intelligence (AI) and Natural Language Processing (NLP))
►▼
Show Figures

Figure 1
Open AccessArticle
GluKDnet: A Lightweight Blood Glucose Prediction Model Based on Heterogeneous Knowledge Distillation
by
Aowei Teng, Xiaoyu Sun, Hongru Li and Xia Yu
Big Data Cogn. Comput. 2026, 10(8), 263; https://doi.org/10.3390/bdcc10080263 - 6 Aug 2026
Abstract
Accurate blood glucose prediction is essential for glycemic management in people with diabetes, but the size of many high-performing models complicates execution on resource-constrained artificial pancreas controllers. We propose GluKDnet, a lightweight glucose-forecasting model for prospective Android-smartphone-based mobile edge controllers. GluKDnet transfers the
[...] Read more.
Accurate blood glucose prediction is essential for glycemic management in people with diabetes, but the size of many high-performing models complicates execution on resource-constrained artificial pancreas controllers. We propose GluKDnet, a lightweight glucose-forecasting model for prospective Android-smartphone-based mobile edge controllers. GluKDnet transfers the representational capacity of a time-series foundation model to a compact causal CNN through heterogeneous knowledge distillation. The teacher model, MOMENT, is adapted to continuous glucose monitoring (CGM) data through risk-event-aware masking, which prioritizes abnormal glucose levels, rapid glucose fluctuations, and CGM-defined dawn phenomenon and Somogyi effect patterns during masked reconstruction. A transient-state and steady-state distillation module jointly aligns ordered patch-level dynamics and day-level summaries between teacher and student. Using DLCP3 for teacher pretraining and leave-one-patient-out evaluation on OhioT1DM, GluKDnet achieves RMSE values of 20.04, 32.04, and 45.33 mg/dL for 30, 60, and 120 min prediction, respectively, with about 53K parameters. Auxiliary evaluation on T1D-UoM shows a similar offline accuracy–parameter count pattern. On a vivo V2072A Android smartphone, the 30 min model achieved a mean API inference latency of 0.470 ms (P95: 0.855 ms), a maximum sampled process proportional-set-size memory of 47.06 MiB, and a median incremental device energy estimate of 0.277 mJ per inference. These device measurements characterize the exported student model under one hardware and software configuration; insulin dosing and prospective closed-loop clinical evaluation remain outside the scope of this study.
Full article
(This article belongs to the Special Issue Artificial Intelligence-Driven Analysis of Big Health Data)
►▼
Show Figures

Figure 1
Open AccessArticle
Secure Knowledge Retrieval for English-Teaching Agents: A Multi-Stage Auditing and Knowledge Purification Method
by
Jiming Yin, Xianfeng Xie, Shanyi Guo, Jiawei Chen and Jie Cui
Big Data Cogn. Comput. 2026, 10(8), 262; https://doi.org/10.3390/bdcc10080262 - 6 Aug 2026
Abstract
►▼
Show Figures
English-teaching agents use external knowledge retrieval to update instructional content, broaden domain coverage, and personalize support beyond standalone large language models (LLMs). However, open sources may introduce harmful, biased, or misleading content into retrieval-augmented generation (RAG) pipelines, affecting learners’ judgment, cultural understanding, and
[...] Read more.
English-teaching agents use external knowledge retrieval to update instructional content, broaden domain coverage, and personalize support beyond standalone large language models (LLMs). However, open sources may introduce harmful, biased, or misleading content into retrieval-augmented generation (RAG) pipelines, affecting learners’ judgment, cultural understanding, and value formation. To address this problem, this study proposes a multi-stage secure knowledge retrieval method for English-teaching agents. The method coordinates safeguards across knowledge-source access, retrieval execution, and model output. At the access stage, custom rules and Semgrep-based static scanning perform preliminary risk screening. At the retrieval stage, LLM-based dynamic evaluation identifies tool-description contamination and cross-file data-flow risks. At the output stage, semantic-embedding pre-screening, LLM review, and bounded knowledge purification detect and rewrite risky responses. Our experiments use public safety benchmarks, a mixed corpus of benign and poisoned passages, synthetic purification cases, and controlled end-to-end teaching scenarios. Compared with vanilla RAG, the framework reduces Poison Exposure@5 from 92.0% to 3.0% and retrieval attack success from 86.0% to 2.0% while preserving retrieval coverage. These results provide preliminary evidence that the framework can empower English teaching by enabling agents to deliver safer materials and trustworthy support for classroom questioning, academic writing, and intercultural learning.
Full article

Figure 1
Open AccessPerspective
Cognitive Entanglement: Toward a Developmental Framework of the Human-AI Coevolutionary Leap
by
Xiao-Kun Wu, Min Chen and Giancarlo Fortino
Big Data Cogn. Comput. 2026, 10(8), 261; https://doi.org/10.3390/bdcc10080261 - 5 Aug 2026
Abstract
Large language models have become routine participants in everyday cognition. Their role has widened from retrieval and text generation to helping users define problems, organize arguments, make judgments, and interpret themselves. Yet their cognitive consequences are strikingly divergent. For some users, generative AI
[...] Read more.
Large language models have become routine participants in everyday cognition. Their role has widened from retrieval and text generation to helping users define problems, organize arguments, make judgments, and interpret themselves. Yet their cognitive consequences are strikingly divergent. For some users, generative AI appears to reduce critical engagement, independent judgment, and tolerance for difficulty. For others, the same class of systems becomes a medium for conceptual expansion, reflective questioning, and higher-order learning. This divergence cannot be explained by model capability alone. Mental effort is often treated as a cost to be reduced. Yet repeated delegation may also reduce opportunities to practice the processes required for independent judgment. The key issue is developmental: how sustained AI use changes users’ cognitive capacities over time. This perspective proposes cognitive entanglement as a framework for understanding the developmental consequences of sustained human-AI coupling. Cognitive entanglement refers to a relation in which human and AI activity become mutually shaping, irreducible to either party alone and organized across different developmental levels. The framework examines how repeated interaction with AI changes the ways users formulate problems, evaluate reasons, and make judgments. Unlike theories that locate the boundaries of cognition (the extended mind, enactivism) or explain the mechanisms of consciousness (global workspace, higher-order, predictive-processing, and integrated-information theories), cognitive entanglement examines whether sustained AI use preserves, weakens, or reorganizes users’ cognitive capacities. The article argues that current AI systems are often optimized for fluency, immediacy, and user satisfaction, and this may reduce the productive difficulty that supports higher-order cognitive development. If AI is to support human cognitive growth, design must move beyond answer provision and efficiency maximization toward the organization of productive human-AI relations: relations that challenge users’ initial assumptions while providing support appropriate to the task and the user’s level of expertise. The argument draws on philosophy of mind, cognitive science, and learning science, and compares divergent approaches to coupling in order to specify which forms of relation carry which developmental consequences. The concept shifts attention from AI as a tool or automation system to the developmental consequences of sustained human-AI interaction.
Full article
(This article belongs to the Topic Learning to Live with Gen-AI)
►▼
Show Figures

Figure 1
Open AccessArticle
SETTA: Parameter-Free Test-Time Adaptation for Graph Neural Networks via Spectral-Energy-Guided Semantic Refinement
by
Dongyang Yu, Xia Cui and Rong Xiao
Big Data Cogn. Comput. 2026, 10(8), 260; https://doi.org/10.3390/bdcc10080260 - 4 Aug 2026
Abstract
Node classification is a central graph data mining task, yet repeated message passing can over-smooth representations and degrade frozen graph neural network (GNN) predictions after deployment. We present SETTA (Spectral-Energy Test-Time Adaptation), a prediction-level graph test-time adaptation framework that refines frozen outputs without
[...] Read more.
Node classification is a central graph data mining task, yet repeated message passing can over-smooth representations and degrade frozen graph neural network (GNN) predictions after deployment. We present SETTA (Spectral-Energy Test-Time Adaptation), a prediction-level graph test-time adaptation framework that refines frozen outputs without test labels, gradients, parameter updates, or learnable adaptation parameters. SETTA denoises features for semantic-neighbor construction, adds complementary semantic routes while preserving observed edges, monitors a smoothness-energy proxy during diffusion, and accepts refinements through entropy-based gating. Configurations are fixed by a dataset-level protocol or selected using validation data only. Across six mostly homophilic benchmarks with 2708–19,717 nodes, SETTA improved a frozen two-layer GCN on every dataset and achieved the highest mean accuracy among the evaluated methods on five, with gains of 4.61, 3.08, and 2.01 percentage points on Cora, CiteSeer, and PubMed, respectively. Positive mean gains were also observed across all 30 dataset–backbone settings. Ablations and transition analyses indicate that semantic injection is most beneficial on sparse citation graphs and that selective refinement limits harmful changes. The current dense implementation supports benchmark-scale, amortized refinement; scalability and robustness on heterophilic graphs remain open.
Full article
(This article belongs to the Special Issue Theories and Applications on Data Mining in Graph Neural Networks)
►▼
Show Figures

Figure 1
Open AccessArticle
Adapting Large-Scale Foundation Models for Turkic Speech-to-Speech Translation: Fine-Tuned Cascade and Direct Approaches
by
Aidana Karibayeva, Vladislav Karyukin, Oleg Myssov, Dina Amirova, Balzhan Abduali and Adina Karybayeva
Big Data Cogn. Comput. 2026, 10(8), 259; https://doi.org/10.3390/bdcc10080259 - 4 Aug 2026
Abstract
This paper presents a novel approach to speech-to-speech (STS) translation for low-resource Turkic languages. Today, STS has progressed rapidly for high-resource languages; the Turkic family remains significantly underrepresented. Consequently, it is quite challenging to develop a reliable speech translation system for these languages.
[...] Read more.
This paper presents a novel approach to speech-to-speech (STS) translation for low-resource Turkic languages. Today, STS has progressed rapidly for high-resource languages; the Turkic family remains significantly underrepresented. Consequently, it is quite challenging to develop a reliable speech translation system for these languages. To address this issue, we have developed two speech translation systems (STS) specifically tailored to Turkic languages with limited resources. The first, TurkicCascadeSTS, is a cascaded system that combines a fine-tuned Whisper-medium speech recognition model, GPT translation, and separate speech synthesis models for each language. The second system is a direct speech translation model based on a fine-tuned SeamlessM4Tv2. Both systems have been tested for translation into Turkic languages, using 24,656 audio recordings per language. The TurkicCascadeSTS system delivered far better results: the average BLEU score rose from 4.68 to 30.60; the METEOR score rose from 15.03 to 44.42; and the word error rate (WER) also fell significantly. These improvements are due to the fact that each module of the system was individually fine-tuned to account for the specific characteristics of each language. Although SeamlessM4Tv2 sometimes produces clearer audio, TurkicCascadeSTS generally delivers higher speech and translation quality for all language pairs. This demonstrates that modular, specially tuned systems are an effective solution for translation into Turkic languages, particularly given their complex structure and limited linguistic resources. Such systems could benefit more than 200 million native speakers of Turkic languages.
Full article
(This article belongs to the Section Artificial Intelligence and Multi-Agent Systems)
►▼
Show Figures

Figure 1
Open AccessArticle
A Governed NL-to-SQL Architecture for Reliable Clinical Data Querying and Outpatient Schedule Monitoring
by
Isaac Daroch, Matías Rojas Cabrera, Rodrigo Muñoz Andrade, Alejandra Fernández and Juan Pablo Vásconez
Big Data Cogn. Comput. 2026, 10(8), 258; https://doi.org/10.3390/bdcc10080258 - 3 Aug 2026
Abstract
►▼
Show Figures
The deployment of natural-language-to-SQL (NL-to-SQL) systems in primary healthcare requires more than accurate query generation: it also requires governed data access, robustness to local terminology, and reliable handling of ambiguous user requests. This study evaluated a pilot proof-of-concept integrating a Spanish-language NL-to-SQL assistant
[...] Read more.
The deployment of natural-language-to-SQL (NL-to-SQL) systems in primary healthcare requires more than accurate query generation: it also requires governed data access, robustness to local terminology, and reliable handling of ambiguous user requests. This study evaluated a pilot proof-of-concept integrating a Spanish-language NL-to-SQL assistant with a governed, read-only outpatient scheduling repository derived from the Rayen information system used in a Centro de Salud Familiar (CESFAM) setting in Renca, Chile. The data used by the prototype were accessed through an external company responsible for data management in this context. The prototype was implemented with MindsDB as an artificial intelligence (AI)-enabled database layer and operated on anonymized, delayed secondary scheduling data. Evaluation was conducted through a Slack interface using 252 audited interactions from 42 users, with six assigned interactions per user and up to three exchanges per interaction. SQL correctness reached 240/252 (95.2%), whereas both query correctness and answer correctness reached 144/252 (57.1%). These findings suggest that governed pilot deployment for outpatient schedule monitoring may be feasible under controlled institutional conditions, while indicating that the main remaining barriers are semantic rather than purely syntactic, specifically ambiguity handling, institution-specific operational language, and faithful answer verbalization. The study therefore contributes deployment-oriented pilot evidence and clarifies where operational Spanish NL-to-SQL remains fragile under real institutional constraints.
Full article

Figure 1
Open AccessArticle
PPRL-Stack: A Novel Stacking Architecture for Efficient and Secure Record Linkage
by
Fatima Zahrae Saber, Ali Choukri, Mohammed Amnai and Abderrahim Waga
Big Data Cogn. Comput. 2026, 10(8), 257; https://doi.org/10.3390/bdcc10080257 - 3 Aug 2026
Abstract
Record linkage is a fundamental step in ensuring the quality of data by detecting duplicate records within different databases. Nevertheless, dealing with big, imbalanced databases and ensuring data confidentiality is still difficult in terms of performance and precision. This paper introduces a new
[...] Read more.
Record linkage is a fundamental step in ensuring the quality of data by detecting duplicate records within different databases. Nevertheless, dealing with big, imbalanced databases and ensuring data confidentiality is still difficult in terms of performance and precision. This paper introduces a new Privacy-Preserving Record Linkage (PPRL) method named PPRL-Stack, which uses the Bloom filter encoding technique to hide information and a Stack Ensemble structure for classification. The proposed model consists of Support Vector Machine (SVM) as a base learner and Logistic Regression (LR) as a meta-classifier in combination with the application of Sorted Neighborhood Method (SNM) technique to bring down the time complexity to O(N log N). Experiments conducted on the Freely Extensible Biomedical Record Linkage (FEBRL) and North Carolina Voter Registration (NCVR) databases prove that the proposed PPRL-Stack can obtain nearly perfect discrimination with an F1-score of 0.9921. Particularly, our proposed architecture is more than 340 times and 40 times faster than the latest Siamese Bidirectional Long Short-Term Memory (Bi-LSTM) architecture in training and validation stages, respectively.
Full article
(This article belongs to the Topic Advances in Integrative AI, Machine Learning, and Big Data for Transformative Applications)
►▼
Show Figures

Figure 1
Open AccessArticle
Cognition Orientation Risk Evaluation—A Personality Driven Integrated Model for Phishing Susceptibility
by
Chih-Hong Kao, Chia-Wei Tsai, Yu-Ting Kao and Chao-Lung Chou
Big Data Cogn. Comput. 2026, 10(8), 256; https://doi.org/10.3390/bdcc10080256 - 3 Aug 2026
Abstract
Phishing attacks exploit human vulnerabilities through social engineering techniques; therefore, a multidimensional framework integrating personality and cognition is crucial for tailoring personalized defenses against phishing threats. To prevent phishing attacks, focusing on psychological mechanisms has become the primary approach to address the human-centric
[...] Read more.
Phishing attacks exploit human vulnerabilities through social engineering techniques; therefore, a multidimensional framework integrating personality and cognition is crucial for tailoring personalized defenses against phishing threats. To prevent phishing attacks, focusing on psychological mechanisms has become the primary approach to address the human-centric nature of these threats. In response, we propose an integrated framework that synthesizes dimensions from the Five-Factor Model (FFM) and the Myers–Briggs Type Indicator (MBTI), grounded in Dual Process Theory to explore the cognitive drivers underlying decision-making under threat. To operationalize this framework, we developed a decision tree classifier to quantify the predictive significance of various personality traits and their hierarchical interactions. The results indicate that the personality types associated with the highest phishing risk profiles are ENFP, ESFP, ESTP, and ENFJ. These hierarchical classification results are projected onto the C.O.R.E. Quadrant (Cognition Orientation Risk Evaluation Quadrant) which enables the systematic representation of risk patterns across all 16 personality types. By providing a structured visualization of personality-driven risk patterns, the C.O.R.E. Quadrant offers a practical foundation for developing personalized defense mechanisms and tailored cybersecurity training strategies, moving beyond one-size-fits-all security protocols.
Full article
(This article belongs to the Topic New Trends in Cybersecurity and Data Privacy)
►▼
Show Figures

Figure 1
Open AccessArticle
A Cloud-Based Multidimensional Big Data Framework for Healthcare Analytics: Bridging OLAP and Cognitive Insights in Chronic Pain Management
by
Alfredo Cuzzocrea and Abderraouf Hafsaoui
Big Data Cogn. Comput. 2026, 10(8), 255; https://doi.org/10.3390/bdcc10080255 - 2 Aug 2026
Abstract
By considering the real-life research project Pain-RELife, which focuses attention on big data management and analytics tools over patients suffering from chronic pain, located in the Region Lombardy of North Italy, this paper provides a relevant journey from theory to practice about
[...] Read more.
By considering the real-life research project Pain-RELife, which focuses attention on big data management and analytics tools over patients suffering from chronic pain, located in the Region Lombardy of North Italy, this paper provides a relevant journey from theory to practice about so-called Multidimensional Big Data Analytics tools over (real-life) big healthcare datasets, as dictated by the effective project goals. This constitutes an effective contribution to the state-of-the-art research. Our conceptual and theoretical results are corroborated by a comprehensive campaign of real-life experimental results focused on advanced tools such as Multidimensional Clustering and Multidimensional Regression.
Full article
(This article belongs to the Special Issue Advanced Software and Machine Learning Techniques for System Architectures and Big Data)
►▼
Show Figures

Figure 1
Open AccessArticle
ExPAM: Explainable Personality Assessment Method Using Heterogeneous Linguistic Features and Off-the-Shelf LLMs
by
Elena Ryumina, Dmitry Ryumin, Maxim Markitantov and Alexey Karpov
Big Data Cogn. Comput. 2026, 10(8), 254; https://doi.org/10.3390/bdcc10080254 - 1 Aug 2026
Abstract
►▼
Show Figures
Many organizations increasingly adopt personalization techniques to enhance user satisfaction. However, current systems generally cannot automatically infer and interpret individual personality traits (PTs), although these traits are key drivers of user behavior. While Large Language Models (LLMs) are widely used, they remain poorly
[...] Read more.
Many organizations increasingly adopt personalization techniques to enhance user satisfaction. However, current systems generally cannot automatically infer and interpret individual personality traits (PTs), although these traits are key drivers of user behavior. While Large Language Models (LLMs) are widely used, they remain poorly suited to reliable and explainable Personality Assessment (PA). To address this gap, we propose ExPAM, a novel Explainable Personality Assessment Method that combines hybrid feature fusion with in-context learning in off-the-shelf LLMs to predict Big Five PTs from text. ExPAM explicitly grounds its predictions in interpretable linguistic patterns without requiring LLM fine-tuning. Its hybrid fusion is designed to improve both predictive performance and interpretability in PA. Transformer-based embeddings encode local contextual information, whereas features extracted using the Linguistic Inquiry and Word Count (LIWC) dictionary provide complementary global and local linguistic indicators of PTs. These interpretable feature patterns are included in prompts that guide the LLM to produce both PT predictions and human-understandable explanations. ExPAM shows competitive performance compared with multi-task models on the ChaLearn First Impressions v2 (FIv2) corpus and single-task models on the PANDORA corpus that rely on a single feature set. On FIv2, it achieves a mean accuracy (mAC) of 0.891 and a Concordance Correlation Coefficient (CCC) of 0.333. On PANDORA, it achieves a mean Pearson Correlation Coefficient (PCC) of 0.240 and a CCC of 0.101. Prompting the LLM with hybrid global–local patterns further improves CCC by 9.9% on FIv2 and 15.8% on PANDORA, while changes in mAC and mean PCC remain marginal. Qualitative interpretability analysis reveals trait-specific linguistic patterns, highlighting the potential of ExPAM for psychological research, computational linguistics, and paralinguistic studies.
Full article

Figure 1
Open AccessArticle
STDPatch: A Three-Stream Framework for Long-Term Time Series Forecasting via SG-Filter-Based Decomposition and Patch Refactor
by
Lanlan Li, Di Liu, Shengfa Miao, Ahmed Zahir, Yongkang Mu, Hualong Deng, Xin Jin, Qian Jiang, Puming Wang, Hua Jiang and Shaowen Yao
Big Data Cogn. Comput. 2026, 10(8), 253; https://doi.org/10.3390/bdcc10080253 - 1 Aug 2026
Abstract
Driven by non-stationary factors in real-world sensor-driven applications, time series streams from energy meters, traffic detectors, weather stations, and industrial monitors often exhibit complex patterns composed of long-term trends and multi-scale seasonal fluctuations. Accurately disentangling and modeling these heterogeneous components remains a fundamental
[...] Read more.
Driven by non-stationary factors in real-world sensor-driven applications, time series streams from energy meters, traffic detectors, weather stations, and industrial monitors often exhibit complex patterns composed of long-term trends and multi-scale seasonal fluctuations. Accurately disentangling and modeling these heterogeneous components remains a fundamental challenge in long-term time series forecasting (LTSF). To address this issue, we propose STDPatch, a novel three-stream forecasting framework that combines structural decomposition with architecture specialization. First, we introduce an SG-Filter-Based seasonal–trend decomposition module that employs polynomial fitting to extract shape-preserving trends while reducing seasonal noise. Second, we design a trend decomposition module that further separates the trend component into ascending and descending segments to capture fine-grained evolutionary dynamics. Third, we propose a patch refactor module that adaptively aggregates adjacent patches according to structural similarity, thereby preserving temporal semantic continuity and reducing spurious correlations. Finally, we develop a three-stream architecture that leverages convolutional, linear, and Transformer branches to model seasonal patterns, smooth trends, and non-stationary sub-trends, respectively, with each branch built from efficient, channel-independent components. Extensive experiments on seven real-world sensor-derived benchmark datasets demonstrate that STDPatch consistently outperforms state-of-the-art methods for long-term time series forecasting.
Full article
(This article belongs to the Special Issue Deep Learning in Sensor Networks and Real-Time and Embedded Applications)
►▼
Show Figures

Figure 1
Open AccessSystematic Review
A Systematic Review of Smart Home IoT Security: Applications, Threat Taxonomy, Privacy Risks, and Emerging Defensive Solutions
by
Dalibor Radovanovic, Nikola Savanovic, Jelena Janackovic and Petar Kresoja
Big Data Cogn. Comput. 2026, 10(8), 252; https://doi.org/10.3390/bdcc10080252 - 1 Aug 2026
Abstract
The rapid proliferation of Internet of Things (IoT) technologies has transformed the modern home into a complex cyber–physical ecosystem encompassing hundreds of millions of connected devices globally. Smart homes support automation, energy management, and healthcare monitoring, but they also introduce a broad and
[...] Read more.
The rapid proliferation of Internet of Things (IoT) technologies has transformed the modern home into a complex cyber–physical ecosystem encompassing hundreds of millions of connected devices globally. Smart homes support automation, energy management, and healthcare monitoring, but they also introduce a broad and evolving range of security and privacy challenges. This review examines 233 sources published between 2018 and May 2025, selected through a PRISMA-informed process covering five major academic databases and relevant standards and technical reports. It discusses communication protocols, including Matter, develops a Threat-Layer-Defense synthesis matrix covering ten attack categories; examines the practical limitations of AI-based anomaly detection and blockchain-based trust management; and derives recommendations for manufacturers, platform providers, users, and regulators. Privacy challenges, regulatory frameworks, and user behavior are considered alongside technical threats. The findings suggest that scalable smart home security requires coordinated progress in protocol standardization, enforceable device update lifecycles, gateway-level anomaly detection, and privacy-preserving local analytics rather than reliance on a single technical solution.
Full article
(This article belongs to the Special Issue Scalability and Interoperability in the Artificial Intelligence of Things)
►▼
Show Figures

Figure 1
Journal Menu
► ▼ Journal Menu-
- BDCC Home
- Aims & Scope
- Editorial Board
- Reviewer Board
- Topical Advisory Panel
- Instructions for Authors
- Special Issues
- Topics
- Sections
- Article Processing Charge
- Indexing & Archiving
- Editor’s Choice Articles
- Most Cited & Viewed
- Journal Statistics
- Journal History
- Journal Awards
- Conferences
- Editorial Office
Journal Browser
► ▼ Journal BrowserHighly Accessed Articles
Latest Books
E-Mail Alert
News
Topics
Topic in
AI, BDCC, Future Internet, Information, Sustainability
Big Data and Artificial Intelligence, 3rd Edition
Topic Editors: Miltiadis D. Lytras, Andreea Claudia SerbanDeadline: 30 August 2026
Topic in
Computers, Electronics, Future Internet, IoT, Network, Sensors, JSAN, Technologies, BDCC
Challenges and Future Trends of Wireless Networks
Topic Editors: Stefano Scanzio, Ramez Daoud, Jetmir Haxhibeqiri, Pedro SantosDeadline: 30 September 2026
Topic in
Applied Sciences, Electronics, Information, Drones, Sensors, BDCC, IJGI
Advances in Integrative AI, Machine Learning, and Big Data for Transformative Applications
Topic Editors: Peiying Zhang, Athanasios V. VasilakosDeadline: 31 October 2026
Topic in
Applied Sciences, Future Internet, AI, Analytics, BDCC
Data Intelligence and Computational Analytics
Topic Editors: Carson K. Leung, Fei Hao, Xiaokang ZhouDeadline: 30 November 2026
Conferences
Special Issues
Special Issue in
BDCC
Machine Learning Methodologies and Applications in Cybersecurity Data Analysis
Guest Editors: Biao Han, Xiaoyan Wang, Xiucai Ye, Na ZhaoDeadline: 31 August 2026
Special Issue in
BDCC
Artificial Intelligence Applications for Cultural Heritage: Innovations, Challenges, and Opportunities
Guest Editors: Giuseppe Maria Luigi Sarnè, Fabrizio Messina, Domenico RosaciDeadline: 31 August 2026
Special Issue in
BDCC
Visual Media Literacy in the Age of AI-Generated Content
Guest Editor: Jacob GroshekDeadline: 31 August 2026
Special Issue in
BDCC
Intelligent Integration of Sensing-Communication-Computing Continuum for Next-Generation Wireless Networks
Guest Editors: Li Zhu, Lei LiuDeadline: 31 August 2026




