Next Issue
Volume 13, June
Previous Issue
Volume 13, April
 
 

Informatics, Volume 13, Issue 5 (May 2026) – 12 articles

Cover Story (view full-size image): Automated monitoring of honey bee traffic at hive entrances provides critical information for environmental informatics and precision agriculture, yet current frameworks rely on simple, unidirectional entrance and exit assumptions. In this paper, we address a key algorithmic gap by systematically demonstrating that complex, compound movements such as U-turns represent a substantial portion of bee activity. Forcing these complex trajectories into simpler categories introduces significant error. We propose extended classification methods utilizing threshold, displacement, and angular cues to explicitly model bidirectional movement patterns. Validated against a manually annotated dataset, our threshold-based and displacement-based bidirectional algorithms achieve near-perfect trajectory classification, significantly improving the fidelity of automated behavioral analyses. View this paper
  • Issues are regarded as officially published after their release is announced to the table of contents alert mailing list.
  • You may sign up for e-mail alerts to receive table of contents of newly released issues.
  • PDF is the official format for papers published in both, html and pdf forms. To view the papers in pdf format, click on the "PDF Full-text" link, and use the free Adobe Reader to open them.
Order results
Result details
Section
Select all
Export citation of selected articles as:
28 pages, 892 KB  
Article
System Quality, Perceived Compulsion, and Tax Literacy as Determinants of Continuous Usage Intention: Evidence from Indonesia’s Mandatory Coretax Platform
by Adi Prasetyo Tedjakusuma, Waiphot Kulachai, Phakawanaporn Phisuthisuwan and Andri Dayarana K. Silalahi
Informatics 2026, 13(5), 76; https://doi.org/10.3390/informatics13050076 - 21 May 2026
Viewed by 1716
Abstract
Governments worldwide are mandating digital tax platforms, yet little is understood about what sustains taxpayer engagement beyond legally compelled minimum use. This study extends the Technology Acceptance Model (TAM) with perceived compulsion and tax literacy to examine continuous usage intention toward Coretax, Indonesia’s [...] Read more.
Governments worldwide are mandating digital tax platforms, yet little is understood about what sustains taxpayer engagement beyond legally compelled minimum use. This study extends the Technology Acceptance Model (TAM) with perceived compulsion and tax literacy to examine continuous usage intention toward Coretax, Indonesia’s mandatory Core Tax Administration System. Using survey data from 535 active users analysed with PLS-SEM, six of eight hypotheses are supported: system quality drives perceived ease of use, which amplifies perceived usefulness, and both usefulness and user satisfaction independently predict continuous usage intention. Contrary to predictions derived from self-determination theory, perceived compulsion positively influences satisfaction, suggesting institutional acceptance of a mandate redirects evaluative attention toward system performance rather than generating resistance. Tax literacy does not moderate the usefulness–continuance pathway but independently increases engagement intentions, pointing to literacy programmes as direct engagement levers rather than amplifiers. These findings extend TAM into mandatory post-adoption contexts and propose institutional acceptance as a boundary condition for coercion theory in IS research. Full article
Show Figures

Figure 1

14 pages, 472 KB  
Article
Robust Multi-View Ensemble Broad Learning for Semi-Supervised Classification
by Ziyang Dong, Mianfen Lin and Zhiwen Yu
Informatics 2026, 13(5), 75; https://doi.org/10.3390/informatics13050075 - 21 May 2026
Viewed by 571
Abstract
In semi-supervised learning scenarios, the presence of limited labeled data and abundant unlabeled samples poses significant challenges to model robustness and generalization. Although the semi-supervised broad learning system (SSBLS) effectively exploits manifold structure through graph Laplacian regularization, its optimization is typically formulated under [...] Read more.
In semi-supervised learning scenarios, the presence of limited labeled data and abundant unlabeled samples poses significant challenges to model robustness and generalization. Although the semi-supervised broad learning system (SSBLS) effectively exploits manifold structure through graph Laplacian regularization, its optimization is typically formulated under the mean square error (MSE) criterion, which is sensitive to noise and outliers. To address this limitation, this paper introduces the maximum mixture correntropy criterion (MMC) into the SSBLS framework and proposes a model termed M2C-SSBLS. By replacing the conventional MSE loss with a mixture correntropy-based objective, the proposed method enhances robustness against non-Gaussian noise and abnormal samples while preserving the computational efficiency and analytical solution property of the BLS. Furthermore, to improve representation diversity and reduce model variance, a multi-view ensemble extension, named EC-SSBLS, is proposed. This method constructs multiple feature views through a random feature subspace strategy, and independently trains an M2C-SSBLS base learner on each subspace. Finally, the predicted results of each view are fused through a voting mechanism. Experiments on benchmark UCI datasets under noise-free, 10% and 20% label noise settings demonstrate that the proposed M2C-SSBLS consistently outperforms conventional SSBLS and other advanced semi-supervised learning approaches. The ensemble extension EC-SSBLS further enhances performance, particularly in noisy environments, validating the effectiveness of combining MMC-based optimization with multi-view ensemble learning. Full article
Show Figures

Figure 1

23 pages, 6792 KB  
Article
Exploring Shifts in User Behavior Through Longitudinal Data from a Digital Platform for Art and Culture
by Minas Pergantis
Informatics 2026, 13(5), 74; https://doi.org/10.3390/informatics13050074 - 18 May 2026
Viewed by 675
Abstract
Digital repositories have been an important gateway for the dissemination of information regarding objects of art and cultural heritage throughout the World Wide Web, but the vast number of available artifacts, both historical and modern, makes their discovery by interested users an arduous [...] Read more.
Digital repositories have been an important gateway for the dissemination of information regarding objects of art and cultural heritage throughout the World Wide Web, but the vast number of available artifacts, both historical and modern, makes their discovery by interested users an arduous task. Often deviating from general-purpose search behavior, people searching online for art and culture adjust their habits to address this challenge. In this study, real-world data from the federated search engine and online art and culture repository ArtBoulevard are used to explore this evolution throughout a period of three years. By collecting and analyzing a large amount of user session data, this research aims to investigate user engagement, query formulation, search behavior, and in-platform and outbound engagement in order to outline the longitudinal behavioral patterns of the platform’s user base. Over the period of the analysis, shifts and trends are identified and discussed within the ever-evolving context of behavioral analysis in the field. This process leads to useful insights that are not only indicative of the platform’s limited but global user base, but which can be useful to all stakeholders active in content dissemination and may also be relevant to broader discussions about the changes in the discovery pathways in art and cultural heritage. Full article
(This article belongs to the Section Social Informatics and Digital Humanities)
Show Figures

Figure 1

15 pages, 1281 KB  
Article
An Empirical Study of Federated BERT for Decentralized Twitter Sentiment Analysis
by Oumaima Louzar, Abdelaziz Elbaghdadi, Ahmed El Oualkadi, Ouafae Baida and Abdelouahid Lyhyaoui
Informatics 2026, 13(5), 73; https://doi.org/10.3390/informatics13050073 - 18 May 2026
Viewed by 985
Abstract
Twitter/x has become a key platform for analyzing public opinion on a large scale; however, traditional centralized approaches raise significant concerns regarding privacy and data governance. To address these challenges, this paper presents an empirical study of a federated learning approach based on [...] Read more.
Twitter/x has become a key platform for analyzing public opinion on a large scale; however, traditional centralized approaches raise significant concerns regarding privacy and data governance. To address these challenges, this paper presents an empirical study of a federated learning approach based on a BERT model for decentralized sentiment analysis at the tweet level. This study focuses on evaluating the effectiveness of transformer-based models under realistic non-independent and identically distributed (non-IID) data distributions across distributed clients. The proposed approach enables collaborative model training without sharing raw tweet data, thereby preserving user privacy while leveraging knowledge from multiple sources. The model is evaluated over 100 communication rounds using the Sentiment140 dataset, distributed among four clients with heterogeneous data distributions. Experimental results demonstrate stable convergence and robust performance, with an accuracy of 95.00%, an F1 score of 95.00%, and a PR-AUC of 96.76%. It should be noted that the federated model performs within 1.2% of a centralized baseline, indicating minimal performance degradation despite data sharing constraints. Full article
Show Figures

Figure 1

30 pages, 4692 KB  
Article
TERA: A Trade-Off Evaluation and Resource-Aware Framework for Spam and Phishing Email Detection
by Chanankorn Jandaeng, Peeravit Koad, Mohamad Fadli Zolkipli and Jurairat Phuttharak
Informatics 2026, 13(5), 72; https://doi.org/10.3390/informatics13050072 - 12 May 2026
Cited by 1 | Viewed by 1341 | Correction
Abstract
Email spam and phishing detection is typically evaluated using accuracy-centric metrics under implicitly unconstrained computational settings. However, in practical deployment scenarios—particularly in real-time and resource-constrained environments—models with comparable predictive performance may differ substantially in inference latency and resource usage, directly affecting their operational [...] Read more.
Email spam and phishing detection is typically evaluated using accuracy-centric metrics under implicitly unconstrained computational settings. However, in practical deployment scenarios—particularly in real-time and resource-constrained environments—models with comparable predictive performance may differ substantially in inference latency and resource usage, directly affecting their operational feasibility. This paper introduces TERA, a deployment-aware evaluation framework that formulates model assessment as a constraint-aware decision problem. Instead of aggregating performance and efficiency into a single objective, TERA treats predictive performance as a feasibility requirement that defines an admissible set of models. Within this feasible region, operational factors such as latency and resource usage are used to differentiate among candidates through structured, multi-dimensional analysis. Experiments on benchmark email datasets show that multiple models achieve comparable detection performance, forming a region of predictive equivalence. Within this region, significant variations in latency and resource consumption are observed, indicating that predictive equivalence does not imply deployment equivalence. These findings demonstrate that accuracy-based evaluation alone may provide limited guidance for deployment-oriented model selection. By explicitly separating feasibility constraints from preference-based trade-offs, TERA enables transparent and deployment-aligned model evaluation. The framework supports consistent comparison and selection among accuracy-comparable models without altering the role of detection effectiveness as a primary requirement, thereby complementing existing evaluation practices with a structured decision-oriented perspective. Full article
Show Figures

Figure 1

11 pages, 837 KB  
Article
Enhancing the Efficiency of Blockchain Verification Through Resource-Weighted Node Selection
by Vedika Jorika and Nagaratna Medishetty
Informatics 2026, 13(5), 71; https://doi.org/10.3390/informatics13050071 - 8 May 2026
Viewed by 1231
Abstract
Blockchain technology has emerged as a foundational paradigm for building decentralized, transparent, and secure systems, particularly in environments that operate without centralized authority. At the core of these systems are consensus mechanisms that ensure transaction validity and maintain trust among distributed participants. However, [...] Read more.
Blockchain technology has emerged as a foundational paradigm for building decentralized, transparent, and secure systems, particularly in environments that operate without centralized authority. At the core of these systems are consensus mechanisms that ensure transaction validity and maintain trust among distributed participants. However, the efficiency of a blockchain network is strongly influenced by how verifier (or validator) nodes are selected, particularly in sharded architectures where transaction processing is distributed across multiple shards. A critical challenge in blockchain design is selecting appropriate nodes for transaction verification in a manner that is efficient, fair, and resilient to adversarial behavior, while also minimizing communication overhead. Existing approaches often rely primarily on resource availability or on the ability to create blocks, particularly in sharded blockchain architectures. Building on these ideas, this paper proposes a Resource Weighted–Block Score selection algorithm, which integrates a node’s block score with its computational resource availability to guide verifier node selection. Simulation-based evaluation demonstrates that the proposed approach significantly reduces transaction verification latency and improves overall node utilization, thereby enhancing network performance and scalability in sharded blockchain systems. Full article
Show Figures

Figure 1

18 pages, 1241 KB  
Article
Cross-Lingual Transfer of Named Entity Markup with Large Language Models
by Vladimir Barakhnin, Rustam Mussabayev, Davlatyor Mengliev, Alexander Krassovitskiy, Alymzhan Toleu, Daniil Lyutaev, Iskander Akhmetov and Bahodir Ibragimov
Informatics 2026, 13(5), 70; https://doi.org/10.3390/informatics13050070 - 7 May 2026
Viewed by 1448
Abstract
This paper investigates the problem of cross-lingual named entity recognition (NER), which involves automatically identifying entities such as persons, organizations, locations, and other structured elements in text. High-quality NER typically requires manually annotated corpora; however, for many low-resource languages, such data are scarce [...] Read more.
This paper investigates the problem of cross-lingual named entity recognition (NER), which involves automatically identifying entities such as persons, organizations, locations, and other structured elements in text. High-quality NER typically requires manually annotated corpora; however, for many low-resource languages, such data are scarce and costly to produce. The study addresses the following question: can annotated sentences in one language be used to transfer NER markup to their machine-translated counterparts in other languages? To explore this, we propose an approach based on a large language model (LLM) that performs two tasks simultaneously: translating a source sentence and generating BIOES-formatted entity tags for the translated output. To improve robustness and reduce semantic drift, a back-translation step is incorporated to verify meaning preservation by comparing the reconstructed source sentence with the original. The proposed method is compared with two baseline approaches: (1) annotation projection via machine translation and (2) automatic tagging using pre-existing NER tools. Performance is evaluated using standard metrics, including precision, recall, and F1-score. Experimental results demonstrate that the LLM-based approach provides a practical and efficient mechanism for transferring NER annotations across languages. While the method achieves strong and balanced performance, its quality remains influenced by translation accuracy and adherence to annotation constraints. Methodologically, the approach can be considered relatively language-independent, as it relies on general LLM capabilities, a universal tagging scheme, and multilingual semantic representations rather than language-specific model training. Full article
Show Figures

Figure 1

29 pages, 17309 KB  
Article
A Lightweight Hybrid CNN–CBAM Model for Multistage Acute Lymphoblastic Leukemia Classification from Peripheral Blood Smear Images
by Kittipol Wisaeng
Informatics 2026, 13(5), 69; https://doi.org/10.3390/informatics13050069 - 30 Apr 2026
Viewed by 1981
Abstract
Accurate and efficient classification of hematological malignancies from peripheral blood smear (PBS) images remains challenging due to the scarcity of annotated datasets, staining variability, and subtle morphological differences among blood cancer subtypes. To address these limitations, this study proposes an Advanced Lightweight Deep [...] Read more.
Accurate and efficient classification of hematological malignancies from peripheral blood smear (PBS) images remains challenging due to the scarcity of annotated datasets, staining variability, and subtle morphological differences among blood cancer subtypes. To address these limitations, this study proposes an Advanced Lightweight Deep Learning (ALDL) framework for the multi-class classification of Acute Lymphoblastic Leukemia (ALL) across four clinically significant stages: Benign, Pro-B, Pre-B, and Early Pre-B. The framework integrates EfficientNetV2-S with Convolutional Block Attention Modules (CBAM) to enhance spatial and channel-wise feature refinement. At the same time, Focal Loss is employed to mitigate class imbalance by prioritizing hard-to-classify samples. A robust preprocessing pipeline, including CLAHE contrast enhancement, Reinhard stain normalization, and data augmentation, improves feature visibility and dataset generalization. Lesion segmentation is performed using RGB-based thresholding and watershed overlay, followed by lesion-level cropping to ensure consistency across inputs. Experimental evaluations on the ALL-DB dataset demonstrate the superior performance of the proposed method, achieving an average accuracy of 96.11%, an F1-score of 95.99%, and an AUC of 0.9875. Comparative analyses against MobileNetV3, ResNet50, DenseNet121, VGG16, and InceptionV3 confirm that the proposed segmentation-guided EfficientNetV2-S + CBAM + Focal Loss framework consistently outperforms conventional CNN architectures across both 70:30 and 60:40 train–test splits. Furthermore, a detailed investigation of color spaces (RGB, HSV, LAB, and HED) indicates that RGB yields the most reliable segmentation and classification results. At the same time, HED enhances lesion visualization at the expense of higher computational cost. The proposed ALDL framework demonstrates strong potential for real-world application as a computer-aided diagnostic (CAD) system for early leukemia detection, offering improved diagnostic reliability, reduced error rates, and practical scalability for clinical environments. Full article
(This article belongs to the Section Health Informatics)
Show Figures

Figure 1

21 pages, 8110 KB  
Article
Beverage Stain Classification Using Hyperspectral Imaging with an L-BFGS-B-Optimized Autoencoder and a Channel-Attention 1D CNN
by Jitendra Shit, Muzaffar Ahmad Dar, Manikandan V M and Partha Pratim Roy
Informatics 2026, 13(5), 68; https://doi.org/10.3390/informatics13050068 - 28 Apr 2026
Viewed by 1675
Abstract
Hyperspectral imaging (HSI) provides rich spectral information and serves as a non-destructive technique for forensic stain analysis. Conventional approaches often exhibit degraded performance due to the high dimensionality and spectral redundancy inherent in hyperspectral data. To address this challenge, a hyperspectral dataset comprising [...] Read more.
Hyperspectral imaging (HSI) provides rich spectral information and serves as a non-destructive technique for forensic stain analysis. Conventional approaches often exhibit degraded performance due to the high dimensionality and spectral redundancy inherent in hyperspectral data. To address this challenge, a hyperspectral dataset comprising nine beverage stains—papaya, coffee, pomegranate, orange, tea, wine, whisky, rum, and brandy—is developed. Building on this dataset, an ensemble framework that combines an optimized autoencoder (AE), channel-attention (CA)-enhanced one-dimensional convolutional neural networks (1D CNNs), and a Limited Memory Broyden–Fletcher–Goldfarb–Shanno (L-BFGS-B)-based weighted fusion strategy is proposed. The autoencoder learns compact latent representations from the 204-band hyperspectral vectors, reducing redundancy while preserving discriminative spectral features. CA emphasizes informative spectral bands and improves stain separability. Multiple 1D CNN models are trained using different latent dimensionalities, and their class probability outputs are fused through an optimized L-BFGS-B weighting scheme, where higher-performing models contribute more strongly to the final decision. Experimental results demonstrate classification accuracies of 96.54%, 97.19%, and 97.86% for the AE32 CA, AE64 CA, and AE128 CA models, respectively, with the optimized ensemble achieving an accuracy of 98.28%. Additionally, the time-dependent evolution of beverage stain reflectance is systematically analyzed using overlapped, normalized reflectance signatures acquired at time intervals of 0 min, 1 h, 2 h, 3 h, 4 h, and 5 h. The results confirm that AE-based latent compression, CA, and L-BFGS-B optimized ensemble fusion enhance hyperspectral beverage stain classification, providing an effective and extensible framework for forensic trace evidence analysis. Full article
(This article belongs to the Section Machine Learning)
Show Figures

Figure 1

20 pages, 17822 KB  
Article
The Evolution of Artificial Intelligence in Marketing: A Bibliometric Analysis of Three Decades (1992–2025)
by Weiming Wang and Zijia Li
Informatics 2026, 13(5), 67; https://doi.org/10.3390/informatics13050067 - 27 Apr 2026
Viewed by 2920
Abstract
Over the past three decades, artificial intelligence (AI) has substantially reshaped marketing research and practice, yet the discipline has not established a systematic understanding of its evolutionary trajectory and intellectual structure. A bibliometric analysis of 1923 Scopus publications (1992–2025) was conducted using CiteSpace [...] Read more.
Over the past three decades, artificial intelligence (AI) has substantially reshaped marketing research and practice, yet the discipline has not established a systematic understanding of its evolutionary trajectory and intellectual structure. A bibliometric analysis of 1923 Scopus publications (1992–2025) was conducted using CiteSpace to explore collaboration patterns, conceptual development, and thematic organization. It identified six evolutionary stages with accelerating innovation cycles, starting with neural networks (1992–2000) and ending with generative AI (2024–2025), with research attention per stage compressing from approximately 9 years to just 2 years. The analysis of the collaboration network shows that the key contributors are India, China, the USA, and the UK. Co-citation analysis indicates that there are three thematic dimensions with seven clusters, namely: (i) AI technological foundations and capabilities, (ii) AI marketing applications and transformation, and (iii) responsible AI governance and ethics. It suggests a Three-Force Evolutionary Framework, which combines technology-push, market-pull, and governance-moderator forces to describe the dynamics of the field. This framework shows that the Regulatory Awakening of 2018 (e.g., GDPR and the Cambridge Analytica incident) guided, not limited, innovation, and highlighted the critical personalization–privacy paradox on which modern developments are based. It identifies three priority research directions: generative AI in creative marketing, consumer trust in the personalization–privacy paradox, and organizational adaptation to fast innovation cycles. This study provides scholars with a comprehensive knowledge map, practitioners with strategic imperatives for responsible AI adoption, and policymakers with evidence that well-designed regulation accelerates innovation by balancing commercial value with societal concerns. Full article
Show Figures

Figure 1

22 pages, 1909 KB  
Article
Intelligent Question-Answering System for New Energy Vehicles Integrating Deep Semantic Parsing and Knowledge Graphs
by Yaqi Wu, Pengcheng Li, Tong Geng, Yi Wang, Haiyu Zhang and Shixiong Li
Informatics 2026, 13(5), 66; https://doi.org/10.3390/informatics13050066 - 24 Apr 2026
Viewed by 1681
Abstract
The new energy vehicle (NEV) industry generates massive multi-source heterogeneous data. To overcome traditional database limitations in terminology disambiguation and multi-hop reasoning, this paper proposes a knowledge graph (KG)-based question-answering (QA) architecture. Three primary domain challenges are addressed: First, to tackle the poor [...] Read more.
The new energy vehicle (NEV) industry generates massive multi-source heterogeneous data. To overcome traditional database limitations in terminology disambiguation and multi-hop reasoning, this paper proposes a knowledge graph (KG)-based question-answering (QA) architecture. Three primary domain challenges are addressed: First, to tackle the poor semantic extraction of informal diagnostic texts, a deep semantic parsing network (BERT-BiLSTM-CRF) is integrated to extract high-precision knowledge from 150,000 real-world maintenance records. Second, to solve topological redundancy, the Labeled Property Graph (LPG) specification is employed to encapsulate parameters of 2157 vehicle models as internal attributes, significantly streamlining complex multi-hop reasoning. Finally, to enhance limited reasoning capabilities, an intent classification module (TextCNN) automatically translates natural language into graph queries, enabling deep fault tracing across up to five semantic levels. Experimental results demonstrate 98% and 93% accuracy in entity-relation recognition and intent classification, respectively. The resulting KG (8274 nodes, 14,488 edges) establishes a scalable paradigm for intelligent diagnostic reasoning in complex vertical domains. Full article
(This article belongs to the Section Machine Learning)
Show Figures

Figure 1

20 pages, 4455 KB  
Article
The Relevance of Compound Events in Bee Traffic Monitoring
by Andrea Nieves-Rivera, Marie Lluberes-Contreras and Rémi Mégret
Informatics 2026, 13(5), 65; https://doi.org/10.3390/informatics13050065 - 23 Apr 2026
Viewed by 2115
Abstract
Bees are essential pollinators for agricultural systems, making accurate, automated monitoring of their behavior critical for assessing colony health and ecosystem stability. Recent advances in computer vision and artificial intelligence have enabled large-scale bee traffic monitoring at hive entrances; however, most existing event [...] Read more.
Bees are essential pollinators for agricultural systems, making accurate, automated monitoring of their behavior critical for assessing colony health and ecosystem stability. Recent advances in computer vision and artificial intelligence have enabled large-scale bee traffic monitoring at hive entrances; however, most existing event classification methods focus exclusively on simple entrance and exit events. This simplification overlooks compound movements—such as U-turns and guarding behaviors—that represent a substantial portion of bee activity and can lead to inaccurate trajectory reconstruction and misleading behavioral interpretations. In this work, we systematically analyze existing event classification strategies used in automatic bee traffic monitoring, evaluating their performance on both simple and compound movements. We then propose extended classification methods that explicitly model compound events by incorporating bidirectional movement patterns derived from positional and angular cues. Using a manually annotated dataset of computer-vision-based hive entrance recordings, we compare threshold-based, displacement-based, and angle-based approaches under simple and mixed-event conditions. Our results demonstrate that compound events account for over one-third of all detected movements and that classification methods explicitly designed to handle bidirectional behavior substantially outperform traditional approaches in both accuracy and robustness. In particular, threshold-based bidirectional classification achieves near-perfect performance when full trajectories are available, while displacement-based methods provide a reliable alternative under partial observations. These findings highlight the importance of modeling compound behaviors in automated bee monitoring systems and contribute to more accurate flight reconstruction, behavioral analysis, and AI-driven decision support for precision agriculture and pollinator management. Full article
Show Figures

Figure 1

Previous Issue
Next Issue
Back to TopTop