Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (571)

Search Parameters:
Keywords = open set classification

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
19 pages, 350 KB  
Article
A Hybrid Intelligent Decision Support Method for Abnormal Situation Management
by Rasul A. Kochkarov, Sergey V. Matseevich, Aleksandr V. Timoshenko and Aleksandr S. Zakharov
Big Data Cogn. Comput. 2026, 10(9), 311; https://doi.org/10.3390/bdcc10090311 - 11 Sep 2026
Viewed by 169
Abstract
Nowadays, the volume of heterogeneous data in situational analysis centers is growing exponentially, leading to information overload for decision makers (DMs) and a decrease in the effectiveness of traditional decision support systems (DSS). Intelligent DSSs (ISDSS) demonstrate potential, but face challenges in explainability, [...] Read more.
Nowadays, the volume of heterogeneous data in situational analysis centers is growing exponentially, leading to information overload for decision makers (DMs) and a decrease in the effectiveness of traditional decision support systems (DSS). Intelligent DSSs (ISDSS) demonstrate potential, but face challenges in explainability, heterogeneous data integration, cognitive load, and scalability. This paper proposes a method for intelligent decision support focused on identifying and generating options for resolving emergency situations—conditions that have no exact precedents in the knowledge base. The method includes formalizing the situation using a vector of normalized parameters St, separating it into independent and dependent variables with the construction of a dependency tree, neural network classification of three types of conditions (normal, abnormal, and emergency) with a forecast for a lead interval τ, the synthesis of solutions for emergency situations based on an analysis of proximity graphs to known emergency precedents and evolutionary optimization. A computational experiment was conducted on the open dataset of the Tennessee Eastman Process simulation model with 28 failure types and 200 repeated simulations. The neural network classifier achieved an accuracy of 0.88 and a macro-averaged F1-score of 0.87 on a test set of 200 situations. Graphs of nearby emergency precedents were constructed for 50 synthetic emergency situations; analysis demonstrated the stability of topological characteristics (vertex degree 5.62 ± 1.18, closeness centrality 0.43 ± 0.09), substantiating the applicability of graph neural networks for accelerated control action synthesis. The proposed method reduces dependence on expert assessments and improves the adaptability and explainability of decisions, while the demonstrated stability of the graph-based precedent retrieval lays the groundwork for future full-scale validation of control-action synthesis in next-generation hybrid IDSS. Full article
(This article belongs to the Section Cognitive System)
28 pages, 1332 KB  
Systematic Review
Decarbonisation Strategies in the Olive Oil Supply Chain: A Systematic Literature Review and ESG-Oriented Framework
by Emrah Karapinar, Roberto Leonardo Rana, Leonardo Orsitto, Mariarosaria Lombardi and Christian Bux
Sustainability 2026, 18(18), 9322; https://doi.org/10.3390/su18189322 - 10 Sep 2026
Viewed by 190
Abstract
Sustainability policies introduced under the European Green Deal have strengthened climate-related disclosure requirements for agri-food companies. In particular, the Corporate Sustainability Reporting Directive requires in-scope companies to transparently disclose information on their environmental performance. However, the academic literature on decarbonisation in the olive [...] Read more.
Sustainability policies introduced under the European Green Deal have strengthened climate-related disclosure requirements for agri-food companies. In particular, the Corporate Sustainability Reporting Directive requires in-scope companies to transparently disclose information on their environmental performance. However, the academic literature on decarbonisation in the olive oil sector remains fragmented. This systematic literature review synthesises findings by considering cultivation, milling and retail, and waste management as interconnected stages of the olive oil supply chain and by developing a matrix linking decarbonisation strategies to the relevant European Sustainability Reporting Standards (ESRS) environmental, social and governance (ESG) topics. Following the PRISMA protocol, 42 peer-reviewed studies from Scopus and Web of Science were included in the final synthesis, covering cultivation (RQ1), milling and retail (RQ2), and waste management (RQ3). The cultivation stage represents an important part of the emission profile of the chain while also offering potential for carbon sequestration through sustainable management practices, such as reduced tillage, cover crops, organic amendments and biochar application. In the downstream stages, the mill and its retail interface rely on a different set of measures, including two-phase extraction, rooftop photovoltaic systems, thermal recovery from pits, and lighter bottles transported in bulk. Waste management also offers opportunities to recover value from pomace, mill wastewater and pruning waste through biogas, biochar, compost or phenolic extracts. The potential for a net-negative carbon balance is context-dependent and varies with system boundaries, the balancing period, functional units, and the methods used to account for carbon sequestration. The matrix offers a clear classification of decarbonisation strategies and ESRS topics, opening valuable avenues for upcoming studies to extend its practical utility. Full article
23 pages, 1905 KB  
Article
Semantic-Driven Adversarial Reconstruction Learning for Open-Set Recognition in Remote Sensing Imagery
by Xing Zhang, Junlin Zhang, Jin Li and Mingqian Liu
Remote Sens. 2026, 18(18), 3101; https://doi.org/10.3390/rs18183101 - 10 Sep 2026
Viewed by 206
Abstract
Open-Set Recognition (OSR) in Remote Sensing Scene Images (RSSIs) is severely hindered by complex backgrounds, which obscure the generative failures of unknown classes in traditional reconstruction-based methods. To address this, we propose a Semantic-Driven Adversarial Reconstruction (SDAR) framework that shifts the OSR paradigm [...] Read more.
Open-Set Recognition (OSR) in Remote Sensing Scene Images (RSSIs) is severely hindered by complex backgrounds, which obscure the generative failures of unknown classes in traditional reconstruction-based methods. To address this, we propose a Semantic-Driven Adversarial Reconstruction (SDAR) framework that shifts the OSR paradigm to targeted semantic verification. A two-stage training protocol is utilized: the encoder is first trained via classification to extract high-level semantic features, after which the decoder is unfrozen for joint training to achieve semantic-level reconstruction. Driven by these semantics, a Class Activation Map (CAM) is used to dynamically mask only the core semantic elements, and a background-agnostic loss forces the model to exclusively reconstruct these critical parts rather than complex backgrounds. Consequently, the model solely masters the generative essence of known semantics; when presented with unknown categories lacking these identified semantic rules, it inevitably fails to reconstruct their masked cores, yielding a highly discriminative metric for OSR. Full article
(This article belongs to the Section AI Remote Sensing)
Show Figures

Figure 1

20 pages, 2001 KB  
Article
Algorithmic Diffusion on YouTube: A Machine Learning Analysis of Channel-Level Information Spread and Its Cross-Platform Generalisability
by Dana Tyulemissova, Aigul Shaikhanova, Oleksandr Kuznetsov, Aigerim Sambetova, Kainizhamal Iklassova and Aisanim Sarsenbayeva
Mach. Learn. Knowl. Extr. 2026, 8(9), 272; https://doi.org/10.3390/make8090272 - 6 Sep 2026
Viewed by 178
Abstract
(1) Background: Information diffusion models developed for graph-based platforms such as Reddit and broadcast architectures such as Telegram identify temporal features—particularly the timing of peak spread—as dominant predictors of coverage. Whether these predictors generalise to platforms where content is distributed through algorithmic recommendation [...] Read more.
(1) Background: Information diffusion models developed for graph-based platforms such as Reddit and broadcast architectures such as Telegram identify temporal features—particularly the timing of peak spread—as dominant predictors of coverage. Whether these predictors generalise to platforms where content is distributed through algorithmic recommendation rather than social-graph contagion remains an open question. (2) Methods: We analyse the YouNiverse dataset, comprising 133,364 English-language YouTube channels observed weekly from January 2015 to September 2019 (18.9 million observations). We derive channel-level diffusion features—including time-to-peak, post-peak decay rate, diffusion volatility, and upload frequency—and train three machine learning models (Linear Regression, Random Forest, and LightGBM) on two tasks: predicting peak weekly view growth (regression) and identifying viral channels (classification). A single-feature naive baseline (subscriber count alone) establishes the marginal contribution of the broader feature set beyond subscriber count alone, and a temporal split experiment (training on channels peaking before 2018, testing on 2018–2019) assesses cross-temporal stability. Because subscriber count and subscriber rank are measured at the October 2019 crawl, this is a retrospective characterisation rather than a strict real-time forecasting design. (3) Results: LightGBM achieves R2=0.776 (5-fold CV: 0.778±0.003) compared with R2=0.548 for the naive baseline, a net gain of +0.228R2. Because subscriber rank and subscriber count are near-perfectly collinear, we interpret them jointly as a channel-size dimension (42.2% of total mean absolute SHAP attribution), rather than as independent effects. Time-to-peak ranks fourteenth (1.1%), in contrast to its dominant role on Reddit (r=0.995, rank #1). For virality classification, LightGBM achieves ROC-AUC =0.967. Under the temporal split, Random Forest (R2=0.703) outperforms LightGBM (R2=0.683), showing greater cross-temporal stability within this retrospective split. (4) Conclusions: Within the 2015–2019 data, the results are consistent with algorithmic recommendation weakening the relationship between temporal diffusion dynamics and coverage magnitude at the channel level. Time-to-peak is weakly informative in this setting, while generalisation to the current recommendation system requires validation on newer data. Full article
(This article belongs to the Section Learning)
Show Figures

Figure 1

37 pages, 9215 KB  
Article
A Hybrid SMOTE-CTGAN and VAE-LSTM Framework for Interpretable Intrusion Detection in Imbalanced Network Traffic
by Felicia Maake, Justice Nkoana, Vekani Reviet Baloyi and Sello Mokwena
Big Data Cogn. Comput. 2026, 10(9), 304; https://doi.org/10.3390/bdcc10090304 - 5 Sep 2026
Viewed by 268
Abstract
The increasing sophistication of cyber threats and severe class imbalance in network traffic continue to challenge traditional intrusion detection systems. This study proposes a hybrid framework that integrates SMOTE and CTGAN for minority-class augmentation, a Bidirectional Long Short-Term Memory (Bi-LSTM) network for supervised [...] Read more.
The increasing sophistication of cyber threats and severe class imbalance in network traffic continue to challenge traditional intrusion detection systems. This study proposes a hybrid framework that integrates SMOTE and CTGAN for minority-class augmentation, a Bidirectional Long Short-Term Memory (Bi-LSTM) network for supervised traffic classification, and a benign-trained Variational Autoencoder (VAE) for validating low-confidence predictions. The framework was evaluated on the CSE-CIC-IDS2018 dataset. On a final holdout test partition of 2,759,227 network flows, it achieved a binary accuracy of 98.51%, an F1-score of 92.59%, and a ROC-AUC of 0.9935. In the 15-class evaluation, it achieved 98.50% accuracy and a weighted F1-score of 98.42%, while stratified 10-fold cross-validation yielded a mean accuracy of 98.79%. The VAE was activated for 519,389 low-confidence flows, representing 18.82% of the complete holdout test partition, and primarily reduced false-positive predictions, although this improvement was accompanied by a measurable reduction in attack recall. SHAP analysis was applied to the supervised Bi-LSTM component to provide feature-level interpretability. The framework also achieved a mean inference latency of 0.5195 ms per network flow. These findings demonstrate strong aggregate detection performance, stable generalisation, effective false-alarm reduction, and low inference latency, while highlighting continuing challenges in rare-class and open-set detection. Full article
Show Figures

Figure 1

19 pages, 457 KB  
Article
Beyond the Belief–Practice Gap: Competing Models of Pedagogical Justice in STEM Education
by Eduarda Ferreira and Maria João Silva
Societies 2026, 16(9), 281; https://doi.org/10.3390/soc16090281 - 3 Sep 2026
Viewed by 245
Abstract
Social inequalities often persist despite widespread endorsement of universal norms such as fairness and equality. One explanation is that actors interpret these principles differently in practice, generating divergent forms of action within the same institutional setting. This study examines this process in the [...] Read more.
Social inequalities often persist despite widespread endorsement of universal norms such as fairness and equality. One explanation is that actors interpret these principles differently in practice, generating divergent forms of action within the same institutional setting. This study examines this process in the context of gender equity in STEM education by analysing how teachers organise beliefs, awareness, pedagogical practices, responsibility, and the perceived value of gender equality when reasoning about pedagogical fairness. Using an exploratory case study with a theory-guided, person-centred typological classification approach, questionnaire data from teachers in a Portuguese school cluster (Grades 1–12) were analysed through quantitative classification and interpretive analysis of open-ended responses. Rather than differing only in degree of commitment to equality, teachers’ reported reasoning was organised into three theory-guided orientations: Procedural Impartiality (fairness as identical treatment), Compensatory Intervention (fairness as corrective action), and Reflexive Professional Responsibility (fairness as situated professional judgement). The resulting profiles showed differentiated configurations across the five analytical dimensions, particularly in the relative organisation of awareness, pedagogical practices, responsibility, and the perceived value of gender equality. Descriptive subgroup distributions also suggested tentative variation across disciplinary field, sex, and teaching experience, although these patterns should not be interpreted as systematic or generalisable associations. The findings indicate that teachers’ reported reasoning can be organised into distinct theory-guided orientations and suggest that shared commitments to gender equality may coexist with different professional interpretations of fairness and legitimate pedagogical intervention. The study therefore proposes a framework for examining how equity is interpreted as a professional category in educational contexts, while recognising that further research is needed to examine how such orientations relate to observed classroom practices and student outcomes. Full article
(This article belongs to the Section Science, Technology, and Society)
Show Figures

Figure 1

28 pages, 625 KB  
Article
Zero-LUT Open-Set Plankton Classification on Edge MPSoC and GPU Platforms
by David Sosa-Trejo, Martín González-García, Antonio Bandera and Santiago Hernández-León
Electronics 2026, 15(17), 3933; https://doi.org/10.3390/electronics15173933 - 1 Sep 2026
Viewed by 291
Abstract
Autonomous plankton monitoring needs classifiers that run on low-power edge hardware and reject out-of-distribution (OOD) inputs. Our prior in situ plankton image biomass estimation (IsPlanktonBIO) framework identifies organisms by nearest-neighbour search over a similarity gallery, a memory-resident lookup that stalls 8-bit integer (INT8) [...] Read more.
Autonomous plankton monitoring needs classifiers that run on low-power edge hardware and reject out-of-distribution (OOD) inputs. Our prior in situ plankton image biomass estimation (IsPlanktonBIO) framework identifies organisms by nearest-neighbour search over a similarity gallery, a memory-resident lookup that stalls 8-bit integer (INT8) acceleration and grows with the corpus. We replace it with a zero look-up-table (Zero-LUT) linear energy head: a matrix-multiply-only scorer over a supervised-contrastive backbone that needs no database and shrinks the footprint sevenfold, to 1.3 MB. We assess recognition and rejection jointly with an open-set protocol based on the correct-classification rate at a fixed OOD-leakage budget, showing that closed-set accuracy alone misleads. Our hardware–software co-design on an AMD Kria KV260 Multiprocessor System-on-Chip (MPSoC) runs the INT8 pipeline in real time at 86.5% end-to-end accuracy; its calibrated gate accepts 82.9% of in-distribution images, rejects 89.0% of OOD inputs, and correctly labels 97.6% of those accepted. INT8 conversion raises two issues: classification varies for background-dominated patterns, and OOD thresholds must be recalibrated on the compressed model, not inherited from full precision. An NVIDIA Jetson Orin Nano Super port reproduces this behaviour: the GPU is faster and more energy-efficient, whereas the MPSoC’s reconfigurable logic can host sensing and control beside the accelerator. Full article
(This article belongs to the Special Issue Advanced Techniques in Real-Time Image Processing)
Show Figures

Figure 1

23 pages, 620 KB  
Article
Automated Writer and Acquisition-Condition Classification of Digitally Captured Handwriting Using Statistical Dynamic Features and Support Vector Machines
by Long-Huang Tsai, Hsiang-Ju Lai, Wen-Chao Yang, Jiajun Jiang and Chung-Hao Chen
Appl. Sci. 2026, 16(17), 8696; https://doi.org/10.3390/app16178696 - 1 Sep 2026
Viewed by 255
Abstract
Digitally captured handwriting preserves pen trajectories and dynamic signals, but it also records hardware- and input-dependent properties that can confound forensic interpretation. This study revises a support vector machine (SVM) screening framework using 16,500 samples from 30 writers, 11 writing-content categories, and five [...] Read more.
Digitally captured handwriting preserves pen trajectories and dynamic signals, but it also records hardware- and input-dependent properties that can confound forensic interpretation. This study revises a support vector machine (SVM) screening framework using 16,500 samples from 30 writers, 11 writing-content categories, and five acquisition conditions spanning three tablets and stylus or finger input. Twenty-four raw and derived time-series variables were summarized by maximum, minimum, mean, median, and standard deviation, yielding 120 features; the mode statistic was removed. Writing direction and angular velocity were recalculated with atan2-based vector formulas. Unavailable device/API channels were encoded as zero, and Z-score parameters were estimated only from training folds. Writer and content evaluations used rotating pooled-“other” categories as rejection-class proxies, whereas acquisition-condition classification remained closed-set. Every outer five-fold split contained an inner five-fold forward-selection loop; RBF-SVM hyperparameters were fixed a priori (C = 1.0, gamma = scale, balanced class weights, and random seed 42). Writer classification achieved 94.85% accuracy (descriptive 95% CI: 94.69–95.01%) and 86.44% pooled-other recall. Content classification achieved 96.64% accuracy (95% CI: 95.74–97.53%) and 98.74% pooled-other recall. Acquisition-condition classification achieved 99.99% accuracy (99.98–100.00%), with one error among 16,500 outer-test predictions. The acquisition result is interpreted primarily as evidence that channel availability and device-specific measurement scales are strongly encoded in the feature space. Because the folds were sample-level, the samples were collected contemporaneously, and pooled-other writers were represented during training, these results do not establish session-disjoint, cross-device writer, or strict open-set generalization. The proposed workflow should therefore be regarded as an experimental triage aid that supports, rather than replaces, examiner-led comparison. Full article
Show Figures

Figure 1

17 pages, 8033 KB  
Article
Reproducible Semi-Automated Quantification of Vascularization in Bone Sections Using CD31 Immunohistochemistry and Trainable Weka Segmentation
by Nick Mattern, Holger Freischmidt, Matthias Schulte, Alma Aubert, Sanja Kalmus, Jan Makogon, Paul Alfred Grützner, Jonas Armbruster and Felix Lamadé-Dootz
J. Imaging 2026, 12(9), 409; https://doi.org/10.3390/jimaging12090409 - 1 Sep 2026
Viewed by 183
Abstract
Quantitative assessment of vascularization is important in bone regeneration research, but CD31-immunohistochemically stained sections are often evaluated manually or semi-quantitatively, limiting reproducibility and comparability. The aim of this study was to establish and validate a reproducible, open-source workflow for semi-automated quantification of CD31-positive [...] Read more.
Quantitative assessment of vascularization is important in bone regeneration research, but CD31-immunohistochemically stained sections are often evaluated manually or semi-quantitatively, limiting reproducibility and comparability. The aim of this study was to establish and validate a reproducible, open-source workflow for semi-automated quantification of CD31-positive area fraction in bone sections using Fiji/ImageJ and Trainable Weka Segmentation (TWS). CD31-immunohistochemically stained rat bone sections from defect/regenerating tissue, femur, tibia, and spine were analyzed. The workflow combined standardized image acquisition, predefined regions of interest, pixel-based TWS classification, extraction of the CD31-positive class, and CD31-positive area normalized to tissue area (CD31.Ar/T.Ar). Manual reference measurements were performed by two independent observers in repeated runs. Manual CD31.Ar/T.Ar measurements showed good retest reliability, with mean coefficients of variation (CV) of 7.17% and 6.84% for observer 1 and observer 2, respectively, and good interobserver agreement. Independently trained TWS classifiers produced highly stable CD31.Ar/T.Ar values, with an overall mean CV of 1.98%. Manual assessment required 3:17 ± 1:35 min per section, whereas the TWS-based workflow separated an initial classifier training step from rapid repeated analysis of larger image sets. This study provides a transparent, reproducible, and time-efficient open-source workflow for semi-automated quantification of CD31-positive vascular area fraction in bone sections and supports its use as a scalable method for vascular histomorphometry in preclinical bone regeneration research. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Figure 1

20 pages, 30932 KB  
Article
PASM-Net: A Coordinate-Aware Manifold-Regularized Framework for Open-Set UAV RF Fingerprint Identification
by Mingjun Jiang, Zhongqiang Luo and Jia Yuan
Big Data Cogn. Comput. 2026, 10(9), 292; https://doi.org/10.3390/bdcc10090292 - 31 Aug 2026
Viewed by 215
Abstract
To address the challenges of classifying known categories and rejecting unknown categories in the radio-frequency fingerprinting of uncooperative UAVs in open, low-altitude environments, this paper proposes PASM-Net, an open-set recognition method based on coordinate awareness and manifold continuity regularization. Existing deep learning-based RFFI [...] Read more.
To address the challenges of classifying known categories and rejecting unknown categories in the radio-frequency fingerprinting of uncooperative UAVs in open, low-altitude environments, this paper proposes PASM-Net, an open-set recognition method based on coordinate awareness and manifold continuity regularization. Existing deep learning-based RFFI methods still suffer from two shortcomings in open scenarios: First, while conventional convolutions and global pooling yield compact representations, they may weaken position-related information in the time–frequency spectrum, thereby affecting the model’s ability to characterize frequency-hopping trajectories and local time–frequency structures; on the other hand, decision mechanisms that rely solely on Softmax classifiers or simple distance metrics are susceptible to signal amplitude fluctuations and intra-class distribution dispersion, leading to the misclassification of unknown samples into known classes. To address these issues, this paper designs the PASM-Net feature-learning framework. First, the model employs multiscale hollow convolutions to extract local time–frequency features across different receptive fields and introduces a coordinate attention mechanism (Coordinate Attention, CoordAtt) prior to global pooling, thereby enhancing the model’s ability to represent position-related information in both the time and frequency domains. Second, the model introduces manifold continuity regularization (MCR) in the fused semantic space. By constraining feature variations within local neighborhoods via a Laplacian regularization term, MCR reduces the dispersion of known-class feature distributions. Finally, the model employs ArcFace to enhance the angular separability among known classes and performs open-set classification based on the cosine distance between test samples and the centers of known classes. The experimental results across six open-set scenarios on the DroneRFb-Spectra dataset showed that PASM-Net achieved an average true unknown rate (TUR) of 99.48%, an unknown accuracy of 97.9%, and an unknown-class precision (UP) of 87.3%, while maintaining a high performance in known-class recognition. Under the current experimental setup, this effectively improves the trade-off between known-class classification and unknown-class rejection. Full article
Show Figures

Figure 1

26 pages, 1647 KB  
Article
Noise-Dependent Robustness of XGBoost, LightGBM, and CatBoost
by Arun Morampudi, Praveena Padi, Pradeep Kumar Dolabehera Kakitapelli and Rakesh Kumar Surapani
Algorithms 2026, 19(9), 730; https://doi.org/10.3390/a19090730 - 30 Aug 2026
Viewed by 469
Abstract
Gradient-boosted decision trees (GBDTs) are among the leading methods for tabular data; however, their comparative robustness to corrupted labels remains unclear. This study resolves that uncertainty with a controlled benchmark protocol in which library identity is the only free variable, so that an [...] Read more.
Gradient-boosted decision trees (GBDTs) are among the leading methods for tabular data; however, their comparative robustness to corrupted labels remains unclear. This study resolves that uncertainty with a controlled benchmark protocol in which library identity is the only free variable, so that an observed difference can be attributed to the implementation rather than to tuning, data splits, or noise realizations, and in which capacity, noise floor, and dataset-selection confound checks are mandatory before any ranking is reported. We benchmarked Extreme Gradient Boosting (XGBoost), Light Gradient-Boosting Machine (LightGBM), and Categorical Boosting (CatBoost) under symmetric, asymmetric pair-flip, and instance-dependent label noise across 15 datasets from the OpenML Curated Classification benchmark suite 2018 (OpenML-CC18), four noise rates, and three random seeds. The clean data-tuned configurations were fixed across the noise conditions. Predictive performance is the macro-averaged F1 score (macro-F1) on a clean test set, and degradation is the absolute drop in macro-F1 relative to each library’s own clean-data score on the same dataset, so the recoveries reported below are absolute percentage points of macro-F1. The ranking depends on both the noise rate and the noise model: no library separated at a 10% rate under symmetric or asymmetric noise; separation emerged from 20% under symmetric noise and only at 40% under asymmetric noise; and under instance-dependent noise, it weakened as the rate rose. At 40% symmetric and asymmetric noise, CatBoost showed significantly lower per-dataset degradation than LightGBM (Friedman tests, both p<0.001), with mean ranks of 1.20 versus 2.47 and 1.20 versus 2.67, respectively, and with XGBoost intermediate in both cases (2.33 and 2.13) and separable from CatBoost under symmetric noise only. At 40% instance-dependent noise, the ranking disappeared: no significant library ranking was detected (p=0.63), and performance gaps were smaller than twice the pooled seed standard deviation, which shows that GBDT robustness conclusions are noise-model-dependent. Rankings also varied by evaluation dimension: LightGBM had the most stable feature importances in observed means, calibration rankings depended on the noise model, and CatBoost’s training loss was the strongest mislabel-detection signal. Under symmetric noise only, small-loss reweighting significantly improved all three libraries, while early stopping recovered up to 8.5 pp for LightGBM; mitigations were not evaluated under asymmetric or instance-dependent noise. For practitioners, this means the library and the mitigation should be chosen together with the noise process that is expected: prefer CatBoost when label-conditional noise is likely and accuracy is the objective; expect no library to buy robustness under feature-dependent noise at high rates; apply early stopping or small-loss reweighting under symmetric noise, where both give significant gains and early stopping helps LightGBM most; and select on calibration, mislabel detectability, or importance stability instead when the deployment depends on those, because the ranking differs by dimension. Full article
Show Figures

Figure 1

29 pages, 3038 KB  
Article
Character-Based Arabic Offline Handwritten Text Recognition Using Faster R-CNN
by Sofiane Medjram and Ruwaidah Saud Alnejaidi
Appl. Sci. 2026, 16(17), 8550; https://doi.org/10.3390/app16178550 - 27 Aug 2026
Viewed by 208
Abstract
Offline handwritten word recognition has progressed from whole-word classification to sequence transcription, yet many systems depend on large annotated corpora and exploit lexical regularities over explicit character evidence. This paper presents an alternative formulation for Arabic offline handwritten word recognition, treating characters as [...] Read more.
Offline handwritten word recognition has progressed from whole-word classification to sequence transcription, yet many systems depend on large annotated corpora and exploit lexical regularities over explicit character evidence. This paper presents an alternative formulation for Arabic offline handwritten word recognition, treating characters as spatial objects detected via a Faster Region-Based Convolutional Neural Network rather than symbols generated by a one-dimensional decoder. We construct and release a character-level annotated subset of 2153 handwritten word images from a standard Arabic benchmark, exporting matched detection, sequence, and word-class labels. We also introduce an open-source subword exchange toolkit that creates a controlled structural-generalization benchmark by swapping subwords while preserving handwriting style. Experiments compare the proposed detector against whole-word and sequence-based baselines on both the original held-out split and the perturbed benchmark. Results show sequence models degrade sharply under structural recombination, whereas the proposed detector remains stable, achieving a 26.56% character error rate and 70.0% word accuracy on the perturbed benchmark. These findings demonstrate that explicit character localization provides a robust, data-efficient alternative for Arabic handwritten text recognition in low-resource settings. Full article
(This article belongs to the Section Computing and Artificial Intelligence)
Show Figures

Figure 1

16 pages, 2666 KB  
Article
Systematic Benchmarking of a Dry Electrode EEG Prototype Against Wet Electrode EEG Systems in Electrophysiological/Cognitive Scenarios
by Boli Pan, Shuo Ding, Yingbo Geng, Jiacheng Liang, Fali Li, Gang Wang, Xirong Li, Yanbin Dong and Rihui Li
Biosensors 2026, 16(9), 467; https://doi.org/10.3390/bios16090467 - 27 Aug 2026
Viewed by 346
Abstract
Objective: Recent advances in dry electrode EEG have enabled rapid setup and recording in unconventional scenarios. However, past developments were primarily driven by brain–computer interfaces (BCI), leaving their comparability to wet electrodes in clinical and daily life applications an open question. Here, we [...] Read more.
Objective: Recent advances in dry electrode EEG have enabled rapid setup and recording in unconventional scenarios. However, past developments were primarily driven by brain–computer interfaces (BCI), leaving their comparability to wet electrodes in clinical and daily life applications an open question. Here, we developed a new dry EEG system and systematically benchmarked its performance against a commercial wet EEG system across various tasks. Methods: Participants (n = 19) underwent simultaneous recording using both devices. We first collected resting-state EEG under both eyes-closed and eyes-open conditions, followed by a steady-state visual evoked potential (SSVEP) task at different flicker frequencies and a motor imagery (MI) task. System performance was evaluated using power spectral density (PSD), signal to noise ratio (SNR), event-related spectral perturbation (ERSP), and single-trial classification accuracy. Results: The two systems performed similarly across different tasks. During the resting state, no statistically significant differences were observed between the two systems in the PSD of the five frequency bands (p > 0.05 in all cases). Similarly, SNR in the SSVEP task showed no significant differences at 8 Hz, 10 Hz, and 12 Hz after correction. For cognitive tasks, classification accuracies were comparable (SSVEP: dry 80.08% ± 7.1% vs. wet 81.10% ± 6.5%; MI: dry 72.46% ± 3.89% vs. wet 70.7% ± 2.37%). Conclusions: The developed dry EEG system can effectively record electrophysiological measurements commonly employed in research and clinical settings, with quality comparable to that of traditional wet EEG systems. Full article
(This article belongs to the Special Issue Biosensors for Physiological Signal Monitoring)
Show Figures

Figure 1

34 pages, 1154 KB  
Article
Metabodeconplus—An R Package for Automated Deconvolution and Alignment of 1D NMR Metabolomics Data
by Tobias Schmidt, Maximilian Sombke, Helena U. Zacharias, Peter J. Oefner, Rainer Spang and Wolfram Gronwald
Metabolites 2026, 16(9), 604; https://doi.org/10.3390/metabo16090604 - 24 Aug 2026
Viewed by 376
Abstract
Background: In one-dimensional NMR spectra of complex biofluids such as urine and plasma, extensive signal overlap obscures individual metabolite signals. Resolving this overlap by deconvolution is only the first step: turning a set of spectra into a table for subsequent statistical analysis also [...] Read more.
Background: In one-dimensional NMR spectra of complex biofluids such as urine and plasma, extensive signal overlap obscures individual metabolite signals. Resolving this overlap by deconvolution is only the first step: turning a set of spectra into a table for subsequent statistical analysis also requires the alignment of signals across samples and their integration into a single feature matrix. Methods: Here, metabodeconplus is presented, an R package that unifies this entire path into a single reproducible end-to-end workflow. From raw one-dimensional spectra, it deconvolutes overlapping signals, aligns resulting signals across samples, and integrates them into a data matrix for built-in sample classification or downstream statistical analysis. Automated parameter optimization removes manual tuning, and a Rust computational backend with parallelization leads to fast runtimes. Results: On the simulated Sim3 spectra, a combined score of correctly identified signals and reconstruction accuracy (maximum 1) rose from 0.712 for the predecessor package to 0.801 for metabodeconplus. For the urinary AKI dataset, metabodeconplus reached a classification accuracy of 73.7 ± 2.20% and an AUC=0.827±0.025, which is comparable to the binning baseline. An advantage is the potential unambiguous metabolite assignment of the deconvoluted signals. Conclusions: The package is freely available as open source on GitHub and on CRAN. Full article
(This article belongs to the Special Issue Advances in NMR-Based Metabolomics for Biomedical Research)
Show Figures

Figure 1

29 pages, 6297 KB  
Article
Do We Have an Agreement? A Comparative Analysis of the ESCOX Skill Extraction Tool with Expert-Labeled EU Labour Market Data
by Dimitrios Christos Kavargyris, Konstantinos Georgiou and Lefteris Angelis
Appl. Sci. 2026, 16(17), 8388; https://doi.org/10.3390/app16178388 - 23 Aug 2026
Viewed by 288
Abstract
Labour markets across Europe increasingly describe workers through skills rather than job titles, and a growing number of large language model (LLM)-based tools now claim to extract these skills automatically from unstructured text at scale. Among these, ESCOX has gained particular traction for [...] Read more.
Labour markets across Europe increasingly describe workers through skills rather than job titles, and a growing number of large language model (LLM)-based tools now claim to extract these skills automatically from unstructured text at scale. Among these, ESCOX has gained particular traction for its open-source, taxonomy-aligned design, yet like any LLM-based system it remains susceptible to hallucination, prompt sensitivity, and non-deterministic output, risks that are rarely quantified before such tools are deployed in practice. The European Skills, Competences, Qualifications, and Occupations (ESCO) classification provides the standardised reference against which this risk can be measured, but no study has yet benchmarked an ESCO-aligned LLM extractor against an independent, expert-labelled dataset at scale. This study addresses that gap. Candidate skills generated by ESCOX are compared against reference skills already assigned to job vacancies on the EURES portal by national labour-market experts, using job-by-skill matrices to quantify agreement and skill co-occurrence networks to characterise how the two sets diverge structurally. Results reveal the extent to which ESCOX’s automatic output aligns with expert judgement and where systematic divergences occur. These findings offer HR practitioners, policymakers, and labour-market researchers an evidence-based basis for deciding when ESCOX’s output can be trusted directly and when expert oversight remains necessary. Full article
(This article belongs to the Special Issue Application of Information Systems: Second Edition)
Show Figures

Figure 1

Back to TopTop