Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (3,331)

Search Parameters:
Keywords = deep ResNet-152

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
39 pages, 1615 KB  
Article
A Robust and Fair Multimodal Recommender System Under Structured Modality Missingness: The Trust-Based Evaluation Framework
by Musa Mbedzi and Thulane Paepae
Information 2026, 17(9), 873; https://doi.org/10.3390/info17090873 - 9 Sep 2026
Abstract
The growing complexity of digital real estate platforms demands intelligent recommendation systems (RS) capable of operating in data-sparse and heterogeneous environments. While transfer learning (TL) has proven effective in general RS, its application to real estate (RE) remains limited, particularly regarding the operationalization [...] Read more.
The growing complexity of digital real estate platforms demands intelligent recommendation systems (RS) capable of operating in data-sparse and heterogeneous environments. While transfer learning (TL) has proven effective in general RS, its application to real estate (RE) remains limited, particularly regarding the operationalization of multi-dimensional evaluation frameworks. This study addresses these gaps by developing a TL-based real estate recommender system (RERS) utilizing a pre-trained ResNet50 architecture, trained on a locally curated dataset from Gauteng, South Africa, providing rare, data-driven insights into a pivotal emerging market economy. By transitioning from traditional label-based retrieval to high-dimensional visual feature alignment, the model mitigates class imbalance and data redundancy in fragmented property markets. The framework is validated using the proposed Trust-based Evaluation (T-EVAL) methodology, demonstrating the efficacy of deep learning architectures in providing reliable and trustworthy property recommendations within emerging market economies. Full article
15 pages, 5864 KB  
Article
Detection of Abalone Freshness Based on Smart Phone Image and Deep Learning
by Yizhan Yu, Jialin Li, Junlong Lai, Zhihong Zheng, Haisheng Lin, Wenhong Cao, Jialong Gao and Xiaoyu Xia
Foods 2026, 15(18), 3189; https://doi.org/10.3390/foods15183189 - 9 Sep 2026
Abstract
Rapid and nondestructive freshness evaluation of abalone is important for quality control during cold-chain distribution, yet conventional chemical and microbiological methods are destructive and labor-intensive. In this study, a smartphone image-based deep learning strategy was developed for abalone freshness classification under refrigerated storage. [...] Read more.
Rapid and nondestructive freshness evaluation of abalone is important for quality control during cold-chain distribution, yet conventional chemical and microbiological methods are destructive and labor-intensive. In this study, a smartphone image-based deep learning strategy was developed for abalone freshness classification under refrigerated storage. Abalone samples stored at 4 °C were imaged daily under natural light, and freshness labels were assigned according to total volatile basic nitrogen (TVB-N) measurements. A total of 1867 images were used to develop binary classification models, and a transfer learning-based ResNet50 model was further interpreted using Gradient-weighted Class Activation Mapping (Grad-CAM). TVB-N increased progressively during storage and exceeded the spoilage threshold on day 5 (15.63 ± 0.43 mg/100 g), which was used to define fresh (days 1–4) and spoiled (days 5–7) classes. Among the evaluated architectures, ResNet50 achieved the best overall performance, with a validation accuracy of 0.9611, precision of 0.9649, recall of 0.9091, and F1-score of 0.9362. On the test set, the model correctly classified 522 fresh and 220 spoiled images, yielding an overall accuracy of 96.11%. Grad-CAM visualization showed that the model mainly focused on the abalone body and marginal contour, indicating that predictions were driven by intrinsic appearance changes rather than background interference. These results demonstrate that smartphone imaging combined with deep learning provides a rapid, low-cost, and nondestructive approach for abalone freshness assessment and has potential for digital quality monitoring in shellfish cold chains. Full article
(This article belongs to the Special Issue Advances in Analytical Techniques for Food Safety Assessment)
Show Figures

Figure 1

18 pages, 2935 KB  
Article
Generalization of Defense Effects Learned from a Single Adversarial Attack
by Dongxian Niu and Lin Shi
Computation 2026, 14(9), 211; https://doi.org/10.3390/computation14090211 - 9 Sep 2026
Abstract
Adversarial attacks misled deep neural networks by injecting perturbations into input images. Training networks with adversarial examples defended against adversarial attacks. However, training with specific adversarial examples only defended against the corresponding attacks. To generalize the defense effect from one specific attack to [...] Read more.
Adversarial attacks misled deep neural networks by injecting perturbations into input images. Training networks with adversarial examples defended against adversarial attacks. However, training with specific adversarial examples only defended against the corresponding attacks. To generalize the defense effect from one specific attack to other attacks, we proposed a method called Gradient Vicinity Adversarial Training (GVAT), which generated adversarial examples along directions sampled in the vicinity of the gradient. The defense effects of GVAT were evaluated using three attack methods: fast gradient sign method (FGSM), projected gradient descent (PGD), and Carlini–Wagner (CW) under the L2-norm constraint. A three-layer convolutional network was trained on the MNIST dataset, and two WideResNet-28-10 networks were trained on the CIFAR-10 and CIFAR-100 datasets respectively. Under the transfer-based black-box setting, the results showed that GVAT not only defended against the corresponding attacks that generated adversarial examples but also defended against other attacks. In other words, the defense effect of GVAT was generalized to other attacks under the transfer-based black-box setting. Full article
(This article belongs to the Special Issue Computational Methods for Multi-View Representation Learning)
Show Figures

Graphical abstract

27 pages, 2859 KB  
Article
Conditional Latent Diffusion for Synthetic Brain MRI in Alzheimer’s Disease: A Preprocessing-Focused Pipeline
by Soheil Fallah and Nitsa J. Herzog
J. Imaging 2026, 12(9), 426; https://doi.org/10.3390/jimaging12090426 - 9 Sep 2026
Abstract
Deep learning for Alzheimer’s disease (AD) detection from structural magnetic resonance imaging (MRI) needs large, labelled datasets, yet many cohorts hold only a few hundred participants, for which conventional augmentation adds little anatomical diversity. In a two-stage pipeline, a variational autoencoder compressed 256 [...] Read more.
Deep learning for Alzheimer’s disease (AD) detection from structural magnetic resonance imaging (MRI) needs large, labelled datasets, yet many cohorts hold only a few hundred participants, for which conventional augmentation adds little anatomical diversity. In a two-stage pipeline, a variational autoencoder compressed 256 × 256 coronal slices to a 32 × 32 × 8 latent space, and a class-conditional latent diffusion model under classifier-free guidance generated AD and cognitively normal (CN) images using 295 participants from the Alzheimer’s Disease Neuroimaging Initiative (ADNI). The pipeline reached a Kernel Inception Distance (KID) of 0.030 ± 0.002 and a bias-corrected Fréchet Inception Distance (FID) of 43.82. A controlled ablation varying preprocessing alone improved KID by 0.0147 and precision by 0.069, both with 95% intervals excluding zero. FID did not separate the configurations. A ResNet-18 trained only on synthetic slices and tested on 44 held-out real participants (18 AD, 26 CN), each scored as the mean probability over twenty slices, reached an area under the curve of 0.779 ± 0.031 against 0.869 ± 0.027 for real data; the difference was not distinguishable at this sample size. No instance memorisation was found among 880 samples, and a size-matched control exposed a 27.7-percentage-point inflation in the standard memorisation metric. Preprocessing, therefore, measurably affects synthesis quality at the small-cohort scale, though not on every measure. Full article
(This article belongs to the Section Medical Imaging)
Show Figures

Figure 1

28 pages, 6101 KB  
Article
Intelligent Visual Prioritization for Retinal Prostheses via Context-Aware Object Ranking and Depth-Aware Phosphene Generation
by Xinwei Li, Irshad Khalil, Faisal Rahman and Muhammad Nawaz Khan
Biomimetics 2026, 11(9), 649; https://doi.org/10.3390/biomimetics11090649 - 9 Sep 2026
Abstract
Images from high-resolution cameras are mapped onto a sparse pattern of low spatial resolution and intensity in the retina, which limits visual perception in retinal prosthetic vision. When the entire scene is converted into phosphenes, it may allow unnecessary background information to be [...] Read more.
Images from high-resolution cameras are mapped onto a sparse pattern of low spatial resolution and intensity in the retina, which limits visual perception in retinal prosthetic vision. When the entire scene is converted into phosphenes, it may allow unnecessary background information to be retained and may cause visual clutter, which may make it hard for prosthetic vision users to interpret the scene. In order to tackle this issue, this paper presents a context-, depth-, and user-preference-aware method for selecting the objects of interest in the generation of phosphene images. The proposed method does not show all the objects equally but learns to sort the objects according to their relevance to prosthetic vision. Manual annotation of a subset of COCO images was conducted where the most salient object was selected based on environment type, scene type, user mode, safety, navigation relevance, task importance, and distance. All of the candidate objects are described by full-scene visual features, object-crop features, handcrafted priority features, context embeddings, and monocular depth features. To predict object-level importance scores and identify the Top-1 and Top-4 important objects in unseen scenes, a hybrid deep learning model combining twin ResNet-18 backbones for scene and object feature extraction with embedding-based context encoding was trained. Priority maps and phosphene images were then created using the selected object masks and were depth-weighted. Two types of phosphene representations were also produced: Canny-edge-based and direct full images. The proposed framework is designed to suppress irrelevant background areas and improve important and closer objects in order to obtain a simplified and informative prosthetic-vision representation of the scene. The experimental evaluation, including Top-1 accuracy, Top-3 accuracy, mean reciprocal rank (MRR), and visual comparison, demonstrates the effectiveness of the proposed framework, achieving a Top-1 accuracy of 90.12%, a Top-3 accuracy of 97.45%, and an MRR of 0.9368. Furthermore, the proposed Canny-priority phosphene representation achieved an average human-participant recognition accuracy of approximately 86%. The proposed method offers a user-adaptive strategy for selecting and visualizing the information of a scene under the severe constraint of the bandwidth of retinal prosthetic vision. Full article
Show Figures

Graphical abstract

21 pages, 2740 KB  
Article
Diagnostic Accuracy of a Machine Learning Model for Cervical Vertebra-Based Skeletal Maturity Assessment in Pediatric and Adolescent Patients Using Cephalometric Radiographs
by Narmin Helal, Faisal Al-Malki, Basil Saadi and Osama Basri
Diagnostics 2026, 16(18), 2886; https://doi.org/10.3390/diagnostics16182886 - 8 Sep 2026
Abstract
Background/Objectives: Skeletal maturity assessment is essential for timing orthodontic growth-modification treatment. The cervical vertebral maturation (CVM) method is widely used, but manual staging is subjective and prone to inter-observer variability, and artificial intelligence (AI) may improve its consistency and accuracy. The aim of [...] Read more.
Background/Objectives: Skeletal maturity assessment is essential for timing orthodontic growth-modification treatment. The cervical vertebral maturation (CVM) method is widely used, but manual staging is subjective and prone to inter-observer variability, and artificial intelligence (AI) may improve its consistency and accuracy. The aim of this study was to develop and externally validate a fully automated deep-learning pipeline for CVM assessment and to determine its diagnostic accuracy for three-phase and six-stage classification against calibrated expert staging. Methods: A cross-sectional three-stage deep-learning pipeline was trained and internally cross-validated on 523 cephalometric radiographs from King Abdulaziz University. Images were CLAHE-enhanced; YOLOv8 detected the C2–C4 region; a ResNet-50 classifier assigned the growth phase (Early, Peak, Late); and phase-specific binary classifiers assigned CVM stages (CS1–CS6). External validation used the independent ‘Aariz dataset (n = 150). Results: The detector achieved mAP@0.5 = 0.9939 and mean IoU = 0.8884. On external validation, the three-phase classifier reached 96.0% accuracy, macro-F1 0.919, weighted κ 0.951 and ROC-AUC 0.998, whereas the end-to-end six-stage cascade reached 79.3% accuracy, macro-F1 0.735, weighted κ 0.910 and ROC-AUC 0.923. Most errors occurred between adjacent stages, with 98.7% of predictions within ±1 stage of the reference standard. Conclusions: The automated pipeline showed promising performance for three-phase CVM classification and moderate performance for six-stage classification on one independent external benchmark. Further multicenter prospective validation is required to establish its generalizability and clinical utility. Full article
(This article belongs to the Section Machine Learning and Artificial Intelligence in Diagnostics)
Show Figures

Figure 1

18 pages, 3708 KB  
Article
Learning Compact Multispectral Signatures for Geographical-Origin Authentication of Pinellia ternata via Correlation-Guided Deep Modeling
by Zhihui Fan, Shaowen Jing, Chao Ma, Sen Wang, Zhenzhen Chen, Jiayu Huang and Mingkun Zhang
Molecules 2026, 31(17), 3138; https://doi.org/10.3390/molecules31173138 - 7 Sep 2026
Abstract
Geographical authentication of medicinal plant materials remains challenging because multispectral variables are often highly collinear and sample grouping can complicate reliable model validation. Existing correlation-based feature-selection strategies also require careful adaptation to multiclass problems to avoid artificial ordering of class labels and information [...] Read more.
Geographical authentication of medicinal plant materials remains challenging because multispectral variables are often highly collinear and sample grouping can complicate reliable model validation. Existing correlation-based feature-selection strategies also require careful adaptation to multiclass problems to avoid artificial ordering of class labels and information leakage during model development. Therefore, this study aimed to develop a compact and leakage-controlled multispectral learning framework for geographical-origin discrimination. This study analyzed 800 physical Pinellia ternata samples from Gansu Xihe, Sichuan Neijiang, Sichuan Chengdu, and Chongqing Dianjiang (200 samples per origin). Each physical sample was represented by 31 mean grayscale intensities calculated from Otsu-segmented multispectral regions of interest. A Pearson-correlation-guided deep multilayer perceptron (PCG-DeepMLP) was constructed by estimating one-vs-rest band relevance and inter-band redundancy only within the training data. The key methodological innovation is a unified multiclass-aware, relevance–redundancy spectral-learning framework in which class-specific one-vs-rest Pearson relevance is coupled with inter-band redundancy control and embedded within leakage-controlled grouped model development. By learning the spectral subset exclusively from each training partition before nonlinear classification, the framework produces compact and complementary multispectral signatures while preserving multiclass structure and strict independence of held-out groups. Model and feature-selection settings were chosen by three-fold grouped cross-validation within each training partition. PCG-DeepMLP retained 9–21 bands and achieved the highest mean accuracy (0.9812 ± 0.0135), macro-F1 (0.9812 ± 0.0135), Matthews correlation coefficient (MCC; 0.9752 ± 0.0179), and macro-AUC (0.9994 ± 0.0006) among seven models. Its macro-F1 was higher than that of 1D-CNN, 1D-ResNet, full-band MLP, PLS-DA, and random forest after Holm correction. Performance was estimated through a strict nested group-wise internal validation scheme, with every outer test fold remaining isolated from feature selection, preprocessing, and model optimization. These findings demonstrate that multiclass-aware relevance–redundancy learning can retain complementary Pinellia ternata origin-discriminative information in a compact and stable spectral representation, enabling accurate geographical-origin authentication while providing a principled basis for reduced-channel acquisition and future independent multi-batch validation. Full article
(This article belongs to the Special Issue Analytical Methods for Safety and Quality Control of Functional Food)
Show Figures

Graphical abstract

25 pages, 2308 KB  
Article
Comparative Analysis of CNN and Transformer Models for Multi-Class Diabetic Retinopathy Grading Using Fundus Images
by Maha A. Thafar
Diagnostics 2026, 16(17), 2882; https://doi.org/10.3390/diagnostics16172882 - 7 Sep 2026
Abstract
Background/Objectives: Diabetic retinopathy is a major cause of preventable vision loss worldwide, making early and accurate disease grading crucial for timely treatment. Although convolutional neural network (CNN)- and transformer-based architectures have demonstrated promising performance for retinal image analysis, comprehensive comparisons within a [...] Read more.
Background/Objectives: Diabetic retinopathy is a major cause of preventable vision loss worldwide, making early and accurate disease grading crucial for timely treatment. Although convolutional neural network (CNN)- and transformer-based architectures have demonstrated promising performance for retinal image analysis, comprehensive comparisons within a unified experimental framework remain limited. This study systematically compares representative standard and lightweight CNN- and transformer-based architectures for multi-class DR grading. Methods: Six ImageNet-pretrained deep-learning models, including ResNet50, EfficientNet-B0, MobileNetV2, Vision Transformer (ViT), Swin-Tiny, and Swin Transformer, were evaluated on the APTOS 2019 retinal fundus image dataset under a unified experimental configuration with consistent preprocessing, data augmentation, training, and evaluation settings. All models were fine-tuned and evaluated independently over five runs with different random seeds. Their performance was assessed using accuracy, precision, recall, F1-score, area under the receiver operating characteristic curve (AUC), Quadratic Weighted Kappa (QWK), per-class analysis, computational efficiency, and statistical analysis. Results: Transformer-based models generally achieved higher mean classification performance than the evaluated CNN-based models. Swin-Tiny achieved the highest mean accuracy (82.3%), macro F1-score (64.4%), weighted F1-score (82.1%), and QWK (89.8%) across the five runs. Among the CNN-based models, EfficientNet-B0 achieved the strongest overall classification performance, whereas MobileNetV2 provided the lowest computational complexity. The results also highlighted differences in learning behavior and computational requirements across the evaluated architectures. Repeated experiments demonstrated stable performance across different random seeds, supporting the reliability of the proposed evaluation. Conclusions: Overall, this study provides a comprehensive comparison of representative CNN- and transformer-based architectures under consistent experimental settings and offers practical guidance for selecting suitable deep learning models for automated diabetic retinopathy screening. Full article
Show Figures

Figure 1

23 pages, 766 KB  
Article
Lightweight Shoulder Physiotherapy Exercise Recognition via Efficient Channel Attention and Depthwise Separable Residual Networks on Wrist-Worn IMU
by Sakorn Mekruksavanich and Anuchit Jitpattanakul
Computers 2026, 15(9), 595; https://doi.org/10.3390/computers15090595 - 7 Sep 2026
Abstract
Accurate and subject-independent recognition of shoulder physiotherapy exercises from wrist-worn inertial measurement unit (IMU) signals is essential for automated home-based rehabilitation monitoring, yet existing deep learning models are too parameter-intensive to deploy on resource-constrained smartwatch hardware. This paper presents ECA-ResNet1D-Lite, a lightweight one-dimensional [...] Read more.
Accurate and subject-independent recognition of shoulder physiotherapy exercises from wrist-worn inertial measurement unit (IMU) signals is essential for automated home-based rehabilitation monitoring, yet existing deep learning models are too parameter-intensive to deploy on resource-constrained smartwatch hardware. This paper presents ECA-ResNet1D-Lite, a lightweight one-dimensional depthwise separable residual network augmented with efficient channel attention (ECA), trained and evaluated on the SPARS9x dataset comprising six shoulder exercises recorded from 20 subjects using a commercial wrist-worn smartwatch at 50 Hz. Because 50% window overlap allows adjacent windows to share samples, we report three protocols—window-level five-fold, recording-level grouped five-fold, and leave-one-subject-out (LOSO) cross-validation—with model selection performed throughout on validation data disjoint from the test partition. Under LOSO, the primary protocol, the model attains 99.1 ± 1.0% accuracy while requiring only 13,612 parameters and 0.46 M multiply–accumulate operations per window—the highest accuracy and the lowest between-subject dispersion of the seven architectures trained under an identical protocol, ahead of the strongest unconstrained baseline (InceptionTime, 98.8 ± 1.4%) at 36.3× fewer parameters and 213× fewer operations, and ahead of the parameter-efficient designs TinyHAR (97.4 ± 2.5%) and TinierHAR (97.4 ± 2.1%). An ablation over four attention variants (SE, ECA, CBAM, and multi-head self-attention) shows ECA to be the cheapest, adding six parameters (0.04% overhead), while delivering the largest LOSO gain over the same backbone without attention (+0.2 percentage points) and reducing the cross-subject standard deviation from 1.4% to 1.0%. Per-class analysis further reveals that shoulder girdle stabilization is the most challenging exercise under LOSO (F1-score: 96.4%), despite being the most represented class, attributable to its quasi-static, low-amplitude IMU signature. Deployed to an Apple Watch Ultra 2, the model classifies a 4-s window in 0.24 ms (duty cycle 0.012%), establishing that subject-independent shoulder physiotherapy monitoring is computationally feasible on current smartwatch hardware. Full article
Show Figures

Figure 1

15 pages, 10306 KB  
Article
Non-Invasive Individual Re-Identification of Water Monitors (Varanus salvator) Using Deep Learning
by Chayatorn Thongsub, Chattraphas Pongcharoen, Warong Suksavate, Kornsorn Srikulnath and Prateep Duengkae
Diversity 2026, 18(9), 545; https://doi.org/10.3390/d18090545 - 7 Sep 2026
Abstract
Effective management of urban Asian water monitor (Varanus salvator (Laurenti, 1768)) populations requires precise individual identification, yet traditional physical-marking methods remain invasive and labor-intensive. This study developed a non-invasive, automated photographic re-identification (Re-ID) system using deep learning and computer vision to facilitate [...] Read more.
Effective management of urban Asian water monitor (Varanus salvator (Laurenti, 1768)) populations requires precise individual identification, yet traditional physical-marking methods remain invasive and labor-intensive. This study developed a non-invasive, automated photographic re-identification (Re-ID) system using deep learning and computer vision to facilitate population monitoring in semi-urban environments. We evaluated seven deep learning configurations based on ResNet50 incorporating Squeeze-and-Excitation (SE), Convolutional Block Attention Module (CBAM), and Batch Normalization neck (BNNeck) optimizations and benchmarked them against traditional feature matching (HotSpotter) using an open-set evaluation dataset of 3311 images across 161 Side-IDs focusing on unique lateral head-scale patterns. HotSpotter demonstrated immediate field viability, achieving a Rank-1 accuracy of 99.88% and a mean Average Precision (mAP) of 90.40%. Among the deep learning architectures, the baseline ResNet50 achieved the highest Rank-1 accuracy of 75.39% and mAP of 55.97%. As a decision-support framework, the deep learning pipeline achieved over 87% Rank-5 accuracy, drastically reducing manual screening effort and cognitive load during capture–mark–recapture surveys. This non-invasive framework establishes a scalable, welfare-friendly protocol for long-term urban wildlife management and biodiversity monitoring. Full article
(This article belongs to the Section Biodiversity Conservation)
Show Figures

Figure 1

15 pages, 5812 KB  
Article
Microfluidic Light-Scattering Imaging Coupled with Deep Learning for Label-Free Single-Cell Classification of Lymphoma Cells
by Linyan Xie, Mengfei Wang, Xijia Luo, Shuoxian Xia, Qiongqiong Ren and Xuezhi Zhou
Biosensors 2026, 16(9), 500; https://doi.org/10.3390/bios16090500 - 7 Sep 2026
Abstract
Accurate classification of lymphoma cell subtypes is essential for disease diagnosis and therapeutic decision-making, yet conventional approaches often rely on fluorescence labeling, labor-intensive sample preparation, and specialized instrumentation, limiting their applicability for rapid, label-free single-cell analysis. Here, we present an AI-assisted microfluidic light-scattering [...] Read more.
Accurate classification of lymphoma cell subtypes is essential for disease diagnosis and therapeutic decision-making, yet conventional approaches often rely on fluorescence labeling, labor-intensive sample preparation, and specialized instrumentation, limiting their applicability for rapid, label-free single-cell analysis. Here, we present an AI-assisted microfluidic light-scattering imaging platform for label-free classification of lymphoma cells. The platform integrates hydrodynamic focusing within a microfluidic chip, continuous acquisition of two-dimensional (2D) light-scattering patterns, automated image preprocessing, and transfer learning based on a pretrained ResNet50 network for intelligent optical feature extraction and classification. Human B lymphoma (Daudi) and T lymphoblastic lymphoma (SUP-T1) cells were used to evaluate the proposed framework. The optical imaging system was first validated using standard microspheres, demonstrating reliable acquisition of light-scattering patterns under continuous-flow conditions. A dataset comprising 800 single-cell scattering patterns was subsequently established and evaluated using stratified five-fold cross-validation. The proposed framework achieved an average classification accuracy of 94.75% with an average area under the receiver operating characteristic (ROC) curve of 0.986. By integrating microfluidic optical biosensing with deep learning, this work enables automated interpretation of intrinsic optical scattering signatures and provides a promising AI-enabled strategy for rapid, label-free lymphoma screening and intelligent healthcare applications. Full article
Show Figures

Figure 1

25 pages, 5957 KB  
Article
A Multi-Scale Fractal Feature Extraction Method for CNN-Based Plant Disease Classification
by Egor Savchenko and Anna Maslovskaya
Mach. Learn. Knowl. Extr. 2026, 8(9), 273; https://doi.org/10.3390/make8090273 - 7 Sep 2026
Abstract
Plant diseases, being a subject of interdisciplinary research, significantly reduce crop yield, quality, and economic returns, while the misidentification of pathogens often leads to ineffective treatments and may harm beneficial organisms and ecosystems. This work develops an approach for robust visual classification of [...] Read more.
Plant diseases, being a subject of interdisciplinary research, significantly reduce crop yield, quality, and economic returns, while the misidentification of pathogens often leads to ineffective treatments and may harm beneficial organisms and ecosystems. This work develops an approach for robust visual classification of plant diseases under limited and heterogeneous data based on multi-scale fractal texture descriptors integrated into a convolutional neural network. The proposed method employs wavelet transform modulus maxima to extract two complementary fractal characteristics, local fractal dimension and singularity spectrum width, from leaf images at several spatial scales. These descriptors form multi-channel fractal maps fed into a fractal attention module (FAM) inserted after the third stage of a ResNet-50 architecture. The FAM learns to emphasize spatial regions where fractal properties are most discriminative, while a parallel branch encodes global fractal statistics into an auxiliary vector combined with backbone features at the final classification layer. Experiments are conducted on a large heterogeneous collection of 11 public plant disease datasets under 5-shot, 50-shot, and full-scale training regimes. The fractal-augmented model raises classification accuracy from 57.06% to 67.73% on 5 shots and from 80.81% to 86.11% on 50 shots, red outperforming the plain ResNet-50 in these settings, converges within 1–2 epochs versus 25–40, and shows markedly better resilience to color distortions, random occlusions, and grayscale conversion in most cases. The generated attention maps provide spatially explicit explanations of the model’s decisions, increasing transparency for practical use. The proposed approach demonstrates that fractal analysis, embedded as a modulating signal inside a deep network, can serve as an efficient and interpretable inductive bias, which is particularly valuable under data scarcity and noisy agricultural imagery. Full article
Show Figures

Figure 1

29 pages, 6517 KB  
Article
Explainable Ceramic-Form Classification and Visual Retrieval of Chinese Ceramic Cultural Heritage Using DINOv2 and Morphological Feature Fusion
by Zengcheng Wang, Wuxin Liu, Yaxin Li and Bin Dong
Appl. Sci. 2026, 16(17), 8850; https://doi.org/10.3390/app16178850 - 5 Sep 2026
Viewed by 147
Abstract
The continued digitization of open museum collections provides new opportunities for the intelligent organization and visual discovery of cultural heritage. However, morphological similarity between ceramic forms, intra-class variation, and changing photographic conditions remain challenges for automated classification, model interpretation, and similar-object retrieval. This [...] Read more.
The continued digitization of open museum collections provides new opportunities for the intelligent organization and visual discovery of cultural heritage. However, morphological similarity between ceramic forms, intra-class variation, and changing photographic conditions remain challenges for automated classification, model interpretation, and similar-object retrieval. This study uses 3305 Chinese ceramic objects from the open collection of The Metropolitan Museum of Art (The Met) to develop an explainable and traceable workflow for ceramic-form classification and visual retrieval. Museum metadata were standardized into 13 form categories, from which 2755 objects were used to establish a seven-class primary classification task. Explicit morphological features, handcrafted visual features, ResNet50 representations, DINOv2 representations, and morphology–deep feature fusion were evaluated under a unified data split and evaluation protocol. Explainable artificial intelligence (XAI) methods were further used to examine spatial model responses and feature attributions of explicit morphological variables, while different representations were evaluated for content-based visual retrieval. The results show that deep visual representations effectively support ceramic-form classification, with DINOv2 demonstrating comparatively stable performance across multiple random seeds. Morphology–deep feature fusion did not provide a consistent classification advantage over DINOv2-only, but the fused representation showed clearer complementary value in visual retrieval, achieving the highest Precision@5 (0.819) and mean average precision at 10 (mAP@10; 0.773). XAI analyses further indicated that structurally meaningful spatial responses and explicit geometric descriptors contributed to form discrimination. By linking classification, interpretation, and retrieval outputs to Object IDs and original collection records, the proposed workflow provides a practical computational approach for ceramic-form organization, similar-object discovery, and traceable visual retrieval in digital museum collections. Full article
(This article belongs to the Special Issue Artificial Intelligence Technologies in Cultural Heritage)
Show Figures

Figure 1

31 pages, 3587 KB  
Article
Evaluating an Artificial Immune System-Evolved Decision-Tree Ensemble for Chest X-Ray Classification
by Abdulaziz A. Alsulami, Qasem Abu Al-Haija, Ahmad J. Tayeb, Badraddin Alturki, Ali Alqahtani and Nayef Alqahtani
Electronics 2026, 15(17), 4002; https://doi.org/10.3390/electronics15174002 - 4 Sep 2026
Viewed by 121
Abstract
Timely and accurate classification of lung diseases from chest X-ray images remains an important healthcare challenge. Machine-learning and deep-learning methods can support automated classification, but their evaluation may be affected by class imbalance, feature redundancy, dataset leakage, and computational cost. This paper evaluates [...] Read more.
Timely and accurate classification of lung diseases from chest X-ray images remains an important healthcare challenge. Machine-learning and deep-learning methods can support automated classification, but their evaluation may be affected by class imbalance, feature redundancy, dataset leakage, and computational cost. This paper evaluates an artificial immune system (AIS)-evolved decision-tree ensemble using fold-specific ResNet18 features. All within-dataset experiments use duplicate-family-aware five-fold splits. Within each fold, standardization and adaptive principal component analysis (PCA) are fitted to the training features, and the Synthetic Minority Over-sampling Technique (SMOTE) is applied only to the reduced training data. Each candidate tree is assigned an affinity based on out-of-bag macro-F1. In the primary run, mean within-dataset macro-F1 was 98.07%, 99.48%, and 98.26% for Datasets 1–3, respectively, and 96.66% for the exploratory Dataset 4. Because of extensive cross-dataset image reuse and conflicting labels, Dataset 4 does not provide independent evidence of clinical lung-cancer detection. A five-seed repeated-initialization analysis repeated the complete fold-specific feature and classification pipeline while preserving the same folds. Mean macro-F1 differences between AIS and the prespecified static comparator for each dataset, calculated as AIS minus the comparator, were 0.18, 0.00, 0.35, and 0.13 percentage points for Datasets 1–4, respectively. Using the same sign convention, mean differences between AIS and the fixed random tree ensemble ranged from 0.05 to +0.05 percentage points. Population diagnostics showed that evolution improved individual-tree macro-F1 but reduced pairwise disagreement, without a consistent majority-vote gain. Median latency from an already decoded image to prediction ranged from 38.86 to 61.20 ms on one CPU thread and from 3.43 to 6.36 ms on an RTX 4090. The results do not establish a practically important or consistent predictive advantage from AIS evolution. The study provides a reproducible and duplicate-controlled framework for evaluating AIS-based tree ensembles. Full article
Show Figures

Figure 1

23 pages, 5994 KB  
Article
A Transfer Learning and Data Augmentation Approach for Classifying Field Images of Granite Residual Slope Soils
by Zuohui Qin, Can Wang, Xin Zhou, Tengfei Yao, Wei Yin, Huimin Liang and Jian Ou
Algorithms 2026, 19(9), 755; https://doi.org/10.3390/a19090755 - 4 Sep 2026
Viewed by 162
Abstract
Granite residual and slope-wash soils are important disaster-prone geological bodies in the hilly and mountainous areas of Hunan Province, China. Their engineering classification has long relied on manual visual inspection and laboratory testing, which is inefficient and subjective. In this study, an automatic [...] Read more.
Granite residual and slope-wash soils are important disaster-prone geological bodies in the hilly and mountainous areas of Hunan Province, China. Their engineering classification has long relied on manual visual inspection and laboratory testing, which is inefficient and subjective. In this study, an automatic classification method based on deep learning image recognition is proposed for granite residual and slope-wash soils in the mountainous areas of Hunan Province. First, a three-class primary classification scheme was established, comprising residual clay (RNC), residual sandy clay (RNSC), and residual gravelly clay (RNGC), based primarily on the gravel content of particles larger than 2 mm (RNC < 5%, RNSC 5–20%, RNGC > 20%). Second, 7678 geotechnical test records from 21 counties in Hunan Province were collected, and classification labels were assigned through a strategy combining manual verification and automatic inference using Random Forest (5-fold cross-validation macro F1 = 0.913). From approximately 10,096 original field images, 3380 pure soil image patches were retained after segmentation and screening. A dataset of 43,940 samples was then generated through two-stage preprocessing (including denoising and illumination correction) and 13-fold data augmentation. A CNN image classification model was constructed based on a ResNet18 backbone network pre-trained on ImageNet. On the independent test set (6591 images), the primary classification accuracy reached 91.46%, with a macro F1-score of 0.9128; the per-class F1-scores for RNC, RNSC, and RNGC were 0.921, 0.885, and 0.933, respectively. Grad-CAM visualization analysis demonstrated that the model’s attention was primarily focused on soil particle distribution regions rather than non-soil background areas, confirming the effective learning of mixed-grain features. The study shows that the combined application of transfer learning and 13-fold data augmentation can significantly improve classification performance under limited sample size conditions (an improvement of 15.33 percentage points compared to the baseline of 76.13%), demonstrating promising potential for engineering applications. Full article
Show Figures

Figure 1

Back to TopTop