Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (330)

Search Parameters:
Keywords = EfficientNetV2B2

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
25 pages, 5154 KB  
Article
X-Net: Hybrid Transformer–CNN Framework with Multi-Aspect Attention for Precise Colon Polyp Segmentation
by Benyamin Mirab Golkhatmi and Mohammad Hossein Moattar
AI 2026, 7(9), 336; https://doi.org/10.3390/ai7090336 (registering DOI) - 30 Aug 2026
Abstract
Accurate localization of colon polyps plays a crucial role in the early diagnosis of colon cancer. In this paper, we propose a novel network architecture, X-Net, which integrates the Pyramid Vision Transformer (PVT) and EfficientNet-B0. We selected EfficientNet-B0 due to its optimized design [...] Read more.
Accurate localization of colon polyps plays a crucial role in the early diagnosis of colon cancer. In this paper, we propose a novel network architecture, X-Net, which integrates the Pyramid Vision Transformer (PVT) and EfficientNet-B0. We selected EfficientNet-B0 due to its optimized design and computational efficiency, enabling it to effectively capture fine-grained local features. Meanwhile, we use PVT V2-b1 (Pb1) because of its hierarchical structure and self-attention mechanism, which enhance the model’s ability to preserve multi-scale features. By concatenating these two networks in part of the decoder, the proposed model effectively leverages both local and global features. Additionally, the X-Net architecture introduces three blocks: Shape-Aware Enhancement Block (SAEB), Multi-Scale Feature Block (MSFB), and Multi-Aspect Attention Block (MAAB). The SAEB utilizes differences in initial-layer features extracted from the encoder to filter noise in polyp images. The MSFB extracts and enhances multi-scale features from the backbone, effectively integrating semantic information, which improves polyp segmentation accuracy and enhances robustness against variations in polyp size and shape. Finally, the MAAB employs high-level features from the outputs of the MSFB and SAEB to attend to the foreground, background, and boundaries. To address the class imbalance problem, we propose a hybrid loss function combining Tversky Loss and Binary Cross-Entropy Loss. For evaluation, we trained the proposed network on four polyp datasets: Kvasir-SEG, CVC-ClinicDB, CVC-T, and ETIS. The results demonstrate that the model performs well across different datasets with varying image quality and characteristics. Full article
Show Figures

Figure 1

27 pages, 6197 KB  
Article
Edge-Based Facial Emotion Recognition for Nurse-Assistive Robots Using a Compact CNN
by Quoc-Cuong Pham, Thanh-Long Le, Huy-Hoang Pham and Huu-Dung Nguyen
Technologies 2026, 14(9), 535; https://doi.org/10.3390/technologies14090535 (registering DOI) - 29 Aug 2026
Abstract
Facial emotion recognition (FER) can provide supplementary affective information for human–robot interaction, but deployment on resource-constrained assistive robots requires a balance between recognition performance and computational efficiency. This study presents an edge-based FER framework using a compact CNN operating on 48 × 48 [...] Read more.
Facial emotion recognition (FER) can provide supplementary affective information for human–robot interaction, but deployment on resource-constrained assistive robots requires a balance between recognition performance and computational efficiency. This study presents an edge-based FER framework using a compact CNN operating on 48 × 48 grayscale facial images and retaining all seven FER-2013 expression categories. Square-root-smoothed inverse-frequency weighting is employed to mitigate class imbalance without excessively emphasizing rare classes. On the held-out FER-2013 test set, the proposed model achieves 63.78% Accuracy and 59.32% Macro-F1, achieving higher Accuracy and Macro-F1 than the evaluated ImageNet-pretrained MobileNetV2 and MobileNetV3-Small baselines. INT8 post-training quantization reduces model size by 74.13% relative to FP32, with decreases of only 0.91 and 0.50 percentage points in Accuracy and Macro-F1, respectively. On a Raspberry Pi 3 Model B+, INT8 achieves a mean model-only inference latency of 19.30 ms and a model-only throughput of 51.82 FPS. Using an actor-disjoint RAVDESS protocol comprising 416 videos, EMA stabilization reduces prediction switching by 54.86% on the held-out test actors. These results support the feasibility of compact edge-based FER for assistive robotic interaction while emphasizing that the framework provides supplementary affective cues rather than clinical diagnosis or autonomous decision-making. Full article
(This article belongs to the Special Issue Advances in Automatics, Robotics & Artificial Intelligence)
Show Figures

Figure 1

28 pages, 3426 KB  
Article
Image-Based Identification of Crop and Weed Species at the Seedling Stage: Deep Learning Classifiers Rely on Acquisition Context Beyond Plant Morphology
by Lyna Miloudi, Khaled Rezeg and Mohamed Kotoub Miloudi
Int. J. Plant Biol. 2026, 17(9), 80; https://doi.org/10.3390/ijpb17090080 (registering DOI) - 29 Aug 2026
Abstract
Distinguishing crop from weed species at the seedling stage is a fine-grained morphological discrimination problem. We study it on the public Plant Seedlings benchmark, which contains twelve species (three crops and nine weeds) imaged at the seedling stage. Automated classifiers report near-ceiling accuracy [...] Read more.
Distinguishing crop from weed species at the seedling stage is a fine-grained morphological discrimination problem. We study it on the public Plant Seedlings benchmark, which contains twelve species (three crops and nine weeds) imaged at the seedling stage. Automated classifiers report near-ceiling accuracy on this benchmark. Yet its images are acquired in controlled trays containing soil, gravel, rulers, barcodes, and printed labels, so a model may identify a species from its acquisition context rather than its morphology. We audit two architectures, EfficientNet-B7 and ViT-B/16, trained on the V2 dataset (5539 images) and probed with plant-only, background-only, and background-swapped inputs. Near-ceiling models (96.1% and 96.4% over three seeds) recover the correct species for up to 47.7% of samples from the background alone (chance 8.3%) and lose about 60 points under background swapping. Context reliance is therefore a property of the benchmark, not any single architecture. Tracing this to its source, an independent learned representation of the plant-free background alone identifies the species at 72.4%. The reliance is correctable end to end: a consistency-regularisation scheme retains 92.1% full-image accuracy for the transformer with no segmentation at inference, at an architecture-dependent cost. Reported accuracy thus partly measures acquisition context, not morphology; morphological grounding should be measured and reported alongside accuracy. Full article
(This article belongs to the Section Application of Artificial Intelligence in Plant Biology)
30 pages, 17083 KB  
Article
Edge-Deployable Lightweight Deep Learning for Hypertensive Retinopathy Grading from En-Face OCT: A Patient-Level Feasibility Study
by Süleyman Burçin Şüyun, Mustafa Yurdakul, Şakir Taşdemir and Serkan Biliş
Bioengineering 2026, 13(9), 984; https://doi.org/10.3390/bioengineering13090984 - 26 Aug 2026
Viewed by 116
Abstract
Background: Hypertensive retinopathy (HR) is an early marker of hypertension-related end-organ damage, and multi-grade classification is limited by overlap between the early stages. No prior study has reported edge-deployable HR grading from optical coherence tomography (OCT). We assess patient-level–validated four-grade HR classification from [...] Read more.
Background: Hypertensive retinopathy (HR) is an early marker of hypertension-related end-organ damage, and multi-grade classification is limited by overlap between the early stages. No prior study has reported edge-deployable HR grading from optical coherence tomography (OCT). We assess patient-level–validated four-grade HR classification from en-face OCT with on-device inference. Methods: MobileViT-XXS and EfficientNetV2-B0 were evaluated on 478 en-face OCT images from 221 patients. To avoid leakage, all partitions were patient-level: stratified group 5-fold cross-validation with a patient-disjoint test set. Each model was evaluated as a 5-fold soft-voting ensemble, with inference profiled on an NVIDIA Jetson Orin Nano. Accuracy, weighted F1, quadratic-weighted kappa, and ROC–AUC were reported; the models were compared by McNemar’s test, and clinical utility by referable-HR (Grade ≥ 2) triage. Results: MobileViT-XXS achieved 84.4% accuracy, ROC–AUC 0.944, and quadratic-weighted kappa 0.916, significantly outperforming EfficientNetV2-B0 (McNemar p = 0.013). Mild-grade overlap limited four-class accuracy, but referable-HR triage reached 95.6% accuracy, 95.2% sensitivity, and 95.8% negative predictive value. MobileViT-XXS required only 3.83 MB versus 23.70 MB, and on-device TensorRT FP16 five-fold ensemble inference achieved 10.4 ms (96 FPS) at 8.9 W. Conclusions: Lightweight models enable feasible, power-efficient point-of-care HR triage from en-face OCT; broader multi-center validation is warranted. Full article
Show Figures

Figure 1

27 pages, 15343 KB  
Article
AMFF-Net: An Adaptive Multi-Layer Feature Fusion Network Based on ConvNeXtV2-B for Medical Image Classification
by Min-Seo Kim and Hyoung-Gook Kim
Bioengineering 2026, 13(9), 983; https://doi.org/10.3390/bioengineering13090983 - 26 Aug 2026
Viewed by 167
Abstract
Although ConvNeXtV2 has shown promising performance in medical image classification, approaches relying primarily on final-stage features may underutilize low-level structural and complementary hierarchical information. To address this limitation, we propose an Adaptive Multi-Layer Feature Fusion Network (AMFF-Net) based on ConvNeXtV2-B for medical image [...] Read more.
Although ConvNeXtV2 has shown promising performance in medical image classification, approaches relying primarily on final-stage features may underutilize low-level structural and complementary hierarchical information. To address this limitation, we propose an Adaptive Multi-Layer Feature Fusion Network (AMFF-Net) based on ConvNeXtV2-B for medical image classification. The proposed framework employs a Feature Alignment (FA) module to project multi-stage features into a unified representation space and an Adaptive Multi-Layer Feature Fusion (AMFF) module to compute a single input-dependent scalar weight for each stage and dynamically adjust the relative contributions of hierarchical features. An Efficient Channel Attention (ECA) module is subsequently incorporated to enhance the fused representation through lightweight channel-wise recalibration. AMFF-Net was evaluated on three medical image classification datasets: Kvasir-v2, HAM10000, and ChestXray14. Experimental results demonstrate that AMFF-Net consistently improves classification performance over the baseline ConvNeXtV2-B and achieves competitive performance compared with representative convolutional neural network (CNN)- and Transformer-based architectures, while incurring relatively modest additional computational overhead. Ablation results further support the contribution of FA, AMFF, and ECA to the overall performance of the proposed framework. Full article
Show Figures

Figure 1

25 pages, 13774 KB  
Article
A Feasibility Study of Deep Learning-Based Motor Defect Screening in a Production Line Using an Airborne Acoustic Signal
by Thorikul Huda, Faaris Mujaahid and Min-Fu Hsieh
Appl. Sci. 2026, 16(17), 8418; https://doi.org/10.3390/app16178418 - 24 Aug 2026
Viewed by 217
Abstract
Reliable defect detection in motor production lines is important for maintaining manufacturing quality. This study investigates the feasibility of a deep learning-based airborne acoustic screening approach for controlled no-load end-of-line induction motor inspection. Acoustic signals collected under controlled no-load test conditions were used [...] Read more.
Reliable defect detection in motor production lines is important for maintaining manufacturing quality. This study investigates the feasibility of a deep learning-based airborne acoustic screening approach for controlled no-load end-of-line induction motor inspection. Acoustic signals collected under controlled no-load test conditions were used to evaluate Feedforward Neural Networks (FNNs), Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), and pre-trained models including ResNet50, MobileNetV2, and EfficientNetB0 for detecting rotor unbalance, assembly-induced bearing abnormalities, and their combination. Experimental results show that several models achieved up to 96% classification accuracy, depending on the selected architecture and feature representation. In cross-motor external validation using an unseen motor from the same manufacturer, the CNN with MFCC features achieved the best performance with 96% accuracy and 97% precision, recall, and F1-score, while ResNet50 with spectrogram inputs achieved 93% accuracy and 92% F1-score. These results demonstrate that airborne motor acoustic signals contain discriminative defect-related information under controlled no-load conditions and support the feasibility of low-cost, non-contact airborne acoustic sensing as a complementary screening approach for rapid end-of-line motor inspection. Full article
Show Figures

Figure 1

12 pages, 3198 KB  
Article
Comparative Evaluation of Deep Transfer Learning Models for Ancient Coin Classification
by Omar Alhuniti, Sami Serahan and Imad Salah
Appl. Sci. 2026, 16(16), 8271; https://doi.org/10.3390/app16168271 - 19 Aug 2026
Viewed by 222
Abstract
In this paper, we present deep learning models for classifying ancient-coins using transfer learning applied to a custom dataset curated from a specialized collection. The dataset was meticulously developed through rigorous preparation and quality screening; the retained fine-grained classes are imbalanced. We initially [...] Read more.
In this paper, we present deep learning models for classifying ancient-coins using transfer learning applied to a custom dataset curated from a specialized collection. The dataset was meticulously developed through rigorous preparation and quality screening; the retained fine-grained classes are imbalanced. We initially performed binary classification before progressing to separate fine-grained assessments of DenseNet121, EfficientNetB0, EfficientNetV2S, and ResNet50. In confirmatory experiments using a physical-coin-aware 70/15/15 split and five training seeds, the coarse CITY-versus-NABATAEAN task remained approximately perfectly separable. Fine-grained performance was lower: ResNet50 achieved 68.28 ± 2.45% CITY accuracy, whereas DenseNet121 achieved 72.89 ± 2.79% NABATAEAN accuracy. These results show that near-perfect coarse classification does not by itself imply reliable fine-grained attribution. This analysis underscores the importance of meticulous dataset preparation and strategic model selection in optimizing performance. A matched DenseNet121 ablation further showed that ImageNet initialization substantially improved the difficult fine-grained tasks compared with random initialization. The proposed framework supports digital documentation and analysis of ancient coin collections, while the current single-collection evaluation still limits claims of cross-collection generalization. Full article
(This article belongs to the Special Issue Artificial Intelligence Applications in Tourism)
Show Figures

Figure 1

29 pages, 27131 KB  
Article
Deep Learning-Assisted Quality Control of Histology Teaching Slides: Detection and Localization of Tissue Fold Artifacts in H&E-Stained Images
by Osman Fatih Koparir, Berrin Tarakci Gencer and Abdulkadir Sengur
Bioengineering 2026, 13(8), 937; https://doi.org/10.3390/bioengineering13080937 - 19 Aug 2026
Viewed by 401
Abstract
Background/Objectives: Tissue fold artifacts observed in hematoxylin-eosin (H&E)- stained preparations used in histology education can complicate the assessment of normal tissue architecture and affect students’ accurate interpretation of microscopic structures. This study aimed to automatically detect and localize tissue fold artifacts in [...] Read more.
Background/Objectives: Tissue fold artifacts observed in hematoxylin-eosin (H&E)- stained preparations used in histology education can complicate the assessment of normal tissue architecture and affect students’ accurate interpretation of microscopic structures. This study aimed to automatically detect and localize tissue fold artifacts in histology teaching preparations using deep learning methods. Methods: A total of 2127 hematoxylin and eosin (H&E)-stained histological images of brain, kidney, liver, small intestine, and testis tissues obtained at 10× magnification were used in the study. The dataset consisted of 899 clean/artifact-free images and 1228 images containing tissue fold artifacts. Seven deep learning architectures, including convolutional neural networks (CNN)-based models and a vision transformer-based model, were evaluated for image-level classification: ResNet18, ResNet50, DenseNet121, EfficientNet-B0, EfficientNet-B3, ConvNeXt-Tiny, and Swin-Tiny. Classification performance was evaluated using an organ-based testing approach, a slide-level train–test split in which images from the same histological slide group were retained within a single subset, a general image-level 80/20 train–test split, and five-fold cross-validation. A DeepLabV3-ResNet50-based segmentation model was trained using QuPath-prepared masks to determine fold regions at the pixel level. Grad-CAM was used for qualitative visualization and quantitative comparison with manually annotated fold regions. Results: All classification models showed excellent performance overall. The Swin-Tiny model was the most accurate in terms of general image-level classification, scoring 99.06% for accuracy, 99.19% for F1-score, and 99.98% for AUC. In slide-level classification, EfficientNet-B3 showed the best performance in terms of accuracy (99.54%) and AUC (99.99%). No pairwise statistical difference was found among the models using Holm adjustment. In the five-fold cross-validation setting, ResNet50 achieved the best performance with the accuracy of 98.73 ± 0.54%. In the detailed small-intestine error analysis, classification accuracy was 67.67%, with high sensitivity (98.59%) but low specificity (24.34%), mainly because of false-positive predictions. In the segmentation evaluation, the final DeepLabV3-ResNet50 model trained with combined BCE + Dice loss resulted in Dice of 0.7630 ± 0.2425 and IoU of 0.6661 ± 0.2577 on the independent test set. False positive segmentations were rare in 899 artifact-free images (only 0.33% of images had a tissue fold region of at least 1%). The quantitative Grad-CAM analysis showed poor spatial agreement with the manually annotated tissue fold masks (Dice = 0.2423; IoU = 0.1441). In the independent MPP10 dataset, ResNet50 achieved 89.63% accuracy and 88.49% F1-score. Conclusions: The suggested method was able to detect tissue fold artifacts as well as localize them in the teaching images of histology. The strong performance recorded at the slide level validates the internal findings, but the limited performance in the case of small intestine images and on the external dataset reveals that tissue structure remains a crucial factor. Full article
(This article belongs to the Special Issue Machine Learning-Aided Medical Image Analysis: Second Edition)
Show Figures

Graphical abstract

12 pages, 1764 KB  
Article
Machine Learning-Based Classification of Retinitis Pigmentosa from Color Fundus Images: A Reproducible Benchmark and Screening-Oriented Pipeline
by Francesco Cappellani, Giovanni Rubegni, Andrea Caruso, Alessia Cosentino, Roberta Torrisi, Grazia Pia Raciti, Marco Mastroeni, Gabriella Lupo, Caterina Gagliano and Massimiliano Salfi
Vision 2026, 10(3), 55; https://doi.org/10.3390/vision10030055 - 18 Aug 2026
Viewed by 238
Abstract
Retinitis pigmentosa (RP) is a rare inherited retinal disorder in which fundus changes may be subtle and heterogeneous, limiting detection from color fundus images. This study evaluated multiple machine learning architectures for binary-RP versus healthy-control classification, and developed a reproducible pipeline for research-oriented [...] Read more.
Retinitis pigmentosa (RP) is a rare inherited retinal disorder in which fundus changes may be subtle and heterogeneous, limiting detection from color fundus images. This study evaluated multiple machine learning architectures for binary-RP versus healthy-control classification, and developed a reproducible pipeline for research-oriented screening support. Three publicly available fundus datasets were combined, including 248 RP images and 1045 healthy controls. An 80/20 train–test split was used, with targeted data augmentation applied only to RP images in the training set to address class imbalance. ConvNeXt-Tiny, ResNet101V2, EfficientNet-B0, a baseline classifier, and custom shallow convolutional neural networks were compared using accuracy, precision, recall, F1-score, confusion matrices, and ROC/precision–recall analyses. A compact ShallowCNN provided the best sensitivity–performance trade-off. On the fixed image-level test set, Adam with a learning rate of 0.0005 reached 96.51% accuracy, while SGD with a learning rate of 0.001 achieved 98% RP recall, minimizing false negatives. The trained models were exported to ONNX and integrated into a Windows inference tool. The proposed framework provides an open, reproducible benchmark for technical evaluation, although external validation is required before clinical use. Full article
(This article belongs to the Section Retinal Function and Disease)
Show Figures

Graphical abstract

27 pages, 12347 KB  
Article
Cotton Leaf Disease Detection via Dual-Backbone CNN-Transformer Fusion with Quantitative XAI Comparison
by Naeem Ullah, Ivanoe De Falco and Giovanna Sannino
Electronics 2026, 15(16), 3650; https://doi.org/10.3390/electronics15163650 - 16 Aug 2026
Viewed by 451
Abstract
Deep learning has shown promise for cotton leaf disease detection, yet two critical gaps remain. First, most studies rely on a single model (Convolutional Neural Network-CNN or Transformer) and do not explore how to effectively fuse these complementary architectures. Second, eXplainable AI (XAI) [...] Read more.
Deep learning has shown promise for cotton leaf disease detection, yet two critical gaps remain. First, most studies rely on a single model (Convolutional Neural Network-CNN or Transformer) and do not explore how to effectively fuse these complementary architectures. Second, eXplainable AI (XAI) methods are often used qualitatively, lacking objective benchmarks to guide method selection. To address these gaps, we evaluate six backbone models, comprising four CNNs (ResNet50, EfficientNet-B0, DenseNet121, and MobileNetV2) and two Vision Transformers (ViT-Base and DeiT-Small), on the Kaggle cotton leaf disease dataset, which contains 1711 images across four classes. We then systematically investigate five CNN–Transformer fusion strategies, namely concatenation, attention, weighted, ensemble, and variance-based fusion, to identify the most effective approach for disease classification. The best-performing individual models are DenseNet121 (92.40% accuracy) and ViT-Base (96.49% accuracy). Classification metrics include accuracy, balanced accuracy, precision/recall, F1-score, Cohen’s kappa, MCC, AUC, bootstrap confidence intervals, and McNemar tests. Computational efficiency (FLOPs, inference time, model size) is also reported. Concatenation fusion achieves the highest performance (accuracy = 99.42%, 95% CI: 98.2–100%, weighted F1 = 0.994, MCC = 0.992). For explainability, we quantitatively compare six XAI techniques, GradCAM, GradCAM++, ScoreCAM, LayerCAM, EigenCAM, and AblationCAM, using the pointing game, IoU, AUC, and localization accuracy. EigenCAM yields the best overall explainability score. This study demonstrates that simple feature concatenation between dual backbones (CNN + Transformer) is highly effective for cotton leaf disease detection and provides a benchmark for XAI method selection in plant pathology. Full article
(This article belongs to the Special Issue Advanced Computer Science and Intelligent Systems Innovations)
Show Figures

Figure 1

23 pages, 11726 KB  
Article
Research on Wheat Drought Stress Recognition Based on Improved EfficientNet-B0
by Jianbin Yao, Meijia Wang, Linyuan Li, Xinjie Xue and Jingke Sun
Agronomy 2026, 16(16), 1565; https://doi.org/10.3390/agronomy16161565 - 14 Aug 2026
Viewed by 201
Abstract
Wheat is one of the major staple crops in China, and drought stress can severely affect its growth, development, and yield. Rapid and accurate identification of drought stress levels in wheat is of great significance for agricultural disaster prevention and mitigation, as well [...] Read more.
Wheat is one of the major staple crops in China, and drought stress can severely affect its growth, development, and yield. Rapid and accurate identification of drought stress levels in wheat is of great significance for agricultural disaster prevention and mitigation, as well as for ensuring food security. To address the problems of insufficient fine-grained feature extraction, class imbalance, and unstable training in wheat drought stress image recognition, this study proposes an improved EfficientNet-B0 model for fine-grained wheat drought stress classification. Based on EfficientNet-B0, an improved lightweight Efficient Multi-scale Attention (EMA) module is introduced after the backbone network to enhance both channel-wise and spatial feature representation. PolyLoss is adopted to enhance the learning of low-confidence and difficult samples under the uneven class distribution, while the Sharpness-Aware Minimization (SAM) optimizer is employed to improve the optimization process. Experiments were conducted on a 15-class wheat drought stress image dataset constructed from three key growth stages and five drought severity levels. The proposed model achieved an accuracy of 98.69% and an F1-score of 98.37% on the internal test set, outperforming the baseline EfficientNet-B0 and comparison models including ResNet-50, DenseNet-121, and MobileNetV3. Component-level ablation experiments further showed that the dual-gating structure, projection residual connection, and intra-group Softmax normalization in the proposed EMA module all contributed positively to model performance. These results indicate that the proposed method provides a lightweight and effective approach for wheat drought stress recognition within the current dataset. Full article
Show Figures

Figure 1

29 pages, 3461 KB  
Article
Benchmarking Class Imbalance Mitigation Strategies Across Deep CNN Architectures for Skin Cancer Classification
by Irshad Ahmad, Muhammad Khubaib and Saleh M. Altowaijri
Diagnostics 2026, 16(16), 2571; https://doi.org/10.3390/diagnostics16162571 - 14 Aug 2026
Viewed by 340
Abstract
Background/Objectives: Class imbalance is one of the major challenges in automated skin lesion classification since the number of categories of malignant and clinically significant skin lesions is normally less than the benign ones. However, due to this imbalance, deep convolutional neural networks [...] Read more.
Background/Objectives: Class imbalance is one of the major challenges in automated skin lesion classification since the number of categories of malignant and clinically significant skin lesions is normally less than the benign ones. However, due to this imbalance, deep convolutional neural networks (CNNs) tend to overlook minority classes and fail to recognize them with an acceptable accuracy, which leads to a decrease in diagnostic reliability. A wide range of imbalance mitigation techniques has been suggested, but their effectiveness is found to differ significantly depending on CNN architecture, and detailed comparative studies of these techniques for a consistent experimental setup are still limited. Methods: This study proposes a comprehensive benchmarking framework that tests sixteen class imbalance mitigation methods by applying them to six pretrained CNN architectures—EfficientNet-B0, EfficientNet-B3, ResNet50, DenseNet121, InceptionV3 and MobileNetV2—on the official ISIC 2019 skin lesion dataset. The tested techniques are conventional resampling techniques, synthetic sample generation techniques, algorithm-level learning techniques, data augmentation techniques, and hybrid techniques. The dataset was partition into a separate training set and testing set, and stratified cross-validation was only conducted on the training set to ensure the study was fair and reproducible. Both models have been optimized with the same optimizer, learning rate, batch size, epochs and preprocessing pipeline. The performance of the models was evaluated by computing the accuracy, precision, recall and F1-score. Results: The experimental results show that the effect of class imbalance mitigation is very specific to the underlying CNN architecture. The traditional undersampling and oversampling methods yielded only moderate improvements, while feature space and hybrid methods yielded more consistent results. When coupled with EfficientNet-B3, Balanced MixUp improved the overall performance of the model by achieving an accuracy of 92.39%, an increase in precision of 93.3%, a recall of 91.36%, and an F1-score of 92.33%. However, some architectures such as ResNet50 performed better with iterative learning techniques, such as Cumulative Learning and Yielding Multi-Fold Training, which suggests that there is a diversity in how different network architectures react to imbalance mitigation methods. Conclusions: This paper highlights the importance of selecting appropriate technique–architecture combinations for addressing long-tailed data distributions in medical imaging. The proposed benchmarking framework provides valuable insights for developing robust and reliable deep learning systems for skin lesion classification and other medical imaging tasks affected by severe class imbalance. Full article
Show Figures

Figure 1

13 pages, 13500 KB  
Article
A Lightweight One-Shot Open-Set Metric Learning Framework for Food Recognition and Decision Support in Smart Ovens
by Nurdanur Pehlivan and Resul Kara
Electronics 2026, 15(16), 3533; https://doi.org/10.3390/electronics15163533 - 9 Aug 2026
Viewed by 273
Abstract
Modern smart kitchen automation requires reliable vision-based tools to provide user-advisory decision support during domestic culinary processes. However, standard deep learning models utilizing closed-set Softmax classifiers typically misclassify unknown or Out-of-Distribution (OOD) kitchen objects with high confidence, posing safety and reliability risks. To [...] Read more.
Modern smart kitchen automation requires reliable vision-based tools to provide user-advisory decision support during domestic culinary processes. However, standard deep learning models utilizing closed-set Softmax classifiers typically misclassify unknown or Out-of-Distribution (OOD) kitchen objects with high confidence, posing safety and reliability risks. To address this problem without clous dependency, this study introduces a localized open-set metric learning framework based on a modified MobileNetV2 architecture. The conventional Softmax classification layer is replaced with a feature embedding layer evaluated via Cosine Similarity and a calibrated decision threshold. This architecture tracks targeted food items across four operational stages—counter-raw, in-oven-raw, in-oven-cooked, and counter-cooked—while identifying and rejecting OOD objects. To ensure reproducibility, comprehensive experimental validations were conducted on a dedicated internal dataset, providing direct baseline comparisons against mainstream backbones (ResNet50, EfficientNet-B0, and Vision Transformers) stripped of their Softmax layers and evaluated under identical metric constraints. The results demonstrate that the proposed framework achieves a Macro F1-score 92.6% and ultra-low inference latency of 11.5 ms, ensuring an optimized trade-off between Macro F1-score and inference speed on edge computing environments. This framework establishes a robust, self-contained solution for open-set object recognition in localized smart home appliances. Full article
(This article belongs to the Special Issue AI Technologies and Smart City)
Show Figures

Figure 1

23 pages, 5896 KB  
Article
Hybrid Decision-Level Fusion of CNN-Based Deep and Handcrafted Features for Colon Cancer Classification
by Simona Moldovanu, Adina Cocu, Diana Stefanescu and Cătălin Anghel
J. Imaging 2026, 12(8), 368; https://doi.org/10.3390/jimaging12080368 - 9 Aug 2026
Viewed by 271
Abstract
In recent years, there has been increased attention on classifying histopathological images through hybrid decision-level fusion, and the challenge of exploring data fusion to improve classification accuracy in colon cancer has become significant. This study introduces new elements by incorporating various deep learning [...] Read more.
In recent years, there has been increased attention on classifying histopathological images through hybrid decision-level fusion, and the challenge of exploring data fusion to improve classification accuracy in colon cancer has become significant. This study introduces new elements by incorporating various deep learning (DL) architectures, including EfficientNetB0, DenseNet121, ResNet101V2, NASNetMobile, MobileNetV2, and VGG16 Convolutional Neural Networks (CNNs), as well as Random Forest (RF) and Histogram Gradient Boosting (HGB) Machine Learning (ML) algorithms, along with the original dataset. The proposed hybrid decision-level fusion approach analyzes the LC25000 dataset’s colon histopathological images and handcraft features (HFs) to improve predictive performance. The HFs such as entropy, the Gini index, and the radius of gyration from adenocarcinoma and benign colon tissue (CC) were extracted. The prediction of the proposed models leveraging late fusion was conducted by classifying both deep and HFs. During the experiments, it was demonstrated that the combination of ResNet101V2 with RF classifier and all HFs yielded greater accuracy and consistent performance. The achieved performance metrics include accuracy of 94.1%, F1-score of 94%, Matthews Correlation Coefficient (MCC) of 88.3%, and an area under the curve (AUC) of 0.979. To explain and interpret the decisions made by the DL models, the explainable methods SHapley Additive exPlanations (SHAP) and Gradient-weighted Class Activation Mapping (Grad-CAM) were utilized. Full article
(This article belongs to the Special Issue Learning and Optimization for Medical Imaging2nd Edition)
Show Figures

Figure 1

22 pages, 5276 KB  
Article
Research on an Improved Multi-Model Dynamic Fusion Classification Technique
by Xiang Wan, Youxing He, Xionghai Rao, Yijian Qiu, Ruijian Cheng, Jiang Wei, Xiangping Cheng, Tianci Li and Manqing Zhu
Electronics 2026, 15(15), 3490; https://doi.org/10.3390/electronics15153490 - 6 Aug 2026
Viewed by 342
Abstract
Existing multi-model fusion methods generally adopt a global, fixed fusion strategy, applying all base models and weighting rules uniformly to all samples to be classified and all categories, thereby lacking adaptability to specific samples and category-specific targeting. Most traditional dynamic selection and dynamic [...] Read more.
Existing multi-model fusion methods generally adopt a global, fixed fusion strategy, applying all base models and weighting rules uniformly to all samples to be classified and all categories, thereby lacking adaptability to specific samples and category-specific targeting. Most traditional dynamic selection and dynamic weighting fusion methods only implement global parameter adjustments at the sample level, without considering the significant differences in category-specific capabilities among the base models. When a single base model exhibits superior recognition capabilities for only certain categories, its prediction accuracy and confidence for the remaining categories are low. If all base models are fused directly, poor-quality class predictions can cause negative interference and even dominate the final decision, leading to classification errors. Furthermore, traditional dynamic fusion suffers from computational redundancy, difficulty in suppressing interference from low-confidence samples, and the challenge of balancing dynamic optimization with inference efficiency. To address these issues, this paper proposes an improved multi-model dynamic fusion classification technique that differs from the traditional global dynamic fusion paradigm. By constructing a voting matrix, contribution weights, and a matrix of effective category voting weights, this method establishes a category-level model performance evaluation and differentiated weighting mechanism. This enables the precise selection of superior base models for each sample and category, thereby filtering out interference from low-confidence and suboptimal category predictions. At the same time, in the network architecture design, lightweight models are organized into a branch structure, and effective branches are dynamically activated as needed to participate in decision-making, significantly reducing the computational overhead of inference. To validate the fusion classification technique proposed in this paper, for the experiments, we selected mainstream lightweight models such as MobileNetV2, EfficientNetB0, ShuffleNetv2, MNASNet 0.75, and MobileNetV3_Small for evaluation on the NEU dataset and NASA’s Milling Data Set. The experimental results demonstrate that the fusion method proposed in this paper can fully aggregate the category-specific strengths of different lightweight models, effectively mitigate the risk of misclassification associated with traditional fusion methods, and enhance model robustness while ensuring high classification accuracy. It achieves classification performance comparable to that of large deep models with extremely low computational overhead. This method is not only suitable for application in multiple-criteria decision-making but can also be implemented and extended to multi-source/multi-modal data fusion and deep neural networks, making it of practical value. Full article
(This article belongs to the Special Issue Multimodal Learning and Transfer Learning)
Show Figures

Figure 1

Back to TopTop