A Deep Learning Model for Chili Pepper Fruit Shape Classification Using DenseNet-121 and CBAM
Abstract
1. Introduction
2. Dataset Construction and Data Augmentation
2.1. Dataset Construction
2.2. Data Augmentation
3. Model Development
3.1. Network Framework
- Data Augmentation and Preprocessing Module: This module utilizes an offline targeted data augmentation strategy via the Albumentations computer vision library to construct a class-balanced dataset, thereby enhancing model robustness directly at the data source level.
- Model Architecture Module: This module forms the core of the chili pepper fruit shape classification process. It utilizes DenseNet-121 as the backbone feature extraction network, leveraging dense connections to promote feature reuse and optimize gradient propagation. Subsequently, CBAM is integrated to adaptively recalibrate feature maps by sequentially applying channel and spatial attention mechanisms, thereby directing the model’s focus toward critical regions of fruit shape heterogeneity. Finally, the refined features are processed through a global average pooling (GAP) layer, a dropout layer, and a fully connected (FC) layer to map the outputs into the eight distinct fruit shape categories.
- Learning and Optimization Module: This module governs model weight iteration and parameter optimization. Training is driven by a Stochastic Gradient Descent (SGD) optimizer with an initial learning rate of 1 × 10−3, which is dynamically adjusted via a cosine annealing schedule. Concurrently, a cross-entropy loss function with LS is employed alongside an L2 WD of 1 × 10−4. This configuration mitigates overfitting, facilitating feature learning for highly similar chili pepper fruit shapes.
- Inference and Performance Evaluation Module: This module verifies model generalization capabilities. Once training convergence is achieved, forward inference is executed on an independently reserved test set, where discriminative and classification performance is assessed across multiple dimensions using precision, recall, F1-score, and a confusion matrix.
3.2. Model Construction and Optimization Strategies
3.2.1. DenseNet-121 Backbone
3.2.2. Convolutional Block Attention Module
3.2.3. Global Average Pooling
3.2.4. Dropout
3.2.5. Classification Head and Output Layer
3.2.6. Loss Function: Cross-Entropy with Label Smoothing
3.2.7. Weight Decay
3.3. Comprehensive Evaluation Methods for Model Performance
3.4. Comparative Experimental Design and Baseline Models
4. Results and Analysis
4.1. Evaluation Results of Baseline Models
4.2. Effects of Optimizers on Model Performance
4.3. Effects of Initial Learning Rate Configurations on Model Performance
4.4. Effects of Regularization on Model Performance
4.5. Effects of Attention Mechanisms on Model Performance
4.6. Ablation Study
4.7. Confusion Matrix and Fine-Grained Feature Analysis of Fruit Shape Classes
4.8. Comparison with State-of-the-Art Models
4.9. Visualization of Feature Maps
5. Discussion
6. Conclusions
- To address the difficulty of extracting fine-grained shape features under environmental interference, the CBAM module was introduced to capture key regions such as fruit contours and tip curvature. Grad-CAM visualizations show that the proposed model localizes the morphological boundaries of the chili peppers, avoiding the background dispersion and feature fragmentation observed in certain lightweight networks and the Swin-Tiny model;
- To address the susceptibility of deep convolutional networks to local optima and hard-label overconfidence in fine-grained tasks, cross-entropy loss with LS (LS = 0.1) was introduced to enhance the robustness of the decision boundaries. Quantitative results show that this strategy, combined with joint regularization (Dropout = 0.3, WD = 1 × 10−4) and an initial learning rate of 1 × 10−3, mitigates model overfitting to background noise in the training set;
- The proposed model and four benchmark networks were trained and evaluated on the constructed chili pepper shape dataset. The results show that the proposed model achieved a precision of 90.09%, a recall of 89.60%, an F1-score of 89.53%, and an overall accuracy of 89.74%. Under mild image degradation typical of real-world production environments, the model maintained a robust F1-score of 87.24%. Furthermore, with 7.09 M parameters and a single-frame inference time of 7.35 ms, the model satisfies the low memory footprint and high-throughput requirements for real-time sorting on embedded devices.
Supplementary Materials
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Carrizo García, C.; Barfuss, M.H.J.; Sehr, E.M.; Barboza, G.E.; Samuel, R.; Moscone, E.A.; Ehrendorfer, F. Phylogenetic Relationships, Diversification and Expansion of Chili Peppers (Capsicum, Solanaceae). Ann. Bot. 2016, 118, 35–51. [Google Scholar] [CrossRef] [PubMed]
- Barboza, G.E.; García, C.C.; de Bem Bianchetti, L.; Romero, M.V.; Scaldaferro, M. Monograph of Wild and Cultivated Chili Peppers (Capsicum L., Solanaceae). PhytoKeys 2022, 200, 1–423. [Google Scholar] [CrossRef] [PubMed]
- Cao, Y.; Zhang, K.; Yu, H.; Chen, S.; Xu, D.; Zhao, H.; Zhang, Z.; Yang, Y.; Gu, X.; Liu, X.; et al. Pepper Variome Reveals the History and Key Loci Associated with Fruit Domestication and Diversification. Mol. Plant 2022, 15, 1744–1758. [Google Scholar] [CrossRef] [PubMed]
- Du, H.; Yang, J.; Chen, B.; Zhang, X.; Zhang, J.; Yang, K.; Geng, S.; Wen, C. Target Sequencing Reveals Genetic Diversity, Population Structure, Core-SNP Markers, and Fruit Shape-Associated Loci in Pepper Varieties. BMC Plant Biol. 2019, 19, 578. [Google Scholar] [CrossRef] [PubMed]
- Sánchez-Toledano, B.I.; Cuevas-Reyes, V.; Kallas, Z.; Zegbe, J.A. Preferences in “Jalapeño” Pepper Attributes: A Choice Study in Mexico. Foods 2021, 10, 3111. [Google Scholar] [CrossRef] [PubMed]
- Lillywhite, J.; Tso, S. Consumers within the Spicy Pepper Supply Chain. Agronomy 2021, 11, 2040. [Google Scholar] [CrossRef]
- Lillywhite, J.; Robinson, C. Understanding Chile Pepper Consumers’ Preferences: A Discrete Choice Experiment. Agriculture 2023, 13, 1792. [Google Scholar] [CrossRef]
- Que, J.; Liu, L.; Gao, Y.; Tang, Q.; Liu, R.; Liao, W.; Huang, R.; Zhang, D.; Chen, S.; Peng, J. Metabolomic Profiling of Pepper Germplasm Resources and Its Correlation with Color, Shape, and Pungency Traits. Food Chem. X 2025, 30, 102905. [Google Scholar] [CrossRef] [PubMed]
- Zhu, Q.; Deng, L.; Chen, J.; Rodríguez, G.R.; Sun, C.; Chang, Z.; Yang, T.; Zhai, H.; Jiang, H.; Topcu, Y.; et al. Redesigning the Tomato Fruit Shape for Mechanized Production. Nat. Plants 2023, 9, 1659–1674. [Google Scholar] [CrossRef] [PubMed]
- Huang, H.; Huang, T.; Li, Z.; Lyu, S.; Hong, T. Design of Citrus Fruit Detection System Based on Mobile Platform and Edge Computer Device. Sensors 2021, 22, 59. [Google Scholar] [CrossRef] [PubMed]
- Huynh, Q.-K.; Nguyen, C.-N.; Vo-Nguyen, H.-P.; Tran-Nguyen, P.L.; Le, P.-H.; Le, D.-K.-L.; Nguyen, V.-C. Crack Identification on the Fresh Chilli (Capsicum) Fruit Destemmed System. J. Sens. 2021, 2021, 8838247. [Google Scholar] [CrossRef]
- Barbosa, M.d.O.; Aguiar, F.P.L.; Sousa, S.d.S.; Cordeiro, L.d.S.; Nääs, I.d.A.; Okano, M.T. YOLOv8m for Automated Pepper Variety Identification: Improving Accuracy with Data Augmentation. Appl. Sci. 2025, 15, 7024. [Google Scholar] [CrossRef]
- Wu, C.; Wang, J.; Yang, Y.; Li, D.; Li, X.; Yan, R. Analysis of the development status and mechanization trend of cash crop industry in China. J. Chin. Agric. Mech. 2024, 45, 1–13. [Google Scholar] [CrossRef]
- Zhu, Y.; Zhang, Y.; Piao, H. Does Agricultural Mechanization Improve Agricultural Environment Efficiency? Evidence from China’s Planting Industry. Environ. Sci. Pollut. Res. Int. 2022, 29, 53673–53690. [Google Scholar] [CrossRef] [PubMed]
- Li, X.; Zhu, M. The Role of Agricultural Mechanization Services in Reducing Pesticide Input: Promoting Sustainable Agriculture and Public Health. Front. Public Health 2023, 11, 1242346. [Google Scholar] [CrossRef] [PubMed]
- Li, Z.; Zhao, H.; Jing, Z.; Zhao, Z.; Wang, M.; Gong, M.; Wu, X.; He, Z.; Liao, J.; Liu, M.; et al. Recent Advances in Pepper Fruit Glossiness. Genes 2025, 16, 1319. [Google Scholar] [CrossRef] [PubMed]
- Ali, T.; Rehman, S.U.; Ali, S.; Mahmood, K.; Obregon, S.A.; Iglesias, R.C.; Khurshaid, T.; Ashraf, I. Smart Agriculture: Utilizing Machine Learning and Deep Learning for Drought Stress Identification in Crops. Sci. Rep. 2024, 14, 30062. [Google Scholar] [CrossRef] [PubMed]
- Yan, J.; Wang, X. Machine Learning Bridges Omics Sciences and Plant Breeding. Trends Plant Sci. 2023, 28, 199–210. [Google Scholar] [CrossRef] [PubMed]
- Xiang, H.; Zou, B.; Tang, L.; Chen, W.; Rao, K.; Liu, Y.; Ma, M.; Yang, Y. Phytoplankton recognition based on residual attention network model. Acta Ecol. Sin. 2021, 41, 6883–6892. [Google Scholar]
- Shang, Y.; Yu, Y.; Wu, G. Plant disease recognition based on mixed attention mechanism. J. Tarim Univ. 2021, 33, 94–103. [Google Scholar]
- Joshi, K.; Yadav, Y.; Hooda, S.; Nandal, R.; Singh, B.; Singh, K.; Tuteja, N.; Gill, R.; Gill, S.S. Classification of Cotton Leaf Disease Using YOLOv8 Based K-Fold Cross Validation Deep Learning Method for Precision Agriculture. Sci. Rep. 2025, 15, 35602. [Google Scholar] [CrossRef] [PubMed]
- Bouhouch, Y.; Esmaeel, Q.; Richet, N.; Barka, E.A.; Backes, A.; Steffenel, L.A.; Hafidi, M.; Jacquard, C.; Sanchez, L. Deep Learning-Based Barley Disease Quantification for Sustainable Crop Production. Phytopathology 2024, 114, 2045–2054. [Google Scholar] [CrossRef] [PubMed]
- Dai, M.; Sun, W.; Wang, L.; Dorjoy, M.M.H.; Zhang, S.; Miao, H.; Han, L.; Zhang, X.; Wang, M. Pepper Leaf Disease Recognition Based on Enhanced Lightweight Convolutional Neural Networks. Front. Plant Sci. 2023, 14, 1230886. [Google Scholar] [CrossRef] [PubMed]
- Jung, M.; Song, J.S.; Shin, A.-Y.; Choi, B.; Go, S.; Kwon, S.-Y.; Park, J.; Park, S.G.; Kim, Y.-M. Construction of Deep Learning-Based Disease Detection Model in Plants. Sci. Rep. 2023, 13, 7331. [Google Scholar] [CrossRef] [PubMed]
- Mohanappriya, K.; Vennila, C.; Keerthivasan, N. AgroDualNet: A Dual Deep Learning-Based Crop Disease Forecasting and Fruit Ripening Detection. Sci. Rep. 2026, 16, 16444. [Google Scholar] [CrossRef] [PubMed]
- Zhao, M.; You, Z.; Chen, H.; Wang, X.; Ying, Y.; Wang, Y. Integrated Fruit Ripeness Assessment System Based on an Artificial Olfactory Sensor and Deep Learning. Foods 2024, 13, 793. [Google Scholar] [CrossRef] [PubMed]
- Nguyen, N.H.; Michaud, J.; Mogollon, R.; Zhang, H.; Hargarten, H.; Leisso, R.; Torres, C.A.; Honaas, L.; Ficklin, S. Rating Pome Fruit Quality Traits Using Deep Learning and Image Processing. Plant Direct 2024, 8, e70005. [Google Scholar] [CrossRef] [PubMed]
- Kumari, A.; Singh, J. Banana and Guava Dataset for Machine Learning and Deep Learning-Based Quality Classification. Data Brief 2024, 57, 111025. [Google Scholar] [CrossRef] [PubMed]
- Zhu, L.; Wang, X.; Fu, H.; Feng, Y.; Zhang, J. Fine-grained image classification based on attention mechanism. J. Jilin Univ. (Sci. Ed.) 2023, 61, 371–376. [Google Scholar] [CrossRef]
- Yuan, P.; Ding, Y.; Xu, H. Fine-grained chrysanthemum phenotype recognition based on deep active learning and CBAM. Trans. Chin. Soc. Agric. Mach. 2024, 55, 258–267. [Google Scholar]
- Li, X.; Pan, J.; Xie, F.; Zeng, J.; Li, Q.; Huang, X.; Liu, D.; Wang, X. Fast and Accurate Green Pepper Detection in Complex Backgrounds via an Improved Yolov4-Tiny Model. Comput. Electron. Agric. 2021, 191, 106503. [Google Scholar] [CrossRef]
- Paul, A.; Machavaram, R.; Ambuj; Kumar, D.; Nagar, H. Smart Solutions for Capsicum Harvesting: Unleashing the Power of YOLO for Detection, Segmentation, Growth Stage Classification, Counting, and Real-Time Mobile Identification. Comput. Electron. Agric. 2024, 219, 108832. [Google Scholar] [CrossRef]
- Lu, W.; Yang, Y.; Yang, L. Fine-Grained Image Classification Method Based on Hybrid Attention Module. Front. Neurorobotics 2024, 18, 1391791. [Google Scholar] [CrossRef] [PubMed]
- Zhang, Y. A Fine-Grained Image Classification and Detection Method Based on Convolutional Neural Network Fused with Attention Mechanism. Comput. Intell. Neurosci. 2022, 2022, 2974960. [Google Scholar] [CrossRef] [PubMed]
- Huang, G.; Liu, Z.; van der Maaten, L.; Weinberger, K.Q. Densely Connected Convolutional Networks. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2018. [Google Scholar]
- Woo, S.; Park, J.; Lee, J.-Y.; Kweon, I.S. CBAM: Convolutional Block Attention Module. In Proceedings of the Computer Vision—ECCV 2018; Ferrari, V., Hebert, M., Sminchisescu, C., Weiss, Y., Eds.; Springer International Publishing: Cham, Switzerland, 2018; pp. 3–19. [Google Scholar]
- Li, X.; Zhang, B.; Shen, D. Descriptors and Data Standard for Pepper; China Agriculture Press: Beijing, China, 2006; pp. 12–13. [Google Scholar]
- Srivastava, N.; Hinton, G.; Krizhevsky, A.; Sutskever, I.; Salakhutdinov, R. Dropout: A Simple Way to Prevent Neural Networks from Overfitting. J. Mach. Learn. Res. 2014, 15, 1929–1958. [Google Scholar]
- Lin, M.; Chen, Q.; Yan, S. Network in Network. In Proceedings of the International Conference on Learning Representations (ICLR), Banff, AB, Canada, 14–16 April 2014. [Google Scholar]
- Krogh, A.; Hertz, J. A Simple Weight Decay Can Improve Generalization. In Proceedings of the Advances in Neural Information Processing Systems; Morgan-Kaufmann: San Mateo, CA, USA, 1991; Volume 4. [Google Scholar]
- Szegedy, C.; Vanhoucke, V.; Ioffe, S.; Shlens, J.; Wojna, Z. Rethinking the Inception Architecture for Computer Vision. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 2818–2826. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]
- Simonyan, K.; Zisserman, A. Very Deep Convolutional Networks for Large-Scale Image Recognition. In Proceedings of the International Conference on Learning Representations (ICLR), San Diego, CA, USA, 7–9 May 2015. [Google Scholar]
- Tan, M.; Le, Q.V. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Proceedings of the International Conference on Machine Learning (ICML); PMLR: Long Beach, CA, USA, 2019. [Google Scholar]
- Howard, A.; Sandler, M.; Chu, G.; Chen, L.-C.; Chen, B.; Tan, M.; Wang, W.; Zhu, Y.; Pang, R.; Vasudevan, V.; et al. Searching for MobileNetV3. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 1314–1324. [Google Scholar]
- Ma, N.; Zhang, X.; Zheng, H.-T.; Sun, J. ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 10012–10022. [Google Scholar]
- Mehta, S.; Rastegari, M. MobileViT: Light-Weight, General-Purpose, and Mobile-Friendly Vision Transformer. In Proceedings of the Tenth International Conference on Learning Representations (ICLR), Virtual Event, 25–29 April 2022. [Google Scholar]
- Rumelhart, D.E.; Hinton, G.E.; Williams, R.J. Learning Representations by Back-Propagating Errors. Nature 1986, 323, 533–536. [Google Scholar] [CrossRef]
- Kingma, D.P.; Ba, J. Adam: A Method for Stochastic Optimization. arXiv 2014, arXiv:1412.6980. [Google Scholar]
- Loshchilov, I.; Hutter, F. Decoupled Weight Decay Regularization. In Proceedings of the International Conference on Learning Representations (ICLR), New Orleans, LA, USA, 6–9 May 2019. [Google Scholar]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-Excitation Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; pp. 7132–7141. [Google Scholar]
- Wang, Q.; Wu, B.; Zhu, P.; Li, P.; Zuo, W.; Hu, Q. ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 11534–11542. [Google Scholar]
- Hou, Q.; Zhou, D.; Feng, J. Coordinate Attention for Efficient Mobile Network Design. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 13713–13722. [Google Scholar]
- Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-CAM: Visual Explanations From Deep Networks via Gradient-Based Localization. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 618–626. [Google Scholar]






| Fruit Shape | Original Images | Augmentation Factor | Augmented Images | Training Samples | Validation Samples | Test Samples |
|---|---|---|---|---|---|---|
| Lantern | 152 | 2.63 | 400 | 281 | 59 | 60 |
| Cone | 130 | 3.08 | 400 | 282 | 61 | 57 |
| Horn | 272 | 1.47 | 400 | 287 | 57 | 56 |
| Goat-horn | 254 | 1.57 | 400 | 285 | 59 | 56 |
| Short-finger | 159 | 2.52 | 400 | 283 | 61 | 56 |
| Long-finger | 373 | 1.07 | 400 | 282 | 60 | 58 |
| Linear | 217 | 1.84 | 400 | 283 | 60 | 57 |
| Round | 75 | 5.33 | 400 | 284 | 58 | 58 |
| Component | Specification |
|---|---|
| Operating system | AutoDL (ubuntu22.04) |
| CPU | 25 vCPU Intel(R) Xeon(R) Platinum 8470Q |
| GPU | RTX 5090 (32 GB) |
| CUDA version | 12.8 |
| Python version | 3.12 |
| PyTorch version | 2.8.0 |
| Hyperparameter | Value | Description |
|---|---|---|
| Epochs (Max) | 100 | Maximum number of training epochs |
| Batch_Size | 32 | Mini-batch size for forward propagation |
| Input_Size | 224 × 224 | Standardized spatial resolution of input images |
| Optimizer | SGD | Stochastic Gradient Descent with momentum |
| Initial_Learning_Rate | 1 × 10−3 | Initial learning rate for network parameter updates |
| Learning_Rate_Scheduler | CosineAnnealingLR | Cosine annealing schedule (Tmax = 100) |
| Weight_Decay | 1 × 10−4 | L2 regularization penalty to mitigate overfitting |
| Momentum | 0.9 | Momentum factor to accelerate SGD and escape local minima |
| Patience | 15 | Early stopping patience based on validation F1-score |
| Dropout_Rate | 0.3 | Dropout probability applied after global average pooling |
| Loss_Function | CrossEntropyLoss | Standard classification loss function |
| Label_Smoothing | 0.1 | Smoothing factor to soften hard one-hot target distributions |
| Attention_Module | CBAM | Dual spatial and channel attention mechanism |
| Model | Test Acc (%) | Test F1 (%) | Parameters (M) | Model Size (MB) | Inference Time (ms) | Training Time (mins) | Comprehensive_Score |
|---|---|---|---|---|---|---|---|
| DenseNet-121 | 86.90 | 86.74 | 6.96 | 27.15 | 7.92 | 5.1 | 0.847 |
| EfficientNet-B0 | 83.19 | 83.08 | 4.02 | 15.62 | 3.4 | 1.5 | 0.687 |
| ResNet-50 | 81.44 | 81.42 | 23.52 | 90.05 | 2.49 | 1.6 | 0.539 |
| VGG-16 | 78.17 | 77.65 | 134.29 | 512.3 | 1.38 | 3.8 | 0.102 |
| Optimizer | Test_Recall (%) | Test_F1 (%) | Test_Accuracy (%) | Training_Time (min) |
|---|---|---|---|---|
| SGD_Momentum | 89.40 | 89.34 | 89.52 | 3.59 |
| Adam | 85.27 | 85.21 | 85.37 | 4.29 |
| AdamW | 87.63 | 87.64 | 87.77 | 4.54 |
| Learning_Rate | Test_Recall (%) | Test_F1 (%) | Test_Accuracy (%) |
|---|---|---|---|
| 1 × 10−1 | 63.29 | 62.63 | 63.54 |
| 1 × 10−2 | 83.28 | 83.32 | 83.41 |
| 1 × 10−3 | 88.72 | 88.64 | 88.86 |
| 1 × 10−4 | 85.25 | 85.10 | 85.37 |
| 1 × 10−5 | 82.38 | 82.13 | 82.53 |
| Experiment | LS | Dropout | WD | Overfit_ Gap (%) | Test_ Accuracy (%) | Test_ F1 (%) | Mild_ Robust_F1 (%) | Severe_ Robust_F1 (%) |
|---|---|---|---|---|---|---|---|---|
| Exp0 | 0 | 0 | 0 | 13.39 | 86.90 | 86.70 | 85.47 | 83.25 |
| Exp1 | 0 | 0 | 1 × 10−4 | 13.75 | 86.90 | 86.55 | 85.78 | 78.61 |
| Exp2 | 0.1 | 0 | 1 × 10−4 | 13.68 | 87.77 | 87.54 | 87.53 | 78.05 |
| Exp3 | 0 | 0.3 | 1 × 10−4 | 14.14 | 87.12 | 86.88 | 86.80 | 82.46 |
| Exp4 | 0 | 0.5 | 1 × 10−4 | 13.46 | 89.74 | 89.67 | 85.64 | 76.92 |
| Exp5 | 0 | 0 | 5 × 10−4 | 13.32 | 87.99 | 87.76 | 87.80 | 81.64 |
| Exp6 | 0 | 0 | 1 × 10−4 | 11.37 | 87.55 | 87.42 | 88.13 | 83.14 |
| Exp7 | 0.1 | 0.3 | 5 × 10−4 | 12.12 | 87.99 | 87.78 | 86.36 | 76.11 |
| Exp8 | 0.1 | 0.3 | 1 × 10−4 | 12.63 | 88.43 | 88.11 | 89.53 | 81.29 |
| Attention_Module | Test_Precision (%) | Test_Recall (%) | Test_F1 (%) | Test_Accuracy (%) |
|---|---|---|---|---|
| Baseline | 86.72 | 86.32 | 86.10 | 86.46 |
| CBAM | 90.09 | 89.60 | 89.53 | 89.74 |
| SE | 89.26 | 88.72 | 88.69 | 88.86 |
| CAM | 88.28 | 87.85 | 87.77 | 87.99 |
| SAM | 84.40 | 83.46 | 83.27 | 83.62 |
| ECA | 89.62 | 89.43 | 89.37 | 89.52 |
| CA | 88.74 | 88.05 | 88.00 | 88.21 |
| Experiment | Attention | LS | Dropout | Test_ Precision (%) | Test_ Recall (%) | Test_ F1 (%) | Test_ Accuracy (%) |
|---|---|---|---|---|---|---|---|
| Exp1_Baseline | / | 0.0 | 0.0 | 88.46 | 87.85 | 87.85 | 87.99 |
| Exp2_ + CBAM | CBAM | 0.0 | 0.0 | 88.91 | 88.80 | 88.78 | 88.86 |
| Exp3_ + LS | CBAM | 0.1 | 0.0 | 88.52 | 88.33 | 88.24 | 88.43 |
| Exp4_ + Dropout | CBAM | 0.1 | 0.3 | 90.09 | 89.60 | 89.53 | 89.74 |
| Pepper_Shape | Precision (%) | Recall (%) | F1-Score (%) |
|---|---|---|---|
| Lantern | 0.82 | 0.98 | 0.89 |
| Cone | 0.96 | 0.79 | 0.87 |
| Horn | 0.83 | 0.79 | 0.81 |
| Goat_horn | 0.86 | 1.00 | 0.93 |
| Short_finger | 0.90 | 0.95 | 0.92 |
| Long_finger | 0.84 | 0.73 | 0.78 |
| Linear | 0.98 | 1.00 | 0.99 |
| Round | 1.00 | 0.91 | 0.95 |
| accuracy | 0.90 | 0.90 | 0.90 |
| macro_avg | 0.90 | 0.89 | 0.89 |
| weighted_avg | 0.90 | 0.90 | 0.89 |
| Model | Precision (%) | Recall (%) | F1-Score (%) | Accuracy (%) | Time (ms) |
|---|---|---|---|---|---|
| ShuffleNetV2 | 63.53 | 59.22 | 57.44 | 59.39 | 2.41 |
| MobileNetV3 | 80.29 | 80.45 | 79.96 | 80.57 | 2.46 |
| Swin-Tiny | 84.75 | 84.15 | 83.97 | 84.28 | 5.18 |
| MobileViT-S | 82.65 | 82.38 | 82.24 | 82.53 | 4.14 |
| DenseNet-121 (Proposed) | 90.09 | 89.60 | 89.53 | 89.74 | 7.35 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Li, Z.; Li, Y.; Zhao, H.; Huang, L.; Zhao, Z.; Liao, J.; Wang, M.; Wu, X.; Gong, M.; He, Z.; et al. A Deep Learning Model for Chili Pepper Fruit Shape Classification Using DenseNet-121 and CBAM. Plants 2026, 15, 2103. https://doi.org/10.3390/plants15132103
Li Z, Li Y, Zhao H, Huang L, Zhao Z, Liao J, Wang M, Wu X, Gong M, He Z, et al. A Deep Learning Model for Chili Pepper Fruit Shape Classification Using DenseNet-121 and CBAM. Plants. 2026; 15(13):2103. https://doi.org/10.3390/plants15132103
Chicago/Turabian StyleLi, Zongjun, Yinghua Li, Hu Zhao, Liping Huang, Zengjing Zhao, Jianjie Liao, Meng Wang, Xing Wu, Mingxia Gong, Zhi He, and et al. 2026. "A Deep Learning Model for Chili Pepper Fruit Shape Classification Using DenseNet-121 and CBAM" Plants 15, no. 13: 2103. https://doi.org/10.3390/plants15132103
APA StyleLi, Z., Li, Y., Zhao, H., Huang, L., Zhao, Z., Liao, J., Wang, M., Wu, X., Gong, M., He, Z., Liu, L., & Wang, R. (2026). A Deep Learning Model for Chili Pepper Fruit Shape Classification Using DenseNet-121 and CBAM. Plants, 15(13), 2103. https://doi.org/10.3390/plants15132103

