A Resource-Efficient Hybrid CNN-LSTM Network for Image-Based Bean Leaf Disease Classification
Abstract
1. Introduction
2. Literature Review
2.1. Image-Based Automated Plant Disease Identification
2.2. Image Augmentation in Plant Disease Identification
3. Methodology
3.1. Data Preparation and Augmentation for Custom Lightweight Models
- Brightness: Increase the brightness on a linear scale by factors of 20 and 30. Each scale factor is applied to ≈50% of each dataset class.
- Crop: By ratios of 0.8 and 0.9. Each crop ratio is applied to ≈50% of each dataset class.
- Flip: Horizontally and vertically. Each flip direction is applied to ≈50% of each dataset class.
- Rotation: clockwise and counterclockwise. Each rotation angle is applied to ≈50% of each dataset class.
- Combination: Random combinations of all 4 of the operations above.
| Classes | ||||
|---|---|---|---|---|
| Data Partition |
Angular Leaf Spot | Bean Rust | Healthy | Total |
| Originally downloaded (OD) | 432 (33.36%) | 436 (33.67%) | 427 (32.97%) | 1295 |
| Train (OD) | 345 (33.37%) | 348 (33.66%) | 341 (32.98%) | 1034 |
| Val (OD) | 44 (33.08%) | 45 (33.83%) | 44 (33.08%) | 133 |
| Test (OD) | 43 (33.59%) | 43 (33.59%) | 42 (32.81%) | 128 |
| Unaugmented Train | 304 (33.59%) | 307 (33.92%) | 294 (32.49%) | 905 |
| Unaugmented Val | 60 (30.77%) | 67 (34.36%) | 68 (34.87%) | 195 |
| Unaugmented Test | 68 (34.87%) | 62 (31.79%) | 65 (33.33%) | 195 |
| Brightness | 912 (33.59%) | 921 (33.92%) | 882 (32.49%) | 2715 |
| Crop | 912 (33.59%) | 921 (33.92%) | 882 (32.49%) | 2715 |
| Flip | 912 (33.59%) | 921 (33.92%) | 882 (32.49%) | 2715 |
| Rotation | 912 (33.59%) | 921 (33.92%) | 882 (32.49%) | 2715 |
| Combination | 912 (33.59%) | 921 (33.92%) | 882 (32.49%) | 2715 |
3.2. Custom Deep Learning Architecture
4. Experiments and Results
4.1. Experimental Setup for Custom Lightweight Models
4.2. Results: Bean-CNN vs. Bean-CNN-LSTM
4.3. Results: Effect of Augmentation on Custom Lightweight Models
4.4. Results: The Best-Performing Custom Lightweight Model
4.5. Results: Comparison with Existing Models
4.6. Ablation Experiments
4.6.1. Spatial Traversal Strategy Ablation
- Row-Major (Scan-Line) [Proposed]: Top-to-bottom row scan ().
- Column-Major: Left-to-right column scan ().
- Patch-Wise (): Spatial block grouping ().
4.6.2. Parameter-Matched Model Architecture Analysis
4.6.3. Controlled Data Sub-Sampling and Augmentation Analysis
- Sub-sampled unaugmented runs (25% and 50% of the 905 unaugmented training images): Evaluates minimal sample availability (226 and 452 unaugmented images).
- Full unaugmented dataset (100% of the 905 unaugmented samples): Evaluates baseline performance on the raw unaugmented dataset.
- Small augmented set (2715 “Flip” set): Evaluates the impact of the training data volume on model performance using the same augmented set on which Bean-CNN-LSTM has its best performance.
- Complete larger augmented dataset: Evaluates the complete pipeline incorporating the same geometric and photometric transformations used in the EfficientNet experiments.
- Persistent Architectural Advantage: Bean-CNN-LSTM consistently outperforms Bean-CNN across every run ( at 25% data, at 50% data, at 100% data, on the 2715 augmented set and at the augmented set). This proves that spatial-to-sequential feature integration provides a scale-invariant benefit.
- Justification for Data Augmentation and Parameter Sizing: On limited training subsets (25% to 100% unaugmented and 1275 augmented data), EfficientNet-B7 performances scale from 78.46% to about 96%. It only reaches its peak performance of 99+% when trained on the full augmented dataset, highlighting the heavy data requirement of its 65 M parameters; hence the need for the larger augmented set. Even with the 1275 augmented images, the EfficientNet models could not achieve such high performance, indicating the need for a larger augmented training set.
- Justification for a pre-trained model: Despite the larger dataset, the baseline models (Bean-CNN and Bean-CNN-LSTM) reached their peak performances at around 91% and 96%, respectively, while the EfficientNet models reached 99%. The deep backbone of the EfficientNet models gives them a significant edge over the baseline models, especially when the training dataset is large.
| Architecture | Param. Count | 25% Train—226 Samples | 50% Train—452 Samples | 100% Unaugmented Train—905 Samples | 2715 Augmented Set (Flip) | Augmented Dataset |
|---|---|---|---|---|---|---|
| Bean-CNN | 527.2 K | 68.91%; 68.85% | 76.33%; 75.90% | 83.60%; 83.32% | 88.62%; 88.48% | 91.18%; 90.90% |
| Bean-CNN-LSTM | 155.2 K | 71.79%; 71.25% | 79.41%; 79.10% | 87.62%; 87.36% | 94.44%; 94.34% | 95.74%; 95.62% |
| EfficientNet-B7-FC | 67.8 M | 78.46%; 78.12% | 83.59%; 83.40% | 90.93%; 90.81% | 96.44%; 96.40% | 99.36%; 99.28% |
| EfficientNet-B7-FC | 65.2 M | 78.46%; 78.10% | 83.49%; 83.42% | 90.90%; 90.80% | 95.95%; 95.71% | 99.32%; 99.26% |
5. Discussion
5.1. Implications for Agricultural Resource Management
5.2. Limitations
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Pamela, P.; Mawejje, D.; Ugen, M. Severity of angular leaf spot and rust diseases on common beans in Central Uganda. Uganda J. Agric. Sci. 2014, 15, 63–72. [Google Scholar] [CrossRef]
- Venbrux, M.; Crauwels, S.; Rediers, H. Current and emerging trends in techniques for plant pathogen detection. Front. Plant Sci. 2023, 14, 1120968. [Google Scholar] [CrossRef] [Scilit]
- Mahlein, A.K. Plant disease detection by imaging sensors–parallels and specific demands for precision agriculture and plant phenotyping. Plant Dis. 2016, 100, 241–251. [Google Scholar] [CrossRef] [Scilit]
- Rahunathan, L.; Sivabalaselvamani, D.; Elakkiya, E.; Madhumitha, M.; Kumaresh, K. Recognition of Bean Leaf Diseases Using Neural Network and Machine Learning Techniques. In Proceedings of the 2023 3rd International Conference on Smart Data Intelligence (ICSMDI), Trichy, India, 30–31 March 2023; pp. 520–526. [Google Scholar] [CrossRef] [Scilit]
- Slimani, H. Artificial Intelligence-based Detection of Fava Bean Rust Disease in Agricultural Settings: An Innovative Approach. Int. J. Adv. Comput. Sci. Appl. 2023, 14, 119–128. [Google Scholar] [CrossRef] [Scilit]
- Islam, Z.; Islam, M.; Amanullah, A. A combined deep CNN-LSTM network for the detection of novel coronavirus (COVID-19) using X-ray images. Inform. Med. Unlocked 2020, 20, 100412. [Google Scholar] [CrossRef] [Scilit]
- Donahue, J.; Hendricks, L.A.; Rohrbach, M.; Venugopalan, S.; Guadarrama, S.; Saenko, K.; Darrell, T. Long-Term Recurrent Convolutional Networks for Visual Recognition and Description. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 677–691. [Google Scholar] [CrossRef] [Scilit]
- Önler, E. Feature fusion based artificial neural network model for disease detection of bean leaves. Electron. Res. Arch. 2023, 31, 2409–2427. [Google Scholar] [CrossRef] [Scilit]
- Elfatimi, E.; Eryigit, R.; Elfatimi, L. Beans Leaf Diseases Classification Using MobileNet Models. IEEE Access 2022, 10, 9471–9482. [Google Scholar] [CrossRef] [Scilit]
- Tan, M.; Le, Q. Efficientnet: Rethinking model scaling for convolutional neural networks. In Proceedings of the International Conference on Machine Learning, Long Beach, CA, USA, 9–15 June 2019; pp. 6105–6114. [Google Scholar]
- Jian, Z.; Wei, Z. Support vector machine for recognition of cucumber leaf diseases. In Proceedings of the 2010 2nd International Conference on Advanced Computer Control, Shenyang, China, 27–29 March 2010; Volume 5, pp. 264–266. [Google Scholar] [CrossRef] [Scilit]
- Lu, Y.; Yi, S.; Zeng, N.; Liu, Y.; Zhang, Y. Identification of rice diseases using deep convolutional neural networks. Neurocomputing 2017, 267, 378–384. [Google Scholar] [CrossRef] [Scilit]
- Geetharamani, G.; Arun Pandian, J. Identification of plant leaf diseases using a nine-layer deep convolutional neural network. Comput. Electr. Eng. 2019, 76, 323–338, Correction in Comput. Electr. Eng. 2019, 78, 536. https://doi.org/10.1016/j.compeleceng.2019.04.011.. [Google Scholar] [CrossRef] [Scilit]
- Mohanty, S.P. Using Deep Learning for Image-Based Plant Disease Detection. Front. Plant Sci. 2016, 7, 1419. [Google Scholar] [CrossRef] [Scilit]
- Patil, M.A.; Manohar, M. Plant Leaf Disease Classification Using Optimal Tuned Hybrid LSTM-CNN Model. SN Comput. Sci. 2023, 4, 710. [Google Scholar] [CrossRef] [Scilit]
- Devi, E.; Gopi, S.; Padmavathi, U.; Arumugam, S.R.; Premnath, S.; Muralitharan, D. Plant Disease Classification using CNN-LSTM Techniques. In Proceedings of the 2023 5th International Conference on Smart Systems and Inventive Technology (ICSSIT), Tirunelveli, India, 23–25 January 2023; pp. 1225–1229. [Google Scholar] [CrossRef] [Scilit]
- Haque, M.A.; Deb, C.K.; Gole, P.; Karmakar, S.; Dheeraj, A.; Din Shah, M.U.; Dutta, S.; Kumar, M.K.P.; Marwaha, S. An enhanced vision transformer network for efficient and accurate crop disease detection. Expert Syst. Appl. 2025, 283, 127743. [Google Scholar] [CrossRef] [Scilit]
- Abade, A.; Ferreira, P.A.; de Barros Vidal, F. Plant diseases recognition on images using convolutional neural networks: A systematic review. Comput. Electron. Agric. 2021, 185, 106125. [Google Scholar] [CrossRef] [Scilit]
- Singla, S.; Gupta, R. Deep Learning based Bean Leaf Lesion Classification utilizing EfficientNetV2-S. In Proceedings of the 2024 8th International Conference on I-SMAC (IoT in Social, Mobile, Analytics and Cloud) (I-SMAC), Kirtipur, Nepal, 3–5 October 2024; pp. 1394–1399. [Google Scholar] [CrossRef] [Scilit]
- Rodríguez-Lira, D.C.; Córdova-Esparza, D.M.; Álvarez Alvarado, J.M.; Romero-González, J.A.; Terven, J.; Rodríguez-Reséndiz, J. Comparative Analysis of YOLO Models for Bean Leaf Disease Detection in Natural Environments. AgriEngineering 2024, 6, 4585–4603. [Google Scholar] [CrossRef] [Scilit]
- Hohman, F.; Kery, M.B.; Ren, D.; Moritz, D. Model Compression in Practice: Lessons Learned from Practitioners Creating On-device Machine Learning Experiences. In Proceedings of the CHI ’24: CHI Conference on Human Factors in Computing Systems, Honolulu, HI, USA, 11–16 May 2024; pp. 1–18. [Google Scholar] [CrossRef] [Scilit]
- Sun, H.; Xu, H.; Liu, B.; He, D.; He, J.; Zhang, H.; Geng, N. MEAN-SSD: A novel real-time detector for apple leaf diseases using improved light-weight convolutional neural networks. Comput. Electron. Agric. 2021, 189, 106379. [Google Scholar] [CrossRef] [Scilit]
- Arsenovic, M.; Karanovic, M.; Sladojevic, S.; Anderla, A.; Stefanovic, D. Solving Current Limitations of Deep Learning Based Approaches for Plant Disease Detection. Symmetry 2019, 11, 939. [Google Scholar] [CrossRef] [Scilit]
- Yang, S.; Xiao, W.; Zhang, M.; Guo, S.; Zhao, J.; Shen, F. Image Data Augmentation for Deep Learning: A Survey. arXiv 2023, arXiv:2204.08610. [Google Scholar] [CrossRef] [Scilit]
- Taylor, L.; Nitschke, G. Improving Deep Learning with Generic Data Augmentation. In Proceedings of the 2018 IEEE Symposium Series on Computational Intelligence (SSCI), Bangalore, India, 18–21 November 2018; pp. 1542–1547. [Google Scholar] [CrossRef] [Scilit]
- Shorten, C.; Khoshgoftaar, T.M. A survey on Image Data Augmentation for Deep Learning. J. Big Data 2019, 6, 60. [Google Scholar] [CrossRef] [Scilit]
- AI-Lab-Makerere/ibean. 2020. Available online: https://github.com/AI-Lab-Makerere/ibean/ (accessed on 1 July 2024).
- Muimba-Kankolongo, A. Food Crop Production by Smallholder Farmers in Southern Africa; Academic Press: San Diego, CA, USA, 2018. [Google Scholar]
- Deng, L.; Platt, J.C. Ensemble deep learnig for speech recognition. Interspeech 2014, 1. [Google Scholar] [CrossRef] [Scilit]
- Sainath, T.N.; Vinyals, O.; Senior, A.; Sak, H. Convolutional, Long Short-Term Memory, fully connected Deep Neural Networks. In Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), South Brisbane, Australia, 19–24 April 2015; pp. 4580–4584. [Google Scholar] [CrossRef] [Scilit]
- Ercolano, G.; Rossi, S. Combining CNN and LSTM for activity of daily living recognition with a 3D matrix skeleton representation. Intell. Serv. Robot. 2021, 14, 175–185. [Google Scholar] [CrossRef] [Scilit]
- Visin, F.; Kastner, K.; Cho, K.; Matteucci, M.; Courville, A.; Bengio, Y. ReNet: A Recurrent Neural Network Based Alternative to Convolutional Networks. arXiv 2015, arXiv:1505.00393. [Google Scholar] [CrossRef] [Scilit]
- Van Den Oord, A.; Kalchbrenner, N.; Kavukcuoglu, K. Pixel recurrent neural networks. In Proceedings of the ICML’16: 33rd International Conference on International Conference on Machine Learning—Volume 48, New York, NY, USA, 19–24 June 2016; pp. 1747–1756. [Google Scholar] [CrossRef]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image is Worth 16 × 16 Words: Transformers for Image Recognition at Scale. arXiv 2021, arXiv:2010.11929. [Google Scholar] [CrossRef] [Scilit]
- Simonyan, K.; Zisserman, A. Very Deep Convolutional Networks for Large-Scale Image Recognition. In Proceedings of the 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, 7–9 May 2015. [Google Scholar]
- Zhang, X.; Han, N.; Zhang, J. Comparative analysis of VGG, ResNet, and GoogLeNet architectures evaluating performance, computational efficiency, and convergence rates. Appl. Comput. Eng. 2024, 44, 172–181. [Google Scholar] [CrossRef] [Scilit]
- Nixon, M.S.; Aguado, A.A. Feature Extraction and Image Processing for Computer Vision; Academic Press: London, UK, 2020. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification. In Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 7–13 December 2015; pp. 1026–1034. [Google Scholar] [CrossRef] [Scilit]
- Hochreiter, S.; Schmidhuber, J. Long Short-Term Memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [Scilit]
- Gers, F.A.; Schmidhuber, J.; Cummins, F. Learning to Forget: Continual Prediction with LSTM. Neural Comput. 2000, 12, 2451–2471. [Google Scholar] [CrossRef] [Scilit]
- Tharwat, A. Classification assessment methods. Appl. Comput. Inform. 2021, 17, 168–192. [Google Scholar] [CrossRef] [Scilit]
- Chicco, D.; Jurman, G. The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genom. 2020, 21, 6. [Google Scholar] [CrossRef] [Scilit]
- Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 618–626. [Google Scholar] [CrossRef] [Scilit]
- Goodfellow, I.; Bengio, Y.; Courville, A. Deep Learning; MIT Press: Cambridge, MA, USA, 2016. [Google Scholar]
- Kingma, D.P.; Ba, J. Adam: A method for stochastic optimization. arXiv 2014, arXiv:1412.6980. [Google Scholar]
- Abed, S.H.; Al-Waisy, A.S.; Mohammed, H.J.; Al-Fahdawi, S. A modern deep learning framework in robot vision for automated bean leaves diseases detection. Int. J. Intell. Robot. Appl. 2021, 5, 235–251. [Google Scholar] [CrossRef] [Scilit]
- Singh, V.; Chug, A.; Singh, A.P. Classification of beans leaf diseases using fine tuned cnn model. Procedia Comput. Sci. 2023, 218, 348–356. [Google Scholar] [CrossRef] [Scilit]
- Sunyoto, A.; Ariatmanto, D.; Noviyanto. Innovative Solutions for Bean Leaf Disease Detection Using Deep Learning. In Proceedings of the 2024 IEEE International Conference on Artificial Intelligence and Mechatronics Systems (AIMS), Virtual, 21–23 February 2024; pp. 1–5. [Google Scholar]
- Jain, E.; Aneja, A. Automated Detection and Classification of Bean Leaf Diseases using InceptionV3: A Deep Learning Approach. In Proceedings of the 2025 International Conference on Electronics and Renewable Systems (ICEARS), Tuticorin, India, 11–13 February 2025; pp. 1890–1895. [Google Scholar] [CrossRef] [Scilit]
- Karthik, R.; Aswin, R.; Geetha, K.S.; Suganthi, K. An Explainable Deep Learning Network With Transformer and Custom CNN for Bean Leaf Disease Classification. IEEE Access 2025, 13, 38562–38573. [Google Scholar] [CrossRef] [Scilit]
- Kahatapitiya, K.; Rodrigo, R. Exploiting the Redundancy in Convolutional Filters for Parameter Reduction. In Proceedings of the 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), Virtual, 5–9 January 2021; pp. 1409–1419. [Google Scholar] [CrossRef] [Scilit]
- Barbedo, J.G.A. A review on the main challenges in automatic plant disease identification based on visible range images. Biosyst. Eng. 2016, 144, 52–60. [Google Scholar] [CrossRef] [Scilit]
- Fenu, G.; Malloci, F.M. DiaMOS Plant: A Dataset for Diagnosis and Monitoring Plant Disease. Agronomy 2021, 11, 2107. [Google Scholar] [CrossRef] [Scilit]












| Parameter | Bean-CNN | Bean-CNN-LSTM |
|---|---|---|
| Learning rate | 1 | 1 |
| Epochs | 40 | 40 |
| Optimiser | Adam | Adam |
| Batch size | 32 | 32 |
| Final Conv + Pooling dim. | ||
| Reshape dim. | 6400 (flattened) | |
| Traversal order | – | Row-major (scan-line) |
| Sequential layer (1st and 2nd) | 64 units and 8 units | 64 units and 16 units |
| Dropout rate | 0.3 | 0.3 |
| Bean-CNN | ||||
|---|---|---|---|---|
| Training Set | Accuracy | Loss | F1 Score | MCC |
| Unaugmented | 0.8410 | 0.5180 | 0.8410 | 0.7614 |
| Brightness | 0.8769 | 0.5180 | 0.8751 | 0.8163 |
| Combination | 0.8872 | 0.6561 | 0.8882 | 0.8329 |
| Crop | 0.8821 | 0.4597 | 0.8844 | 0.8267 |
| Flip | 0.8872 | 0.5089 | 0.8867 | 0.8307 |
| Rotation | 0.8923 | 0.5130 | 0.8913 | 0.8391 |
| Bean-CNN-LSTM | ||||
| Training set | Accuracy | Loss | F1 Score | MCC |
| Unaugmented | 0.8769 | 0.3773 | 0.8737 | 0.8173 |
| Brightness | 0.9077 | 0.3902 | 0.9074 | 0.8615 |
| Combination | 0.9231 | 0.4116 | 0.9234 | 0.8856 |
| Crop | 0.9333 | 0.2757 | 0.9332 | 0.9000 |
| Flip | 0.9436 | 0.1769 | 0.9438 | 0.9164 |
| Rotation | 0.9179 | 0.3398 | 0.9168 | 0.8776 |
| Bean-CNN | ||||||
|---|---|---|---|---|---|---|
| Accuracy | F1 Score | MCC | ||||
| Training Set | Mean | Std. | Mean | Std. | Mean | Std. |
| Unaugmented | 0.7808 | 0.0403 | 0.7778 | 0.0443 | 0.6776 | 0.0562 |
| Brightness | 0.8026 | 0.0444 | 0.8021 | 0.0446 | 0.7092 | 0.0622 |
| Combination | 0.8525 | 0.0174 | 0.8518 | 0.0177 | 0.7855 | 0.0214 |
| Crop | 0.8460 | 0.0170 | 0.8492 | 0.0174 | 0.7732 | 0.0246 |
| Flip | 0.8540 | 0.0214 | 0.8533 | 0.0220 | 0.7849 | 0.0314 |
| Rotation | 0.8475 | 0.0266 | 0.8470 | 0.0276 | 0.7740 | 0.0370 |
| Bean-CNN-LSTM | ||||||
| Accuracy | F1 Score | MCC | ||||
| Training set | Mean | Std. | Mean | Std. | Mean | Std. |
| Unaugmented | 0.8275 | 0.0383 | 0.8274 | 0.0381 | 0.7470 | 0.0545 |
| Brightness | 0.8624 | 0.0308 | 0.8636 | 0.0299 | 0.7968 | 0.0436 |
| Combination | 0.8770 | 0.0205 | 0.8760 | 0.0211 | 0.8165 | 0.0341 |
| Crop | 0.8875 | 0.0271 | 0.8744 | 0.0802 | 0.8331 | 0.0401 |
| Flip | 0.8865 | 0.0250 | 0.8854 | 0.0263 | 0.8322 | 0.0362 |
| Rotation | 0.8744 | 0.0183 | 0.8740 | 0.0183 | 0.8136 | 0.0267 |
| Training Set | Accuracy | F1 Score | MCC |
|---|---|---|---|
| Unaugmented | |||
| Brightness | |||
| Combination | |||
| Crop | |||
| Flip | |||
| Rotation |
| Class | Precision (%) | Recall (%) | F1 Score (%) |
|---|---|---|---|
| Angular leaf spot | 100.00 | 91.18 | 95.38 |
| Bean rust | 89.23 | 93.55 | 91.34 |
| Healthy | 94.12 | 98.46 | 96.24 |
| Overall Acc. (%) | 94.36 | ||
| Overall F1 (%) | 94.38 | ||
| Class | Precision (%) | Recall (%) | F1 Score (%) |
|---|---|---|---|
| Angular leaf spot | 97.73 | 100.00 | 98.85 |
| Bean rust | 100.00 | 97.67 | 98.82 |
| Healthy | 100.00 | 100.00 | 100.00 |
| Overall Acc. (%) | 99.22 | ||
| Overall F1 (%) | 99.22 | ||
| Source | Method | Acc.; F1 Score (%) | Train/Test Size | Augmentation |
|---|---|---|---|---|
| Abed et al., 2021 [46] | DenseNet121 | 91.02; – | 9324/259 | flip and rotation |
| Elfatimi et al., 2022 [9] | MobileNetV2 | 92.97; 92.94 | 1034/128 | none |
| Singh et al., 2023 [47] | EfficientnetB6 | 91.74; – | 1034/128 | none |
| Önler 2023 [8] | MobileNetV2 | 99.24; – | 1034/128 | blur, brightness, crop, flip, and rotation |
| Sunyoto et al., 2024 [48] | DenseNet121 | 96.90; 97.00 | 1034/128 | rotation and zoom |
| Jain & Aneja 2025 [49] | InceptionV3 | 91.00; 91.00 | 693/149 | brightness, flip, rotation, and zoom |
| Karthik et al., 2025 [50] | Transformer+CNN | 97.66; 97.67 | 13,442/128 | flip, rotation, and translation |
| Ours | Bean-CNN-LSTM | 94.36; 94.38 | 41,360/128 | blur, brightness, flip, rotation, and scaling |
| EfficientNetB7+FC | 99.22; 99.22 | as above | as above | |
| EfficientNetB7+LSTM | 99.22; 99.22 | as above | as above |
| Traversal Strategy | Tensor Shape | Test Accuracy (%) | Test F1 (%) | Test MCC (%) |
|---|---|---|---|---|
| Row-major (scan-line) [proposed] | 89.74 | 89.64 | 85.11 | |
| Column-major | 88.72 | 88.70 | 83.32 | |
| Patch-wise () | 86.67 | 86.58 | 80.15 |
| Architecture | Classification Head | Param. Count | Model Size (MB) | Test Accuracy (%) | Test F1 (%) | Test MCC (%) |
|---|---|---|---|---|---|---|
| Bean-CNN (Baseline) | Standard FC | 527,171 | 6.11 | 84.62 | 84.48 | 77.24 |
| Bean-CNN-Compact (Control) | Bottleneck FC | 155,219 | 1.86 | 82.56 | 82.48 | 74.12 |
| Bean-CNN-LSTM (Proposed) | Spatial-to-Seq LSTM | 155,219 | 1.86 | 89.74 | 89.64 | 85.11 |
| Models | Model Size (MB) | Total Params | Trainable Params | Best Accuracy/F1 (%) |
|---|---|---|---|---|
| Bean-CNN | 6.11 | 527,171 | 527,171 | 89.23/89.13 |
| Bean-CNN-LSTM | 1.86 | 155,219 | 155,219 | 94.36/94.38 |
| EfficientNetB7+FC | 260.56 | 67,825,595 | 4,039,035 | 99.22/99.22 |
| EfficientNetB7+LSTM | 250.20 | 65,165,139 | 1,377,667/52,844,795 | 99.22/99.22 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Rhee, H.J.; Akinyemi, J.D. A Resource-Efficient Hybrid CNN-LSTM Network for Image-Based Bean Leaf Disease Classification. J. Imaging 2026, 12, 468. https://doi.org/10.3390/jimaging12100468
Rhee HJ, Akinyemi JD. A Resource-Efficient Hybrid CNN-LSTM Network for Image-Based Bean Leaf Disease Classification. Journal of Imaging. 2026; 12(10):468. https://doi.org/10.3390/jimaging12100468
Chicago/Turabian StyleRhee, Hye Jin, and Joseph Damilola Akinyemi. 2026. "A Resource-Efficient Hybrid CNN-LSTM Network for Image-Based Bean Leaf Disease Classification" Journal of Imaging 12, no. 10: 468. https://doi.org/10.3390/jimaging12100468
APA StyleRhee, H. J., & Akinyemi, J. D. (2026). A Resource-Efficient Hybrid CNN-LSTM Network for Image-Based Bean Leaf Disease Classification. Journal of Imaging, 12(10), 468. https://doi.org/10.3390/jimaging12100468

