Next Article in Journal
Modern Imaging and the Future of Prostate Biopsy: Rethinking Systematic Sampling
Previous Article in Journal
Correction: Rathee et al. SIFT-SNN for Traffic-Flow Infrastructure Safety: A Real-Time Context-Aware Anomaly Detection Framework. J. Imaging 2026, 12, 64
Previous Article in Special Issue
DGF-YOLO: A Degradation-Guided Feature Enhancement Method for Small-Scale Pedestrian Detection in UAV Images
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Resource-Efficient Hybrid CNN-LSTM Network for Image-Based Bean Leaf Disease Classification

by
Hye Jin Rhee
and
Joseph Damilola Akinyemi
*
Department of Computer Science, University of York, York YO10 5GH, UK
*
Author to whom correspondence should be addressed.
J. Imaging 2026, 12(10), 468; https://doi.org/10.3390/jimaging12100468
Submission received: 12 August 2026 / Revised: 15 September 2026 / Accepted: 21 September 2026 / Published: 24 September 2026
(This article belongs to the Special Issue AI-Driven Image Analysis and Pattern Recognition)

Abstract

Accurate and resource-efficient automated diagnosis is a cornerstone of modern agricultural expert systems. While Convolutional Neural Networks (CNNs) have established benchmarks in plant pathology, their ability to capture long-range spatial dependencies is often limited by standard pooling layers, and their high memory footprint hinders deployment on portable devices. This paper proposes a lightweight hybrid CNN-LSTM system for bean leaf disease classification. By integrating an LSTM layer to model the spatial–sequential relationships within feature maps, our hybrid architecture achieves a 94.36% accuracy and 94.38% F1 score while maintaining an exceptionally small footprint of 1.86 MB, a 70% reduction in size compared to traditional CNN-based systems. Furthermore, we provide a systematic evaluation of image augmentation strategies, demonstrating that tailored transformations are superior to generic combinations for maintaining the integrity of diagnostic patterns. Results on the ibean dataset confirm that the proposed system achieves state-of-the-art F1 scores of 99.22% with EfficientNet-B7+LSTM, providing a potentially robust and scalable framework for real-time agricultural decision support in resource-constrained environments. The code and augmented datasets used in this study are publicly available on GitHub.

1. Introduction

Automated plant disease identification is a crucial smart farming system that helps identify real-time crop infections. In common bean (Phaseolus vulgaris L.) plants, prevalent fungal diseases, such as bean rust and angular leaf spot, compromise the plants’ photosynthetic capacity, which can cause structural damage to the plant and extensive loss of yield. For instance, the production loss in Uganda in 2012 caused by bean rust and angular leaf spot was estimated at up to 47% and 55%, respectively [1]. Visual detection remains the essential first step in identifying symptoms, and subsequent microscopic or biochemical component analysis confirms the pathogens. This approach can be expensive and requires processing time [2]. Automated disease identification has increasingly been considered an attractive means to achieve timely agricultural disease management [3].
The practical management of common bean crops requires diagnostic tools that are not only accurate but also economically viable for smallholder farmers. While traditional biochemical analysis is a definitive ’gold standard,’ the associated costs and logistical delays often render it inaccessible for real-time field management [2]. By developing ultra-lightweight architectures that maintain high diagnostic precision while minimising computational overhead, this research provides a potential for Edge-AI deployment. Such systems enable immediate, on-site decision-making, allowing farmers to initiate localised treatments before a minor outbreak scales into the severe yield losses observed in regions like Uganda [1]. This shifts the paradigm from reactive agricultural management to a proactive, data-driven approach that is sustainable even in resource-constrained environments.
The development of automated expert systems for agriculture is increasingly essential as global food security faces threats from climate-driven disease outbreaks. Conventional diagnostic processes rely on the availability of human experts, which is often a bottleneck in large-scale farming or remote regions. An effective agricultural expert system must bridge the gap between high-level diagnostic accuracy and operational feasibility on edge devices. This requires a knowledge-based approach to architecture design, where models are not only deep but also resource-efficient. By integrating sequential logic into spatial feature extraction, an expert system can better mimic the human expert’s ability to contextualise localised symptoms within the broader structural geometry of a leaf, thereby providing more reliable decision support for crop management.
Artificial intelligence (AI) techniques, such as machine learning (ML) and deep learning (DL), have been instrumental in automated bean leaf disease detection. Traditional ML algorithms, such as support vector machines (SVMs) [4], have been used to classify bean leaf disease. DL methods, such as CNNs, have been even more successful in accurately detecting and classifying diseases [5]. Nonetheless, exploring alternative tools beyond conventional CNNs is in high demand. Leveraging CNN with LSTM has shown encouraging results in various computer vision tasks, such as medical imaging [6] and human motion recognition [7]. However, this design has not been extensively explored for common bean leaf disease. Furthermore, previous studies that commonly addressed data shortages by implementing data augmentation [8,9] have yet to systematically review the impact of various schemes. Notwithstanding, the suitability of a technique is highly subject to various aspects of ML operations.
Our primary experiments involved additional data generation using augmentation techniques and constructing two DL architectures: Bean-CNN (baseline CNN) and Bean-CNN-LSTM (hybrid CNN-LSTM). We trained these models on original and augmented images, and then compared their performance across different augmentation strategies. In additional experiments, we further expanded the data and trained fine-tuned EfficientNet [10] models using this hybrid design concept.
The rest of the paper is organised as follows. Section 2 reviews related work. Section 3 describes the methodology employed in our experiments. Section 4 and Section 5 present the experimental results and discussion. The final section summarises our findings.

2. Literature Review

2.1. Image-Based Automated Plant Disease Identification

With the growing demand for smart farming systems, researchers have proposed techniques for effective image-based plant disease identification. Early research predominantly used traditional ML algorithms combined with feature engineering, as an SVM-based cucumber leaf disease classifier showed the best result with the radial basis function (RBF) [11]. Like many other image classification tasks, adopting CNNs for this problem has significantly improved prediction quality. For example, Lu et al. [12] showed that their compact CNN model outperformed an SVM-based model in their rice disease classification. Naturally, applying DL methods has become a mainstream approach, as Geetharamani and Pandian [13] proposed a nine-layered CNN trained on the Plant Village dataset [14]. More recently, Patil and Manohar proposed an LSTM-CNN tomato leaf disease classifier [15]. A similar attempt by Devi et al. [16] showed superior results for their CNN-LSTM plant disease classifier to existing models. Haque et al. [17] proposed a Vision transformer (ViT) with triplet multi-head attention to tackle disease detection and achieved a 97.99% accuracy on plants and apples.
Due to data scarcity, the common bean plant is less studied among agricultural crops [18]. In most bean disease research, the transfer learning approach using pre-trained models, such as MobileNet [9] or EfficientNet [19], is common. A more recent attempt using YOLO (You Only Look Once) [20] was successful in detecting the damage by bean leaf beetles. The customisation of new deep learning architectures is starting to gain momentum, as an article [8] proposed a combined model using the histogram of oriented gradients (HOGs) and CNN for bean leaf disease identification.
The main drawback of previous attempts is that these complex models have limited usability in real-world situations under hardware constraints [21]. For real-world deployment, we explored the potential of a hybrid CNN-LSTM as a more efficient deep learning architecture than existing models, and this is rarely attempted for the classification of common bean disease.

2.2. Image Augmentation in Plant Disease Identification

Data augmentation is commonly applied to image classification problems to improve diversity in training data. The classic model-free approach includes geometric and photometric transformations. Geometric transformations, including rotation, cropping, and flipping, manipulate the geometric primitives of an image. On the other hand, photometric transformations adjust the colour space information, such as RGB (Red–Green–Blue) and HSB (Hue–Saturation–Brightness). The most popular method in plant disease research is the multi-technique combination due to practicality and general consensus over its effectiveness. This strategy involves applying several transformations to the image data proportionally or randomly. The application of this technique in single-crop [22], and multi-crop classifiers [23] demonstrated its effectiveness in generating large training datasets. Similarly to other plant disease studies, this combination technique is a popular approach in bean disease research, typically without reviewing its efficacy. However, various articles [24,25,26] have suggested that the effectiveness of any augmentation can be highly dependent upon the context, such as the type of task or the model. Therefore, we attempt to review several techniques for our custom model development.

3. Methodology

In our initial experiments, we constructed lightweight, custom-designed deep learning architectures built from scratch. For these experiments, we expanded the training samples to three times the original training data by implementing data augmentation techniques. We trained and evaluated our models to contrast their performance to determine which architecture and augmentation techniques were superior to others.
Our experiments were carried out on a hardware environment with an Intel Xeon CPU running at 2.2 GHz and an NVIDIA Tesla 4 GPU. We used OpenCV (v 4.10) and Scikit-image (v 0.23.2) packages for augmentation and TensorFlow (v 2.17) for model construction.

3.1. Data Preparation and Augmentation for Custom Lightweight Models

The ibean data [27] is created by Makerere AI Research Lab and contains field-collected bean leaf images showing two disease classes and one healthy class, as shown in Figure 1. Whilst angular leaf spot displays distinctive, angular-shaped damage [28], bean rust exhibits irregular light haloes, which may progress to darker blotchy patterns.
The originally downloaded (OD) ibean dataset contained 1295 images (excluding one corrupted image file), divided into 1034 training images, 133 validation images, and 128 test images. We first combined all splits in the OD dataset into one, shuffled it, and then split it into 70% (905 samples) for training, 15% (195 samples) for validation, and 15% (195 samples) for testing. The new 70% training set is henceforth referred to as the unaugmented training set. Only the unaugmented training set (70%) was augmented; the validation and test subsets did not contain augmented data. This combination and re-splitting of the dataset was carried out to create larger and identical validation and test sets, as seen in Table 1. To obtain the augmented training sets, we applied 5 different augmentation operations to the unaugmented training set. Each augmentation operation was applied to each of the three dataset classes (Angular leaf spot, Bean Rust, and Healthy) using two different parameters to increase variability. Thus, to obtain the five augmented training sets, we applied augmentation operations as follows:
  • Brightness: Increase the brightness on a linear scale by factors of 20 and 30. Each scale factor is applied to ≈50% of each dataset class.
  • Crop: By ratios of 0.8 and 0.9. Each crop ratio is applied to ≈50% of each dataset class.
  • Flip: Horizontally and vertically. Each flip direction is applied to ≈50% of each dataset class.
  • Rotation: 15 ° clockwise and 15 ° counterclockwise. Each rotation angle is applied to ≈50% of each dataset class.
  • Combination: Random combinations of all 4 of the operations above.
Table 1. Dataset distribution: originally downloaded (OD), unaugmented and augmented.
Table 1. Dataset distribution: originally downloaded (OD), unaugmented and augmented.
Classes
Data Partition Angular Leaf
Spot
Bean Rust Healthy Total
Originally downloaded (OD)432 (33.36%)436 (33.67%)427 (32.97%)1295
Train (OD)345 (33.37%)348 (33.66%)341 (32.98%)1034
Val (OD)44 (33.08%)45 (33.83%)44 (33.08%)133
Test (OD)43 (33.59%)43 (33.59%)42 (32.81%)128
Unaugmented Train304 (33.59%)307 (33.92%)294 (32.49%)905
Unaugmented Val60 (30.77%)67 (34.36%)68 (34.87%)195
Unaugmented Test68 (34.87%)62 (31.79%)65 (33.33%)195
Brightness912 (33.59%)921 (33.92%)882 (32.49%)2715
Crop912 (33.59%)921 (33.92%)882 (32.49%)2715
Flip912 (33.59%)921 (33.92%)882 (32.49%)2715
Rotation912 (33.59%)921 (33.92%)882 (32.49%)2715
Combination912 (33.59%)921 (33.92%)882 (32.49%)2715
This resulted in 2715 samples for each augmented training set, all with identical class distributions to that of the training set, as shown in Table 1. In Table 1, the train (OD), val (OD), and test (OD) refer to the data partitions available in the OD dataset, while unaugmented train, unaugmented val, and unaugmented test refer to the unaugmented partitions created from re-splitting. All data splitting, shuffling, and augmentation at this stage were controlled by a random seed of 0. Samples of the effect of each augmentation technique on one image is shown in Figure 2

3.2. Custom Deep Learning Architecture

The idea of integrating a CNN and a recurrent neural network (RNN) was first introduced by Deng and Platt [29] for speech recognition. This approach has been considered efficient [30] and has been shown to be effective for human-activity recognition [31].
The theoretical justification for incorporating an LSTM layer into a computer vision pipeline lies in treating spatial feature maps as a pseudo-temporal sequence, an established paradigm in spatial sequence modelling [32,33]. This is even more recently demonstrated in successful vision literature such as the Vision Transformer [34]. While standard CNNs excel at extracting local hierarchical features through localised convolutional kernels, they often lack an explicit mechanism to capture long-range relational dependencies across the spatial grid without adding deep stacks of pooling or parameter-heavy dense layers. By unrolling the spatial feature grid into a sequence of feature vectors, the LSTM cell acts as a contextual accumulator over adjacent spatial regions. Its gating mechanisms systematically preserve spatial continuity across contiguous rows while modulating transitions across boundaries, allowing the network to model directional Markov dependencies across the leaf geometry. In bean leaf pathology, where lesions like angular leaf spot manifest in sporadic but structurally correlated patterns, this spatial-to-sequential integration contextualises localised symptoms within broader leaf spatial patterns, providing a richer structural representation than isolated fully connected projection layers.
Figure 3 presents our proposed model architecture. This is a VGG-inspired [35] DL network that offers greater flexibility in modifying the fully connected (FC) layers. We selected this stacked architecture because it has the advantage of capturing intricate patterns in an image over residual networks [36]. Small filters are selected for computational efficiency, and padding parameters are used to minimise the boundary effects. We implemented either the stride parameter or max-pooling to reduce the input resolution and improve feature extraction [37]. All convolutional layers employ ReLU combined with the Kaiming (He) initialisation to increase non-linearity and ensure more reliable convergence [38].
To effectively leverage the LSTM’s capability for spatial correlation, we transform the 2D feature maps generated by the final convolutional layer into a structured 1D temporal sequence. Instead of a standard flattening operation, which collapses the spatial hierarchy, the feature maps are reshaped into consecutive time steps that represent a scan-line or patch-wise traversal of the leaf image. This allows the LSTM’s internal gates, specifically the forget ( f t ) and input ( i t ) gates, to learn the contextual dependency between adjacent pixel regions. In agriculture, where disease manifestations such as bean rust spread in sporadic and blotchy clusters, the LSTM acts as a secondary filter that identifies the recurring geometric patterns of pathogens across the entire leaf surface, leading to a potentially improved performance gain over baseline CNNs. Hence, whilst the vectors fed into the densely connected neural network (NN) in Bean-CNN were flattened to preserve spatial information, these vectors were sequentially reshaped in time steps for Bean-CNN-LSTM. An LSTM is a modified variant of RNNs [39] and uses the backpropagation through time (BPTT) technique in its cell-based computations, allowing the gradient update across multiple time steps. The relevant LSTM computations are expressed in the following equations:
f t = σ W f x t + U f h t − 1 + b f
i t = σ W i x t + U i h t − 1 + b i
c ˜ t = tanh W c x t + U c h t − 1 + b c
c t = f t ⊙ c t − 1 + i t ⊙ c ˜ t
o t = σ W o x t + U o h t − 1 + b o
h t = o t ⊙ tanh ( c t )
where f t , i t , o t ∈ R H represent the forget gate [40], input gate, and output gate activation vectors at time step t. c ˜ t ∈ R H is the candidate cell state, and c t ∈ R H is the updated cell memory state vector. h t ∈ R H is the hidden state output vector passed to subsequent layers/steps. W ∗ ∈ R H × D and U ∗ ∈ R H × H denote the input-to-hidden and recurrent hidden-to-hidden weight matrices, respectively. b ∗ ∈ R H represents the bias vectors. σ ( · ) is the element-wise logistic sigmoid function, tanh ( · ) is the hyperbolic tangent activation, and ⊙ denotes the Hadamard (element-wise) product.
To examine the quality of the estimated weights, our proposed models use the softmax and sparse categorical cross-entropy loss functions. The probability calculated by softmax is integrated with categorical cross-entropy to approximate the prediction error. The following equation describes the computation of this loss function L, where v k is the prediction in k classes for the j t h class of the i t h data point and y i is the true distribution.
L = − ∑ i = 1 k y i l o g ( e v i ∑ j = 1 k e v j )
For model evaluation, we collected accuracy, the confusion matrix, the F1 score, and the Matthews correlation coefficient (MCC) for a more balanced analysis [41,42]. The confusion matrix represents the instances of the predicted label against the true label. Correctly predicted instances are true positives (TP) and true negatives (TN). In contrast, false positives (FP) and false negatives (FN) represent erroneous predictions. Based on the confusion matrix, F1 and MCC can be calculated as follows:
P r e c i s i o n = T P T P + F P
R e c a l l = T P T P + F N
F 1 = 2 ∗ P r e c i s i o n ∗ R e c a l l P r e c i s i o n + R e c a l l
M C C = T P ∗ T N − F P ∗ F N ( T F + F P ) ( T P + F N ) ( T N + F P ) ( T N + F N )

4. Experiments and Results

4.1. Experimental Setup for Custom Lightweight Models

To ensure the best performance and a balance between overfitting and underfitting, we optimised each model’s capacity, such as the number of hidden layers and units. Table 2 describes the optimal parameter and hyperparameter values that we found during the iterative fine-tuning process for our custom models. The two values seen for the sequential layer in Table 2 correspond to the number of hidden units for the first and second sequential layers in each custom model.
To evaluate stability and rule out performance artefacts from lucky weight initialisations, training experiments were repeated across 30 independent runs on each of the six training sets (the unaugmented dataset and the five augmented training sets). Each run was governed by a distinct random seed (integers 1 to 30), which controlled both the pseudo-random weight initialisation (using the He initialisation [38]) and mini-batch shuffling order during training. The same fixed sequence of 30 seeds was applied across all experimental conditions (the unaugmented dataset and the five augmented training sets) to ensure rigorous and reproducible comparisons. As only the unaugmented 70% training set was augmented to obtain the five augmented training sets, the test and validation sets remained fixed at 15% each across all runs to maintain consistent evaluation conditions.
We used Gradient-Weighted Class Activation Mapping (Grad-CAM) [43] to visualise the features extracted in the convolutional layers. The weight distribution is represented in a heat map in Figure 4 (in the Viridis scale) and shows that the feature extraction by the custom models becomes more specific in deeper layers.

4.2. Results: Bean-CNN vs. Bean-CNN-LSTM

Figure 5 and Figure 6 visualise the training progress of the best models trained on the unaugmented training set (905 samples). Table 3 describes the best performance across all training sets (the unaugmented and the augmented), which is summarised in Figure 7. Bean-CNN-LSTM consistently outperforms Bean-CNN, and the accuracy varies by 2.56% to 5.64%, depending on the training set. A smaller disparity between accuracy and MCC in Bean-CNN-LSTM implies that this model made fewer false predictions than the Bean-CNN models.
We also analysed the results from all 30 training runs and present summary statistics in Table 4, showing the mean and standard deviation values of the accuracy, F1 score, and MCC. Table 4 shows relatively low standard deviations from the mean (in the order of 10 − 2 ) across all three metrics (accuracy, F1, and MCC), indicating a tight distribution. Across all three metrics, the mean values consistently indicate significant improvement of Bean-CNN-LSTM over Bean-CNN. Also, Figure 8 presents the accuracy values in box-and-whisker plots, showing the key statistics of the distribution. The box plots show relatively tight interquartile ranges (IQR) with few outliers. Together, Table 4 and Figure 8 indicate that the distribution is tight and free of skewness and excessive outlier values. From the box plots, it is clear that the IQR of the accuracies obtained by the CNN-LSTM model consistently ranks higher than that of the baseline CNN model, implying that the hybrid CNN-LSTM is more effective than the baseline CNN, with noticeable gaps.
To further evaluate the statistical stability and significance of the CNN-LSTM model over the baseline CNN model, we performed a paired two-tailed t-test of the metric values (Accuracy, F1 Score, and MCC) across all 30 experiment runs. The results are presented in Table 5, showing high statistical significance generally in the order of 10 − 2 – 10 − 7 . This confirms that the performance gain offered by the CNN-LSTM model is based on statistical significance and mathematical soundness, rather than mere chance or experimental noise.

4.3. Results: Effect of Augmentation on Custom Lightweight Models

Training on the augmented sets, as shown in Table 3, generally increases the initial accuracy to a different extent. Bean-CNN shows the highest accuracy on the rotation set, increasing the initial accuracy by 5.13%. The best Bean-CNN-LSTM improves accuracy by 6.67% on the flip set. Both models on the brightness set show the least increase in accuracy with elevated loss. The multi-technique combination also raises accuracy and loss simultaneously, indicating greater uncertainty in their performance [44].
The average performance results of our custom models also show notable trends. As shown in Figure 9, the crop, flip, rotation, and combination methods equally improve Bean-CNN. The loss improves further with the crop technique, indicating an additional benefit to the performance. For Bean-CNN-LSTM, both the flip and the crop techniques are more effective than others when all metrics are considered. Data analysis of the training sets does not show strong correlations between each training set and performance, although we found that the brightness set has elevated frequency and noise levels, which may have negatively contributed to the performance of all models.
In the box plots of Figure 8 and Figure 9, a few training set runs (such as the crop and flip augmented sets) exhibit overlapping or superimposed outlier points beyond the standard 1.5 × IQR whiskers. These overlapping data points occur because multiple independent experimental runs (30 seeds) converged to identical discrete test metric values (e.g., exact test accuracies due to the finite test set size of 195 samples, where each single sample misclassification changes accuracy in about 0.51% increments). From an optimisation standpoint, these low-performing outliers reflect instances where specific stochastic weight initialisations coincided with aggressive spatial augmentations (such as severe cropping), leading the network into sub-optimal local minima during gradient descent. The fact that these occurrences are rare (yielding superimposed points outside the upper/lower whiskers while the interquartile boxes remain tightly clustered) underscores the overall stability of the models across the vast majority of initialisation seeds.
From a data governance perspective, these results highlight the necessity of principled data curation over exhaustive data expansion. Our findings indicate that while generic augmentation combinations are often viewed as a default for increasing model robustness, they can introduce a form of synthetic noise that degrades the internal consistency of lightweight models. For agricultural datasets, data governance must prioritise transformations that preserve the biological integrity of the specimen. Geometric shifts like flipping and cropping mimic natural variations in camera angles while maintaining the distinctive morphological features of the lesions. In contrast, excessive photometric noise can obscure the subtle colour gradients critical for distinguishing early-stage rust from nutrient deficiencies. Thus, managing the quality of synthetic data through tailored augmentation is as critical to system reliability as the volume of the underlying training set.

4.4. Results: The Best-Performing Custom Lightweight Model

The best-performing custom model was the Bean-CNN-LSTM trained on the flip training set. It showed an accuracy of 94.36%, a weighted average F1 score of 94.38%, and an MCC of 91.64%. Table 6 describes its class-specific performance. The recall score indicates the most false-negative cases in the angular leaf spot class. The precision score shows that the bean rust class has the most false-positive cases. The F1 scores suggest that the most balanced performance is in the healthy class, whereas the least is in the bean rust class. The confusion matrix in Figure 10 visualises this performance and shows a higher error pattern in the bean rust class.

4.5. Results: Comparison with Existing Models

To investigate the robustness of our CNN-LSTM-based method, we conducted further experiments with a pre-trained deep network, EfficientNet-B7 [10], trained on a larger version of the ibean dataset. We selected the EfficientNet for its excellent performance in image recognition, yet improved efficiency compared to other CNN-based models. The B7 model is even smaller/faster and achieves better image recognition accuracy.
To create an even larger dataset suitable for deep learning, the originally downloaded (OD) ibean dataset (see first four rows of Table 1) was augmented to create 40 times larger training and validation sets. To perform this, we used the data splits as found in the OD ibean dataset—1034 training images, 133 validation images, and 128 test images. Using the random seed of 0, we augmented only the training and validation sets and retained the 128 test images available in the OD dataset to facilitate fair comparisons with previous works. Thus, no augmented versions of the same image were contained in different splits. We used geometric and photometric transformations, including random rotations, horizontal and vertical flips, scaling, blurring, and contrast/brightness adjustment. Much like the originally downloaded (OD) dataset, the resulting augmented dataset is nearly balanced across the classes, containing a total of 41,360 training imageshl—13,800, 13,920, and 13,640 images for angular leaf spot, bean rust, and healthy leaves, respectively. It also contains 5320 validation imageshl—1760, 1800, and 1760 images for angular leaf spot, bean rust, and healthy leaves, respectively. The augmented dataset and code are available on GitHub (https://github.com/HJin-R/bean_disease)
To allow for some comparison between conventional Dense layers and LSTM layers, we created two separate models from EfficientNet by replacing its final classification layer. In the first model, we replaced the final classification layer with Dense or fully-connected (FC) layers, interspersed with Dropout (30–50%) and batch normalisation layers. In the second model, the classification layer is replaced with an LSTM layer and a linear classification layer, and a Dropout layer (50%) in between them. We refer to the former model as the EfficientNetB7+FC model and the latter as the EfficientNetB7+LSTM model. For best results, the entire architecture was trained, but the EfficientNet backbone was unfrozen in blocks after every five epochs. Both models were trained with the Adam optimiser at a learning rate of 5 × 10 − 4 , cross-entropy loss function, a batch size of eight and an image size of 380 × 380. The Adam optimiser has proven to be more efficient for image recognition/classification tasks than most other variants of the gradient descent algorithm, such as Stochastic Gradient Descent (SGD) and RMSProp. This is due to Adam’s adaptive learning rate and momentum-based weight updates, which lead to faster convergence [45]. The very low learning rate slowed down the training but allowed for smoother training with more chances of finding the global minima and better generalisation. The batch size of eight was necessary to allow for more memory management and better generalisation while training a very deep network like EfficientNet, but it also slows down training. We chose to increase the image resolution a bit more here to allow for improved performance, i.e., more pixels for convolution and sequencing.
The resulting performances of the models are not unexpected. Table 7 shows the results of both models on the test set; they both reached overall accuracies and F1 scores of 99.22% and had the same scores (precision, recall, and F1) across the classes. Both models reached training and validation accuracies of about 98%, but the EfficientNetB7+FC model reached this performance in only 3 epochs, while the EfficientNetB7+LSTM model reached it in 16 epochs. This means that the FC-based model did not even require the first block of layers of the EfficientNet backbone to be unfrozen before it reached such an excellent classification performance, while the LSTM layer needed some blocks of layers of the backbone to be unfrozen. This shows that the linear layers in the FC-based model are more effective in leveraging the features from the EfficientNet backbone to achieve both high accuracy and efficiency in the bean leaf disease classification task, while the fact that the LSTM-based model achieved the same result with more training indicates that it could not efficiently leverage the huge image features from the EfficientNet backbone for image classification. In fact, to simulate more similarity with the FC-based model, we also included batch normalisation and adaptive average pooling in the LSTM-based model, but the result was rather degraded, reaching only 97.68% test F1 score after 16 epochs. This could also be due to the combined augmentation techniques used in the extended dataset, as our earlier results indicate that the combination of several augmentation techniques did not necessarily produce the best result for both models. However, the fact that the EfficientNetB7+LSTM model still reaches the same accuracy as the EfficientNetB7+FC model after more training epochs indicates that it is suitable for the bean classification task and indeed smaller in size.
Table 8 shows a comparison of the performance of previous models on the ibean dataset, with the last three rows referring to our models. To facilitate a fair comparison, we evaluated our models on the standard 128-sample test split provided in the originally downloaded (OD) ibean dataset, which matches the test benchmark utilised by most prior works. For each reference model in Table 8, we report the best overall accuracy and overall F1 score exactly as published; where an overall metric was not explicitly reported, it is indicated as “–”. For instance, while Ref. [8] reported individual class-level F1 scores, the authors did not publish an aggregated overall F1 score across all three classes, so no overall F1 score is listed in Table 8. This strict adherence ensures no estimated or inferred values are introduced into the comparative baseline. It is important to note that published methodologies on the ibean dataset employ varying data augmentation strategies and training sample sizes to optimise test performance. In benchmark evaluations, data augmentation choices constitute an integral part of each method’s overall training protocol. To ensure rigorous and unbiased evaluation across these diverse pipelines, most of the comparative models in Table 8 are evaluated against the standardised, unaugmented 128-sample test split. Providing the specific training volumes and augmentation types alongside test metrics in Table 8 ensures full transparency regarding the operational conditions under which each reported result was achieved.
It is noteworthy that in many of these works, such as Refs. [8,48,50], the models were trained for a high number of epochs: 100, 50, and 30, respectively. This is most likely because they (except [50]) used the transfer learning approach of training only a few custom layers on top of the pre-trained backbone architecture. While this approach can be computationally less expensive, it does not guarantee the best performance and is prone to overfitting due to the generic image features transferred from the backbone; hence the need for longer training. Standard pre-trained backbones used in prior works range from lightweight edge networks (MobileNetV2: 3.5 M parameters; DenseNet121: 7.9 M parameters) to heavy architectures (InceptionV3: 24 M parameters; EfficientNetB6: 43 M parameters); this places our custom EfficientNetB7 models (65–67 M parameters) in clear context. Overall, both of our pre-trained models (EfficientNetB7+FC, EfficientNetB7+LSTM) achieved better generalisation in fewer training iterations than the previous models.

4.6. Ablation Experiments

To evaluate the specific design choices and empirical assumptions of our proposed framework, we conducted three targeted ablation and control experiments on the standardised Seed 0 test split ( N = 195 ). These controlled evaluations isolate the contribution of each architectural component, validating the choice of spatial traversal strategies, structural parameter scaling, and data volume constraints. First, we examine the effect of spatial traversal strategies on sequence formation to justify our row-major scanning approach. Second, we evaluate a parameter-matched control model to confirm that performance gains stem from sequential feature integration rather than increased capacity. Finally, we analyse model scaling behaviour across varying training set sizes and augmentation pipelines to demonstrate the operational boundaries of lightweight edge architectures compared to deep pre-trained backbones. All evaluations used the 905 / 195 / 195 (Train/Val/Test) partition under Seed 0. As the “Flip” training set consistently gives the best performance for the proposed Bean-CNN-LSTM, all ablation experiments, except the third one, used the “Flip” training set. The third ablation experiment used the unaugmented 905 samples, the 2715 flip set, as well as the augmented 40 × larger training set.

4.6.1. Spatial Traversal Strategy Ablation

To evaluate the effect of spatial feature sequence ordering, we tested three traversal methods to flatten the 10 × 10 × 64 CNN feature map into time steps. We represent the time steps as T and the dimensions as D:
  • Row-Major (Scan-Line) [Proposed]: Top-to-bottom row scan ( T = 100 , D = 64 ).
  • Column-Major: Left-to-right column scan ( T = 100 , D = 64 ).
  • Patch-Wise ( 2 × 2 ): Spatial block grouping ( T = 25 , D = 256 ).
As seen in Table 9, continuous line-by-line scanning strategies outperformed coarse spatial block grouping, and row-major order yielded optimal feature alignment compared to the other two strategies.

4.6.2. Parameter-Matched Model Architecture Analysis

To confirm that performance gains stem from recurrent temporal aggregation rather than arbitrary parameter scaling, we designed Bean-CNN-Compact as a control model. This control model replaces the dense projection layers of the standard Bean-CNN (527,171 parameters/6.11 MB) with a compact bottleneck (four dense layers: 5 units → 247 units → 19 units → 3 units) that matches the exact 155,219 parameters (1.86 MB) of Bean- CNN-LSTM.
As shown in Table 10, reducing parameter capacity in a standard feed-forward CNN (Bean-CNN-Compact) causes performance to drop to 82.56% accuracy. In contrast, Bean-CNN-LSTM uses the same 155,219 parameter footprint to achieve 89.74% accuracy. This proves that the recurrent structure enables parameter reduction while simultaneously increasing feature representation power.

4.6.3. Controlled Data Sub-Sampling and Augmentation Analysis

To confirm that performance trends are consistent across dataset scales and are not artefacts of data volume, all four model configurations were evaluated across controlled sub-sampled training partitions: 25%, 50%, and 100% of the 905 unaugmented training images, the 2715 flip set, and the 40 × larger augmented set. All models were evaluated on the fixed Seed 0 split ( N = 195 ) test set as used in previous ablation experiments.
The training data progression spans five controlled runs as follows:
  • Sub-sampled unaugmented runs (25% and 50% of the 905 unaugmented training images): Evaluates minimal sample availability (226 and 452 unaugmented images).
  • Full unaugmented dataset (100% of the 905 unaugmented samples): Evaluates baseline performance on the raw unaugmented dataset.
  • Small augmented set (2715 “Flip” set): Evaluates the impact of the training data volume on model performance using the same augmented set on which Bean-CNN-LSTM has its best performance.
  • Complete 40 × larger augmented dataset: Evaluates the complete pipeline incorporating the same geometric and photometric transformations used in the EfficientNet experiments.
Table 11 reveals the following three major insights:
  • Persistent Architectural Advantage: Bean-CNN-LSTM consistently outperforms Bean-CNN across every run ( + 2.88 % at 25% data, + 3.07 % at 50% data, + 4.02 % at 100% data, + 5.82 % on the 2715 augmented set and + 4.56 % at the 40 × augmented set). This proves that spatial-to-sequential feature integration provides a scale-invariant benefit.
  • Justification for Data Augmentation and Parameter Sizing: On limited training subsets (25% to 100% unaugmented and 1275 augmented data), EfficientNet-B7 performances scale from 78.46% to about 96%. It only reaches its peak performance of 99+% when trained on the full augmented dataset, highlighting the heavy data requirement of its 65 M parameters; hence the need for the 40 × larger augmented set. Even with the 1275 augmented images, the EfficientNet models could not achieve such high performance, indicating the need for a larger augmented training set.
  • Justification for a pre-trained model: Despite the 40 × larger dataset, the baseline models (Bean-CNN and Bean-CNN-LSTM) reached their peak performances at around 91% and 96%, respectively, while the EfficientNet models reached 99%. The deep backbone of the EfficientNet models gives them a significant edge over the baseline models, especially when the training dataset is large.
Table 11. Test Performance across different training and aumentation sizes. All metric values are presented as “Accuracy; F1 score”.
Table 11. Test Performance across different training and aumentation sizes. All metric values are presented as “Accuracy; F1 score”.
ArchitectureParam. Count25% Train—226
Samples
50% Train—452
Samples
100%
Unaugmented
Train—905
Samples
2715
Augmented Set
(Flip)
40 × Augmented
Dataset
Bean-CNN527.2 K68.91%; 68.85%76.33%; 75.90%83.60%; 83.32%88.62%; 88.48%91.18%; 90.90%
Bean-CNN-LSTM155.2 K71.79%; 71.25%79.41%; 79.10%87.62%; 87.36%94.44%; 94.34%95.74%; 95.62%
EfficientNet-B7-FC67.8 M78.46%; 78.12%83.59%; 83.40%90.93%; 90.81%96.44%; 96.40%99.36%; 99.28%
EfficientNet-B7-FC65.2 M78.46%; 78.10%83.49%; 83.42%90.90%; 90.80%95.95%; 95.71%99.32%; 99.26%
Together, these three ablation experiments validate the core design principles of the proposed framework. The spatial traversal strategy ablation (Table 9) proves that unrolling feature maps in a scan-line sequence maximises pseudo-temporal coherence for spatial feature extraction. The parameter-matched model architecture check (Table 10) confirms that performance gains stem directly from recurrent spatial feature integration rather than increased parameter capacity, allowing Bean-CNN-LSTM to achieve 89.74 % accuracy while matching the ultra-compact 155 K -parameter footprint of a bottlenecked CNN ( 82.56 % ). Finally, the controlled data scaling experiment (Table 11) highlights the strong relationship between model complexity, feature representation, and dataset scale: while Bean-CNN-LSTM consistently outperforms Bean-CNN across all training subsets ( + 3.08 % to + 5.12 % ), the deep pre-trained EfficientNet-B7 backbones require the full 40 × larger augmented dataset to unlock their peak 99 + % accuracy. Ultimately, while transfer learning backbones maximise raw accuracy on large datasets, Bean-CNN-LSTM provides an optimal trade-off, delivering high classification fidelity with less than 0.25 % of the parameter footprint required by heavy backbones, making it potentially desirable for edge devices.

5. Discussion

As summarised in Table 12, adopting an LSTM has a strong advantage for our ultra-lightweight custom models. This hybrid architecture can potentially be an optimal solution in resource-constrained hardware environments such as embedded/edge devices. Note that in Table 12, the EfficientNetB7+LSTM model has two different values in the third column because the number of training parameters increased after more layers of the EfficientNetB7 model were unfrozen during training. However, the remarkably high accuracy despite the smaller model size/number of parameters indicates that the LSTM-based model is an effective approach.
Beyond technical performance, the deployment of this lightweight CNN-LSTM architecture offers potentially significant advantages for integrated agricultural data management. By achieving state-of-the-art results with a model size significantly smaller than traditional CNNs (1.86 MB vs. 6.11 MB for our custom models), this system could potentially facilitate the real-time processing of leaf health data directly on edge devices. In a management context, this decentralised data processing reduces the reliance on costly cloud infrastructure and high-bandwidth connectivity, which are often absent in the rural regions where bean production is most vital. Consequently, this approach could enable more agile farm management decisions, allowing for immediate localised interventions that can mitigate the estimated 47% to 55% yield losses typically associated with bean rust and angular leaf spot. While the compact parameter footprint ( 1.86 MB ) of Bean-CNN-LSTM significantly reduces storage requirements and theoretical computational overhead compared to heavy backbones, we note that physical edge deployment involves additional hardware constraints. Factors such as memory bandwidth, operational latency (ms/sample), and energy consumption vary across physical edge accelerators. Thus, our model represents a parameter-efficient candidate design whose real-time performance and power metrics are subject to validation on physical edge hardware in future work.
The dynamic weight computations of an LSTM clearly showed strength in discovering spatiotemporal correlations within the feature map (Figure 11) in the ibean data. Our experiment with input vectors resulted in a reduction in the number of features. Despite this, Bean-CNN-LSTM required up to 1.7 times longer training time than Bean-CNN. Nevertheless, Bean-CNN is a more complex model with probably redundant connections [51], which may explain the struggle of this model to achieve the same level of performance as Bean-CNN-LSTM.
To understand why the LSTM layer yields different effects across model families (baseline CNN versus EfficientNet), we must consider the feature extraction capability of the underlying backbone architecture. In lightweight models like Bean-CNN, the feature extractor generates relatively low-level, compact spatial feature maps. Converting these maps into a sequential structure allows the LSTM to act as a temporal feature aggregator and regularizer, helping the model capture contextual dependencies across spatial regions that the baseline CNN missed. This explains the consistent performance improvement of Bean-CNN-LSTM over Bean-CNN. In contrast, a deep, highly expressive backbone like EfficientNet-B7 already extracts rich, highly disentangled high-level representations on its own. As a result, both EfficientNetB7+FC and EfficientNetB7+LSTM reach the same upper limit of 99.22% performance, making the sequential modelling of the LSTM redundant. Furthermore, adding the recurrent parameters of the LSTM head increases optimisation complexity, which is why the simpler Fully Connected (FC) head converges in fewer training epochs. Therefore, while sequential feature modelling offers clear benefits for parameter-constrained custom architectures like the baseline CNN of Bean-CNN-LSTM, a standard FC classification head is more practical and computationally efficient when paired with deep, pre-trained backbones.
Despite the strong performance of our compact models, we observed limitations of these models and the superior and more sophisticated performance of the pre-trained models in this task. More precisely, the best Bean-CNN-LSTM showed the weakest performance in detecting bean rust. Through visual examination of the data, we found that the bean rust symptoms are less severe and more varied, which could be more challenging for the model, explaining the higher error rates [52]. Figure 12 shows the misidentification by this model. Similarly to Figure 12a, Figure 12b represents the early onset of infection with additional difficulty in a highly close-up image. Figure 12c represents an error in the image with a busy background. We consider that a busy background in training data can improve generalisability in real-life conditions, as previously attempted [22,53]. Thus, there is still room for improvement in our ultra-lightweight model in this regard.
In addition, our systematic review of augmentation techniques suggests that generic combinations may introduce counterproductive noise for specialised agricultural tasks. Specifically, while standard multi-technique pipelines are often used as a default, our results indicate that they can simultaneously raise both accuracy and loss, signalling increased model uncertainty. We posit that for leaf pathology, geometric transformations, such as flipping and cropping, preserve the essential structural features of disease lesions, such as the distinctive brown angular patterns of leaf spot, whereas excessive photometric shifts (like extreme brightness or random noise) may obscure the subtle halo effects critical for identifying early-stage rust. This underscores the necessity of tailored augmentation strategies that prioritise the preservation of domain-specific visual markers over simple dataset expansion.

5.1. Implications for Agricultural Resource Management

The transition from centralised high-performance computing to decentralised Edge-AI carries significant implications for agricultural resource management. In traditional management models, the latency and cost of data transmission to the cloud often result in delayed interventions. By reducing the model footprint to 1.86 MB, a 70% reduction compared to standard architectures, this research facilitates the integration of sophisticated diagnostics into low-cost mobile management systems. This efficiency allows for precision agriculture practices where farmers can manage specific sections of a plot based on real-time data, optimising the application of fungicides and labour. Such data-driven decision-making not only reduces the operational overhead for smallholder farmers but also mitigates the environmental impact of broad-spectrum chemical use, aligning agricultural productivity with sustainable management goals.

5.2. Limitations

This work has a number of limitations, and caution is necessary not to overgeneralise the findings. The results are subject to the architectures and the ML settings described in this article for experimental convenience. Despite the remarkable performance of the models, the tuning and training process can be time-consuming. Although care was taken to avoid the models’ overfitting, the significantly small training data necessitates more diverse and robust sets to ensure that the proposed model adequately generalises across various datasets. Furthermore, the successful model construction of this work does not directly indicate real-time performance across crop types since the training, validation, and testing are based solely on the common bean plant cultivated in Uganda. Other crop types or bean plant varieties may exhibit different results.

6. Conclusions

Automated image-based disease classification aims to reduce yield loss and ensure global food security. We investigated the feasibility of a hybrid CNN-LSTM architecture for this task. In our two-phase experiments, we introduced LSTM-based ultra-lightweight custom models and pre-trained EfficientNet. Our primary results show that Bean-CNN-LSTM consistently outperforms Bean-CNN with 2.56% to 5.64% higher best test accuracy, depending on the training sets. The best Bean-CNN-LSTM model of 1.86 MB in size achieved an accuracy of 94.36%. This result inspired us to conduct further experiments with a pre-trained image recognition model, EfficientNet. In these experiments, we discovered that our modified EfficientNet models, trained on significantly larger data, improved the performance further. Both an FC-based and an LSTM-based model achieved 99.22% accuracy and F1 score. Although there are no significant gains in training efficiency, integrating an LSTM with CNN contributed to streamlining the model and enhancing performance. Nevertheless, we found that complex backgrounds and disease pattern variations are challenging, especially for our lightweight models. For further improvement, advanced feature extraction methods using specialised filters or segmentation are worth exploring. Also, comparing other RNN families, such as a gated recurrent unit (GRU), could shed light on discovering more efficient models. While the lightweight architecture demonstrates strong theoretical potential for edge deployment, physical hardware testing was beyond the scope of this study. Consequently, empirical evaluations of inference latency, peak RAM utilisation, and energy consumption on physical edge devices remain important targets for future work.

Author Contributions

Conceptualisation, H.J.R.; methodology, H.J.R. and J.D.A.; experiments, H.J.R. and J.D.A.; writing—original draft preparation, H.J.R.; writing—review and editing, H.J.R. and J.D.A.; supervision, J.D.A. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The originally downloaded (OD) ibean dataset is available from [27] while the augmented dataset curated in this study have been made openly available on IEEE Dataport https://dx.doi.org/10.21227/4k7y-vs03 and on Kaggle https://www.kaggle.com/datasets/akinyemijoseph/ibean-aug.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Pamela, P.; Mawejje, D.; Ugen, M. Severity of angular leaf spot and rust diseases on common beans in Central Uganda. Uganda J. Agric. Sci. 2014, 15, 63–72. [Google Scholar] [CrossRef]
  2. Venbrux, M.; Crauwels, S.; Rediers, H. Current and emerging trends in techniques for plant pathogen detection. Front. Plant Sci. 2023, 14, 1120968. [Google Scholar] [CrossRef] [Scilit]
  3. Mahlein, A.K. Plant disease detection by imaging sensors–parallels and specific demands for precision agriculture and plant phenotyping. Plant Dis. 2016, 100, 241–251. [Google Scholar] [CrossRef] [Scilit]
  4. Rahunathan, L.; Sivabalaselvamani, D.; Elakkiya, E.; Madhumitha, M.; Kumaresh, K. Recognition of Bean Leaf Diseases Using Neural Network and Machine Learning Techniques. In Proceedings of the 2023 3rd International Conference on Smart Data Intelligence (ICSMDI), Trichy, India, 30–31 March 2023; pp. 520–526. [Google Scholar] [CrossRef] [Scilit]
  5. Slimani, H. Artificial Intelligence-based Detection of Fava Bean Rust Disease in Agricultural Settings: An Innovative Approach. Int. J. Adv. Comput. Sci. Appl. 2023, 14, 119–128. [Google Scholar] [CrossRef] [Scilit]
  6. Islam, Z.; Islam, M.; Amanullah, A. A combined deep CNN-LSTM network for the detection of novel coronavirus (COVID-19) using X-ray images. Inform. Med. Unlocked 2020, 20, 100412. [Google Scholar] [CrossRef] [Scilit]
  7. Donahue, J.; Hendricks, L.A.; Rohrbach, M.; Venugopalan, S.; Guadarrama, S.; Saenko, K.; Darrell, T. Long-Term Recurrent Convolutional Networks for Visual Recognition and Description. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 677–691. [Google Scholar] [CrossRef] [Scilit]
  8. Önler, E. Feature fusion based artificial neural network model for disease detection of bean leaves. Electron. Res. Arch. 2023, 31, 2409–2427. [Google Scholar] [CrossRef] [Scilit]
  9. Elfatimi, E.; Eryigit, R.; Elfatimi, L. Beans Leaf Diseases Classification Using MobileNet Models. IEEE Access 2022, 10, 9471–9482. [Google Scholar] [CrossRef] [Scilit]
  10. Tan, M.; Le, Q. Efficientnet: Rethinking model scaling for convolutional neural networks. In Proceedings of the International Conference on Machine Learning, Long Beach, CA, USA, 9–15 June 2019; pp. 6105–6114. [Google Scholar]
  11. Jian, Z.; Wei, Z. Support vector machine for recognition of cucumber leaf diseases. In Proceedings of the 2010 2nd International Conference on Advanced Computer Control, Shenyang, China, 27–29 March 2010; Volume 5, pp. 264–266. [Google Scholar] [CrossRef] [Scilit]
  12. Lu, Y.; Yi, S.; Zeng, N.; Liu, Y.; Zhang, Y. Identification of rice diseases using deep convolutional neural networks. Neurocomputing 2017, 267, 378–384. [Google Scholar] [CrossRef] [Scilit]
  13. Geetharamani, G.; Arun Pandian, J. Identification of plant leaf diseases using a nine-layer deep convolutional neural network. Comput. Electr. Eng. 2019, 76, 323–338, Correction in Comput. Electr. Eng. 2019, 78, 536. https://doi.org/10.1016/j.compeleceng.2019.04.011.. [Google Scholar] [CrossRef] [Scilit]
  14. Mohanty, S.P. Using Deep Learning for Image-Based Plant Disease Detection. Front. Plant Sci. 2016, 7, 1419. [Google Scholar] [CrossRef] [Scilit]
  15. Patil, M.A.; Manohar, M. Plant Leaf Disease Classification Using Optimal Tuned Hybrid LSTM-CNN Model. SN Comput. Sci. 2023, 4, 710. [Google Scholar] [CrossRef] [Scilit]
  16. Devi, E.; Gopi, S.; Padmavathi, U.; Arumugam, S.R.; Premnath, S.; Muralitharan, D. Plant Disease Classification using CNN-LSTM Techniques. In Proceedings of the 2023 5th International Conference on Smart Systems and Inventive Technology (ICSSIT), Tirunelveli, India, 23–25 January 2023; pp. 1225–1229. [Google Scholar] [CrossRef] [Scilit]
  17. Haque, M.A.; Deb, C.K.; Gole, P.; Karmakar, S.; Dheeraj, A.; Din Shah, M.U.; Dutta, S.; Kumar, M.K.P.; Marwaha, S. An enhanced vision transformer network for efficient and accurate crop disease detection. Expert Syst. Appl. 2025, 283, 127743. [Google Scholar] [CrossRef] [Scilit]
  18. Abade, A.; Ferreira, P.A.; de Barros Vidal, F. Plant diseases recognition on images using convolutional neural networks: A systematic review. Comput. Electron. Agric. 2021, 185, 106125. [Google Scholar] [CrossRef] [Scilit]
  19. Singla, S.; Gupta, R. Deep Learning based Bean Leaf Lesion Classification utilizing EfficientNetV2-S. In Proceedings of the 2024 8th International Conference on I-SMAC (IoT in Social, Mobile, Analytics and Cloud) (I-SMAC), Kirtipur, Nepal, 3–5 October 2024; pp. 1394–1399. [Google Scholar] [CrossRef] [Scilit]
  20. Rodríguez-Lira, D.C.; Córdova-Esparza, D.M.; Álvarez Alvarado, J.M.; Romero-González, J.A.; Terven, J.; Rodríguez-Reséndiz, J. Comparative Analysis of YOLO Models for Bean Leaf Disease Detection in Natural Environments. AgriEngineering 2024, 6, 4585–4603. [Google Scholar] [CrossRef] [Scilit]
  21. Hohman, F.; Kery, M.B.; Ren, D.; Moritz, D. Model Compression in Practice: Lessons Learned from Practitioners Creating On-device Machine Learning Experiences. In Proceedings of the CHI ’24: CHI Conference on Human Factors in Computing Systems, Honolulu, HI, USA, 11–16 May 2024; pp. 1–18. [Google Scholar] [CrossRef] [Scilit]
  22. Sun, H.; Xu, H.; Liu, B.; He, D.; He, J.; Zhang, H.; Geng, N. MEAN-SSD: A novel real-time detector for apple leaf diseases using improved light-weight convolutional neural networks. Comput. Electron. Agric. 2021, 189, 106379. [Google Scholar] [CrossRef] [Scilit]
  23. Arsenovic, M.; Karanovic, M.; Sladojevic, S.; Anderla, A.; Stefanovic, D. Solving Current Limitations of Deep Learning Based Approaches for Plant Disease Detection. Symmetry 2019, 11, 939. [Google Scholar] [CrossRef] [Scilit]
  24. Yang, S.; Xiao, W.; Zhang, M.; Guo, S.; Zhao, J.; Shen, F. Image Data Augmentation for Deep Learning: A Survey. arXiv 2023, arXiv:2204.08610. [Google Scholar] [CrossRef] [Scilit]
  25. Taylor, L.; Nitschke, G. Improving Deep Learning with Generic Data Augmentation. In Proceedings of the 2018 IEEE Symposium Series on Computational Intelligence (SSCI), Bangalore, India, 18–21 November 2018; pp. 1542–1547. [Google Scholar] [CrossRef] [Scilit]
  26. Shorten, C.; Khoshgoftaar, T.M. A survey on Image Data Augmentation for Deep Learning. J. Big Data 2019, 6, 60. [Google Scholar] [CrossRef] [Scilit]
  27. AI-Lab-Makerere/ibean. 2020. Available online: https://github.com/AI-Lab-Makerere/ibean/ (accessed on 1 July 2024).
  28. Muimba-Kankolongo, A. Food Crop Production by Smallholder Farmers in Southern Africa; Academic Press: San Diego, CA, USA, 2018. [Google Scholar]
  29. Deng, L.; Platt, J.C. Ensemble deep learnig for speech recognition. Interspeech 2014, 1. [Google Scholar] [CrossRef] [Scilit]
  30. Sainath, T.N.; Vinyals, O.; Senior, A.; Sak, H. Convolutional, Long Short-Term Memory, fully connected Deep Neural Networks. In Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), South Brisbane, Australia, 19–24 April 2015; pp. 4580–4584. [Google Scholar] [CrossRef] [Scilit]
  31. Ercolano, G.; Rossi, S. Combining CNN and LSTM for activity of daily living recognition with a 3D matrix skeleton representation. Intell. Serv. Robot. 2021, 14, 175–185. [Google Scholar] [CrossRef] [Scilit]
  32. Visin, F.; Kastner, K.; Cho, K.; Matteucci, M.; Courville, A.; Bengio, Y. ReNet: A Recurrent Neural Network Based Alternative to Convolutional Networks. arXiv 2015, arXiv:1505.00393. [Google Scholar] [CrossRef] [Scilit]
  33. Van Den Oord, A.; Kalchbrenner, N.; Kavukcuoglu, K. Pixel recurrent neural networks. In Proceedings of the ICML’16: 33rd International Conference on International Conference on Machine Learning—Volume 48, New York, NY, USA, 19–24 June 2016; pp. 1747–1756. [Google Scholar] [CrossRef]
  34. Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image is Worth 16 × 16 Words: Transformers for Image Recognition at Scale. arXiv 2021, arXiv:2010.11929. [Google Scholar] [CrossRef] [Scilit]
  35. Simonyan, K.; Zisserman, A. Very Deep Convolutional Networks for Large-Scale Image Recognition. In Proceedings of the 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, 7–9 May 2015. [Google Scholar]
  36. Zhang, X.; Han, N.; Zhang, J. Comparative analysis of VGG, ResNet, and GoogLeNet architectures evaluating performance, computational efficiency, and convergence rates. Appl. Comput. Eng. 2024, 44, 172–181. [Google Scholar] [CrossRef] [Scilit]
  37. Nixon, M.S.; Aguado, A.A. Feature Extraction and Image Processing for Computer Vision; Academic Press: London, UK, 2020. [Google Scholar]
  38. He, K.; Zhang, X.; Ren, S.; Sun, J. Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification. In Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 7–13 December 2015; pp. 1026–1034. [Google Scholar] [CrossRef] [Scilit]
  39. Hochreiter, S.; Schmidhuber, J. Long Short-Term Memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [Scilit]
  40. Gers, F.A.; Schmidhuber, J.; Cummins, F. Learning to Forget: Continual Prediction with LSTM. Neural Comput. 2000, 12, 2451–2471. [Google Scholar] [CrossRef] [Scilit]
  41. Tharwat, A. Classification assessment methods. Appl. Comput. Inform. 2021, 17, 168–192. [Google Scholar] [CrossRef] [Scilit]
  42. Chicco, D.; Jurman, G. The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genom. 2020, 21, 6. [Google Scholar] [CrossRef] [Scilit]
  43. Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 618–626. [Google Scholar] [CrossRef] [Scilit]
  44. Goodfellow, I.; Bengio, Y.; Courville, A. Deep Learning; MIT Press: Cambridge, MA, USA, 2016. [Google Scholar]
  45. Kingma, D.P.; Ba, J. Adam: A method for stochastic optimization. arXiv 2014, arXiv:1412.6980. [Google Scholar]
  46. Abed, S.H.; Al-Waisy, A.S.; Mohammed, H.J.; Al-Fahdawi, S. A modern deep learning framework in robot vision for automated bean leaves diseases detection. Int. J. Intell. Robot. Appl. 2021, 5, 235–251. [Google Scholar] [CrossRef] [Scilit]
  47. Singh, V.; Chug, A.; Singh, A.P. Classification of beans leaf diseases using fine tuned cnn model. Procedia Comput. Sci. 2023, 218, 348–356. [Google Scholar] [CrossRef] [Scilit]
  48. Sunyoto, A.; Ariatmanto, D.; Noviyanto. Innovative Solutions for Bean Leaf Disease Detection Using Deep Learning. In Proceedings of the 2024 IEEE International Conference on Artificial Intelligence and Mechatronics Systems (AIMS), Virtual, 21–23 February 2024; pp. 1–5. [Google Scholar]
  49. Jain, E.; Aneja, A. Automated Detection and Classification of Bean Leaf Diseases using InceptionV3: A Deep Learning Approach. In Proceedings of the 2025 International Conference on Electronics and Renewable Systems (ICEARS), Tuticorin, India, 11–13 February 2025; pp. 1890–1895. [Google Scholar] [CrossRef] [Scilit]
  50. Karthik, R.; Aswin, R.; Geetha, K.S.; Suganthi, K. An Explainable Deep Learning Network With Transformer and Custom CNN for Bean Leaf Disease Classification. IEEE Access 2025, 13, 38562–38573. [Google Scholar] [CrossRef] [Scilit]
  51. Kahatapitiya, K.; Rodrigo, R. Exploiting the Redundancy in Convolutional Filters for Parameter Reduction. In Proceedings of the 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), Virtual, 5–9 January 2021; pp. 1409–1419. [Google Scholar] [CrossRef] [Scilit]
  52. Barbedo, J.G.A. A review on the main challenges in automatic plant disease identification based on visible range images. Biosyst. Eng. 2016, 144, 52–60. [Google Scholar] [CrossRef] [Scilit]
  53. Fenu, G.; Malloci, F.M. DiaMOS Plant: A Dataset for Diagnosis and Monitoring Plant Disease. Agronomy 2021, 11, 2107. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Three classes of bean leaf images in the ibean dataset.
Figure 1. Three classes of bean leaf images in the ibean dataset.
Jimaging 12 00468 g001
Figure 2. Augmentation effect.
Figure 2. Augmentation effect.
Jimaging 12 00468 g002
Figure 3. Custom architecture of lightweight models.
Figure 3. Custom architecture of lightweight models.
Jimaging 12 00468 g003
Figure 4. Activation map by a custom lightweight model.
Figure 4. Activation map by a custom lightweight model.
Jimaging 12 00468 g004
Figure 5. Training Bean-CNN on the unaugmented set.
Figure 5. Training Bean-CNN on the unaugmented set.
Jimaging 12 00468 g005
Figure 6. Training Bean-CNN-LSTM on the unaugmented set.
Figure 6. Training Bean-CNN-LSTM on the unaugmented set.
Jimaging 12 00468 g006
Figure 7. Graph representation of the best test accuracy: Comparison of two custom lightweight architectures.
Figure 7. Graph representation of the best test accuracy: Comparison of two custom lightweight architectures.
Jimaging 12 00468 g007
Figure 8. Box plots showing the average test accuracy on each training set: The red line indicates the median value, and a box represents the distribution between 25 and 75 percentiles. Small hollow circles indicate outliers.
Figure 8. Box plots showing the average test accuracy on each training set: The red line indicates the median value, and a box represents the distribution between 25 and 75 percentiles. Small hollow circles indicate outliers.
Jimaging 12 00468 g008
Figure 9. Box-plot representation: The average performance of our lightweight custom models (Bean-CNN and Bean-CNN-LSTM) from 30 training runs.
Figure 9. Box-plot representation: The average performance of our lightweight custom models (Bean-CNN and Bean-CNN-LSTM) from 30 training runs.
Jimaging 12 00468 g009
Figure 10. Confusion matrix: The best Bean-CNN-LSTM model.
Figure 10. Confusion matrix: The best Bean-CNN-LSTM model.
Jimaging 12 00468 g010
Figure 11. GradCAM feature maps. The heatmaps indicate that the model correctly prioritises the necrotic centres of the lesions (yellow patches indicate high activation) rather than the leaf edges or image background (dark-blue regions), validating the management reliability of the system.
Figure 11. GradCAM feature maps. The heatmaps indicate that the model correctly prioritises the necrotic centres of the lesions (yellow patches indicate high activation) rather than the leaf edges or image background (dark-blue regions), validating the management reliability of the system.
Jimaging 12 00468 g011
Figure 12. Misclassified images. The first column indicates the raw images, while the second column shows the predicted image activation. (a) True label: Bean rust; Predicted label: Healthy. (b) True label: Bean rust; Predicted label: Healthy. (c) True label: Angular leaf spot; Predicted label: Bean rust.
Figure 12. Misclassified images. The first column indicates the raw images, while the second column shows the predicted image activation. (a) True label: Bean rust; Predicted label: Healthy. (b) True label: Bean rust; Predicted label: Healthy. (c) True label: Angular leaf spot; Predicted label: Bean rust.
Jimaging 12 00468 g012
Table 2. Configuration of custom lightweight models.
Table 2. Configuration of custom lightweight models.
ParameterBean-CNNBean-CNN-LSTM
Learning rate1 × 10 − 3 1 × 10 − 3
Epochs4040
OptimiserAdamAdam
Batch size3232
Final Conv + Pooling dim. 10 × 10 × 64 10 × 10 × 64
Reshape dim.6400 (flattened) 100 × 64
Traversal order–Row-major (scan-line)
Sequential layer (1st and 2nd)64 units and 8 units64 units and 16 units
Dropout rate0.30.3
Table 3. Top-1 Test performance of custom lightweight models: Results across training sets—the best shown in bold.
Table 3. Top-1 Test performance of custom lightweight models: Results across training sets—the best shown in bold.
Bean-CNN
Training Set Accuracy Loss F1 Score MCC
Unaugmented0.84100.51800.84100.7614
Brightness0.87690.51800.87510.8163
Combination0.88720.65610.88820.8329
Crop0.88210.45970.88440.8267
Flip0.88720.50890.88670.8307
Rotation0.89230.51300.89130.8391
Bean-CNN-LSTM
Training setAccuracyLossF1 ScoreMCC
Unaugmented0.87690.37730.87370.8173
Brightness0.90770.39020.90740.8615
Combination0.92310.41160.92340.8856
Crop0.93330.27570.93320.9000
Flip0.94360.17690.94380.9164
Rotation0.91790.33980.91680.8776
Table 4. Mean and standard deviation (Std.) values across the 30 experiment runs for Bean-CNN and Bean-CNN-LSTM.
Table 4. Mean and standard deviation (Std.) values across the 30 experiment runs for Bean-CNN and Bean-CNN-LSTM.
Bean-CNN
Accuracy F1 Score MCC
Training Set Mean Std. Mean Std. Mean Std.
Unaugmented0.78080.04030.77780.04430.67760.0562
Brightness0.80260.04440.80210.04460.70920.0622
Combination0.85250.01740.85180.01770.78550.0214
Crop0.84600.01700.84920.01740.77320.0246
Flip0.85400.02140.85330.02200.78490.0314
Rotation0.84750.02660.84700.02760.77400.0370
Bean-CNN-LSTM
AccuracyF1 ScoreMCC
Training setMeanStd.MeanStd.MeanStd.
Unaugmented0.82750.03830.82740.03810.74700.0545
Brightness0.86240.03080.86360.02990.79680.0436
Combination0.87700.02050.87600.02110.81650.0341
Crop0.88750.02710.87440.08020.83310.0401
Flip0.88650.02500.88540.02630.83220.0362
Rotation0.87440.01830.87400.01830.81360.0267
Table 5. Paired 2-tailed t-test of the Accuracy, F1 score and MCC values for Bean-CNN versus Bean-CNN-LSTM across the 30 experiments. N = 30 degree of freedom = 29 .
Table 5. Paired 2-tailed t-test of the Accuracy, F1 score and MCC values for Bean-CNN versus Bean-CNN-LSTM across the 30 experiments. N = 30 degree of freedom = 29 .
Training SetAccuracyF1 ScoreMCC
Unaugmented 3.52 × 10 − 6 2.49 × 10 − 6 2.55 × 10 − 6
Brightness 1.29 × 10 − 5 7.81 × 10 − 6 7.53 × 10 − 6
Combination 1.25 × 10 − 5 2.60 × 10 − 5 3.02 × 10 − 4
Crop 1.33 × 10 − 7 9.96 × 10 − 2 2.11 × 10 − 7
Flip 8.13 × 10 − 6 1.59 × 10 − 5 1.18 × 10 − 5
Rotation 1.19 × 10 − 4 1.25 × 10 − 4 7.14 × 10 − 5
Table 6. Performance of the best Bean-CNN-LSTM model.
Table 6. Performance of the best Bean-CNN-LSTM model.
ClassPrecision (%)Recall (%)F1 Score (%)
Angular leaf spot100.0091.1895.38
Bean rust89.2393.5591.34
Healthy94.1298.4696.24
Overall Acc. (%)94.36
Overall F1 (%)94.38
Table 7. Classification scores of both the EfficientNetB7+FC and EfficientNetB7+LSTM Models.
Table 7. Classification scores of both the EfficientNetB7+FC and EfficientNetB7+LSTM Models.
ClassPrecision (%)Recall (%)F1 Score (%)
Angular leaf spot97.73100.0098.85
Bean rust100.0097.6798.82
Healthy100.00100.00100.00
Overall Acc. (%)99.22
Overall F1 (%)99.22
Table 8. Comparison with previous works.
Table 8. Comparison with previous works.
SourceMethodAcc.; F1 Score (%)Train/Test SizeAugmentation
Abed et al., 2021 [46]DenseNet12191.02; –9324/259flip and rotation
Elfatimi et al., 2022 [9]MobileNetV292.97; 92.941034/128none
Singh et al., 2023 [47]EfficientnetB691.74; –1034/128none
Önler 2023 [8]MobileNetV299.24; –1034/128blur, brightness, crop, flip, and rotation
Sunyoto et al., 2024 [48]DenseNet12196.90; 97.001034/128rotation and zoom
Jain & Aneja 2025 [49]InceptionV391.00; 91.00693/149brightness, flip, rotation, and zoom
Karthik et al., 2025 [50]Transformer+CNN97.66; 97.6713,442/128flip, rotation, and translation
OursBean-CNN-LSTM94.36; 94.3841,360/128blur, brightness, flip, rotation, and scaling
EfficientNetB7+FC99.22; 99.22as aboveas above
EfficientNetB7+LSTM99.22; 99.22as aboveas above
Table 9. Traversal strategies comparison for Bean-CNN-LSTM.
Table 9. Traversal strategies comparison for Bean-CNN-LSTM.
Traversal
Strategy
Tensor ShapeTest Accuracy
(%)
Test F1 (%)Test MCC (%)
Row-major (scan-line) [proposed] 100 × 64 89.7489.6485.11
Column-major 100 × 64 88.7288.7083.32
Patch-wise ( 2 × 2 ) 25 × 256 86.6786.5880.15
Table 10. Parameter-matched architectural check.
Table 10. Parameter-matched architectural check.
ArchitectureClassification
Head
Param. CountModel Size
(MB)
Test Accuracy
(%)
Test F1 (%)Test MCC
(%)
Bean-CNN (Baseline)Standard FC527,1716.1184.6284.4877.24
Bean-CNN-Compact (Control)Bottleneck FC155,2191.8682.5682.4874.12
Bean-CNN-LSTM (Proposed)Spatial-to-Seq LSTM155,2191.8689.7489.6485.11
Table 12. Summary of models.
Table 12. Summary of models.
ModelsModel Size (MB)Total ParamsTrainable ParamsBest Accuracy/F1 (%)
Bean-CNN6.11527,171527,17189.23/89.13
Bean-CNN-LSTM1.86155,219155,21994.36/94.38
EfficientNetB7+FC260.5667,825,5954,039,03599.22/99.22
EfficientNetB7+LSTM250.2065,165,1391,377,667/52,844,79599.22/99.22
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Rhee, H.J.; Akinyemi, J.D. A Resource-Efficient Hybrid CNN-LSTM Network for Image-Based Bean Leaf Disease Classification. J. Imaging 2026, 12, 468. https://doi.org/10.3390/jimaging12100468

AMA Style

Rhee HJ, Akinyemi JD. A Resource-Efficient Hybrid CNN-LSTM Network for Image-Based Bean Leaf Disease Classification. Journal of Imaging. 2026; 12(10):468. https://doi.org/10.3390/jimaging12100468

Chicago/Turabian Style

Rhee, Hye Jin, and Joseph Damilola Akinyemi. 2026. "A Resource-Efficient Hybrid CNN-LSTM Network for Image-Based Bean Leaf Disease Classification" Journal of Imaging 12, no. 10: 468. https://doi.org/10.3390/jimaging12100468

APA Style

Rhee, H. J., & Akinyemi, J. D. (2026). A Resource-Efficient Hybrid CNN-LSTM Network for Image-Based Bean Leaf Disease Classification. Journal of Imaging, 12(10), 468. https://doi.org/10.3390/jimaging12100468

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop