1. Introduction
Leaf starch content serves as a critical storage reserve for carbohydrates, sustaining plant metabolism and driving growth during non-photosynthetic periods [
1]. In early seedling development, leaves function as primary storage organs where photosynthetic products accumulate as starch before being hydrolyzed into monosaccharides to fuel vascular reconnection and tissue expansion [
2]. Normal starch metabolism is essential for maintaining photosynthetic efficiency, delaying senescence, and ensuring high crop yields [
3]. Particularly in grafting research, starch accumulation and mobilization are closely associated with graft union formation and vascular bridge reconnection, serving as early physiological markers for graft compatibility [
4]. Despite its significance, traditional starch determination primarily relies on chemical methods such as spectrophotometry, chromatography, or thermogravimetric analysis. While accurate, these approaches are inherently destructive, labor-intensive, and costly, rendering them unsuitable for large-scale, real-time monitoring in modern breeding programs. Meanwhile, the rapid advancement of computer vision and machine learning has introduced innovative strategies for the intelligent detection of substance content [
5].
Hyperspectral imaging (HSI) has been widely employed for physiological assessment by integrating spatial and spectral information [
6,
7]. Several studies have demonstrated its efficacy in starch quantification across diverse crops. For instance, Zhang et al. [
8] established a rice starch regression model with an
R2 of 0.8029, enabling spatial starch visualization. Frey et al. [
9] utilized HSI to predict leaf starch in red clover with an
R2 of 0.36 to support the breeding of high-starch varieties. Similarly, Bu et al. [
10] achieved
R2 values exceeding 0.99 for both amylose and amylopectin in mixed sorghum samples using data fusion techniques. Furthermore, Polder et al. [
11] confirmed strong correlations between starch levels and near-infrared spectral responses in tomato leaves and fruits. Despite these advancements, the widespread field adoption of HSI remains limited by high equipment costs, complex data processing, and poor portability [
12].
Multispectral imaging (MSI) offers a more practical alternative by capturing data at specific, optimized wavelengths. Conventional MSI systems (e.g., RedEdge-P) typically utilize narrowband filters; however, their fixed spectral bands limit customization for specific target substances, often resulting in lower regression accuracy compared to HSI [
13,
14,
15]. In contrast, multispectral imaging systems based on narrowband light-emitting diodes (LEDs) offer superior flexibility, rapid switching, and high spectral control (within ±5 nm). For example, Wang et al. [
16] developed a 13-channel LED-based multispectral microscopic imaging system. Similarly, Wang et al. [
17] designed a portable high-resolution device equipped with an LED array for capturing dicotyledonous leaf images. Despite these hardware advancements, research specifically targeting non-destructive leaf starch quantification remains sparse, and few studies have explored multimodal deep learning architectures to fuse spatial and spectral features in this context.
In this study, we developed a portable, low-cost multispectral imaging and analysis system based on 12-channel narrowband LEDs. Using watermelon–pumpkin grafted seedlings as experimental subjects, we proposed a hybrid CNN-FCNN-Transformer architecture to collaboratively extract multimodal features. Specifically, a CNN was employed for spatial texture analysis from RGB images, an FCNN for spectral reflectance processing, and a Transformer network for high-dimensional feature fusion and regression. This approach aims to provide a real-time, high-accuracy solution for non-destructive starch detection, offering a theoretical foundation for field-scale plant physiological monitoring. The modular design of our system further allows for rapid adaptation to other target substances, demonstrating significant potential for precision agricultural management.
3. Results and Analysis
All experimental datasets were randomly divided into training and test sets in an 8:2 ratio. The enzymatically assayed starch content was adopted as the reference standard and compared against the detection values. Model performance was evaluated utilizing R2 (coefficient of determination) and RMSE (root mean square error). The development environment was established on a dedicated deep-learning server equipped with an Intel(R) Core(TM) i9-12900K processor, 128 GB DDR4 memory, a 4 TB SSD, and an NVIDIA RTX 3090 graphics card (24 GB VRAM). Validation testing was conducted on a laptop configured with an Intel(R) Core(TM) i5-12400F processor, 16 GB DDR4 memory, a 1 TB SSD, and an NVIDIA GTX 1060 graphics card (6 GB VRAM). Both platforms operated under Python 3.6 and PyTorch 1.13.1.
3.1. Evaluation of Leaf Segmentation Performance Using the DeepLab v3+ Model
Twelve grafted seedling images were randomly selected. Following manual annotation of watermelon and pumpkin leaves, the annotated images were compared with segmentation results generated by the Deeplab v3+ model. The comparison outcomes are summarized in
Table 2. As presented in the table, all randomly selected images from the test set achieved a PA (Pixel Accuracy) exceeding 0.99 and an IoU (Intersection over Union) above 0.88, with average values of 0.9991 and 0.9475, respectively. These results robustly demonstrate that employing the Deeplab v3+ network for leaf segmentation yields a high-performance outcome.
Figure 3 illustrates the segmentation results obtained using the DeepLab v3+ network. In the figure, the dicotyledonous structures correspond to the cotyledons of watermelon seedlings, while the monocotyledonous structures represent the cotyledons of pumpkin seedlings. As shown in
Figure 3, the DeepLab v3+ semantic segmentation network demonstrates robust segmentation performance on both RGB and binary images of watermelon and pumpkin grafted seedlings, effectively distinguishing between the leaf tissues of watermelon and pumpkin.
3.2. Results of Optimal Feature Wavelength Selection
3.2.1. Test Results of Pretreatment Methods
Modeling results employing various preprocessing combinations based on Partial Least Squares regression (PLSR) were evaluated, as presented in
Table 3. The optimal preprocessing method for both watermelon and pumpkin leaves was identified as Gaussian smoothing filtering combined with first-order derivation.
3.2.2. Test Results of Optimal Band Selection Based on CARS Feature Scores
- (1)
Results of Optimal Machine Learning Modeling Method Selection
To identify the optimal machine learning modeling approach, five machine learning models were employed to establish regression models for leaf starch content. The results, as shown in
Table 4, indicate that the RF model consistently achieved the highest R
2 values and the smallest RMSE values. Consequently, random forest was determined to be the most effective modeling method for the preprocessed hyperspectral data of both watermelon and pumpkin leaves.
- (2)
Test Results of CARS Feature Scores
Following the designation of the random forest (RF) algorithm as the core evaluation model, the Competitive Adaptive Reweighted Sampling (CARS) algorithm was employed to perform feature band selection on the preprocessed hyperspectral data. In this procedure, the sampling iteration was set to 50 cycles, and a 5-fold cross-validation method was implemented to pinpoint the optimal feature wavelength combination based on the global minimum value of the Root Mean Square Error (RMSE) from the cross-validation.
Figure 4 illustrates the trend of model RMSE variation with the number of feature wavelengths extracted. As can be observed from the figure, when the count of extracted feature bands was low, the model suffered from inadequate fitting capability due to insufficient spectral information characterizing starch content, resulting in elevated RMSE values accompanied by significant oscillations. As the number of extracted feature wavelengths progressively increased and reached 12, the spectral information most pertinent to starch content was effectively integrated, leading to a substantial reduction in regression error, with the RMSE descending to its global minimum. Further addition of feature bands would introduce redundant bands and noise signals, potentially inducing overfitting of the model, which consequently resulted in a rebound of the RMSE.
Based on the identification of the global minimum RMSE, this study successfully identified precisely 12 optimal feature bands for both watermelon leaf and pumpkin leaf samples. These 12 bands represent the core spectral information that consistently exhibited the highest predictive contribution across successive feature score comparisons. Specifically, the 12 feature bands associated with watermelon leaf starch content were ultimately determined as: 450 nm, 470 nm, 490 nm, 520 nm, 540 nm, 570 nm, 610 nm, 625 nm, 640 nm, 690 nm, 740 nm, and 880 nm. The 12 feature bands relevant to pumpkin leaf starch content were ultimately determined as: 410 nm, 450 nm, 480 nm, 490 nm, 510 nm, 520 nm, 540 nm, 570 nm, 780 nm, 830 nm, 850 nm, and 900 nm.
3.3. Performance Evaluation of the CNN–FCNN–Transformer Regression Model
3.3.1. Performance Evaluation of Image–Spectral Modeling Approaches
We examined the impact of seven Convolutional Neural Network (CNN) architectures on modeling outcomes, comparing two input modalities: image-spectral multimodal inputs versus spectral-only data inputs. Training was conducted using 5-fold cross-validation. The results are presented in
Table 5. Experimental findings demonstrate that ShuffleNet v2 achieved the highest
R2 in starch content regression for watermelon leaves, whereas EfficientNet b1 yielded the highest
R2 in starch content regression for pumpkin leaves. In comparison to models trained solely on spectral data, those utilizing image-spectral multimodal data exhibited significant performance improvements, validating the importance of multi-modal data fusion.
3.3.2. Comparative Evaluation of Deep Learning Regression Networks
To demonstrate the advancement of the Transformer network, a comparative experiment was conducted using a Long Short-Term Memory (LSTM) network. The LSTM parameters were set as follows: hidden size = 1024, number of layers = 1. The results are presented in
Table 6. On the watermelon leaf dataset, the R
2 value for ShuffleNet v2 combined with Transformer was 0.9566, representing an increase of approximately 3.34% compared to the R
2 of LSTM (0.9257) under the same backbone network. On the pumpkin leaf dataset, the R
2 value for EfficientNet b1 combined with Transformer was 0.9670, reflecting an improvement of about 0.91% relative to the R
2 of LSTM (0.9583) under the same backbone network. Compared to the LSTM network, the Transformer network exhibited stronger non-linear fitting capabilities, enabling more effective capture of global information within the data, thereby enhancing the model’s expressive and generalization abilities.
3.4. Performance Test of the Portable Starch Detection Device
The system demonstrated stable operation during continuous testing, facilitating one-click acquisition of 12-band multispectral and RGB images without communication failures. Compared to the benchmark chemical method (average 135 min/sample), the proposed system achieved a total test time of 28.2 s. Specifically, 21.1 s was dedicated to image capture, while the remaining 7.1 s was utilized for automated processing, including DeepLabV3+ leaf segmentation, feature extraction, and Transformer-based regression. Results summarized in
Table 7 indicate that while the actual accuracy (
R2 = 0.928–0.952) slightly deviates from the theoretical model (
R2 = 0.956–0.967) due to hardware-related factors like LED wavelength process deviation, it remains sufficient for rapid field screening.
Furthermore, the system offers significant economic and practical advantages for large-scale agricultural applications. From a cost–benefit perspective, it reduces hardware costs to approximately 10% of commercial multispectral cameras (e.g., RedEdge-P) and eliminates recurring expenses for specialized enzymatic reagent kits. By providing a non-destructive detection method with a high throughput of 120 samples per hour, this portable device enables efficient, large-scale starch monitoring in breeding programs without sacrificing plant integrity.
4. Discussion
A portable method and a corresponding device were proposed for determining the starch content in plant leaves. Currently, content prediction primarily relies on conventional neural network architectures. For instance, Wang et al. (2023) [
12] constructed a multi-source image fusion model based on CNN, fully connected neural networks (FCNN), and recurrent neural networks (RNN). However, the application of Transformer networks in this field remains relatively limited. While machine learning has been widely explored for the multiscale characterization of starch properties and material design in laboratory settings [
29], its application for in situ, non-destructive monitoring in living plants remains less common. By bridging the gap between laboratory-grade starch analysis and field-scale physiological sensing, this study demonstrates a significant shift toward practical, real-time agricultural monitoring. In this study, a Transformer network was introduced to predict starch content by extracting RGB image features via CNN and spectral features via FCNN. This integration enables the model to better handle sequential data and long-range dependencies, thereby further improving the accuracy of content regression. Compared to architectures that focus solely on extracting local spatial features (CNN) or independent spectral features (FCNN), the Transformer, through its self-attention mechanism, is capable of capturing global correlations and long-range dependencies among the fused multimodal features. This deep modeling capability for complex non-linear relationships between features enhances the model’s generalization performance under multi-source heterogeneous data, thereby significantly improving the regression accuracy of starch content prediction. Simultaneously, CNN, FCNN, and Transformer networks were organically combined to construct an end-to-end regression model based on multi-source image fusion.
Current multispectral cameras on the market primarily utilize narrowband filters to acquire spectral images. For example, the RedEdge-P camera provides only six spectral bands: blue (center wavelength 475 nm, bandwidth 32 nm), green (560 nm, 27 nm), red (668 nm, 14 nm), red edge (717 nm, 12 nm), and near-infrared (842 nm, 57 nm). These preset bands cannot be freely selected according to user requirements, and their bandwidths are significantly larger than the ±5 nm range of narrowband LEDs. In contrast, the proposed portable multispectral imaging system is designed based on characteristic spectral bands associated with leaf starch content, which not only improves the accuracy of spectral measurement but also allows for flexible band selection, offering higher flexibility and adaptability. It is worth noting that the selection of CARS parameters, specifically the 50 Monte Carlo sampling iterations and 5-fold cross-validation, was empirically optimized to maximize the accuracy of the multispectral forecasting model. Altering these parameters would disrupt the balance of the exponential decay function, leading to either the retention of noisy bands or the loss of critical spectral information. As demonstrated by the global minimum of RMSE achieved at exactly 12 characteristic bands (
Figure 4), the current parameter configuration represents the optimal setting, ensuring the highest and most stable regression accuracy. Furthermore, while the RedEdge-P camera is priced at approximately 50,000 RMB, the proposed portable multispectral imaging system costs only 5000 RMB. The system adopts a modular architecture, facilitating the free selection of characteristic bands by simply replacing the light source module, thereby significantly reducing equipment costs. Given the relatively limited research on the non-destructive detection of leaf starch, the successful applications of multispectral imaging and deep learning in assessing other physiological indicators, such as chlorophyll [
12] and nitrogen content [
30], provide crucial cross-disciplinary validation. Consequently, the portable system proposed in this study, with its customizable LED modules and robust Transformer architecture, demonstrates high versatility and holds significant potential to serve as a highly efficient, low-cost alternative method for monitoring various other plant biochemical components in precision agriculture.
Although the proposed method and device have been effectively validated for predicting starch content in watermelon–pumpkin grafted seedling leaves, further research is required to extend these predictions to a wider variety of plant species. When addressing starch content prediction for more diverse plant leaves, the selection strategy for characteristic bands may require fine-tuning. Future research will focus on exploring the variation patterns of starch content across more plant species and deeply analyzing the correlation between starch levels and environmental factors—such as temperature, humidity, and light intensity—to enhance the understanding of plant physiological status and nutrient uptake mechanisms under varying environmental conditions.
5. Conclusions
A portable detection method and instrument have been proposed for determining the starch content in plant leaves. The Transformer network was introduced, and an end-to-end content regression model based on multi-source image fusion was constructed by extracting RGB image features through CNN and spectral features through FCNN. This integration enables the model to better handle sequence data and long-range dependencies, thereby further improving the accuracy of content regression. The multispectral acquisition device scheme based on narrowband LEDs can flexibly select bands according to specific needs and is cost-effective. It can replace multispectral cameras in situations where real-time requirements are not high. Compared with traditional destructive enzymatic methods, this system significantly reduces the comprehensive detection cost. Specifically, it eliminates the need for expensive chemical reagents and complex sample pretreatment, and its hardware cost is only about 10% of that of commercial multispectral cameras. In a viable plant breeding situation, the high efficiency and low-cost characteristics of this method make it a more economical choice for the large-scale screening of breeding populations. Despite these advantages, the current system is highly sensitive to ambient light conditions, which restricts its application to controlled laboratory environments. Due to potential interference from solar radiation and complex environmental lighting, the device is not yet suitable for direct use in open agricultural fields. Overall, the proposed portable spectral detection method achieves dynamic non-destructive detection of starch content at low cost and high precision within controlled settings, laying a theoretical foundation for future field-based spectral imaging applications.