2.1. General Concept and Research Methodology
The research methodology is designed as a sequential process for transforming UAV crop imagery into structured information about vegetation cover and candidate problem zones. The experimental workflow combines dataset preparation, semantic segmentation, RGB-based spatial interpretation, vectorization, and decision-support analysis. A complementary UAV acquisition and geospatial processing architecture is proposed for subsequent field deployment. This approach is consistent with current trends in UAV monitoring, in which individual data collection and analysis operations are integrated into a single information process [
2,
17,
18,
21].
The primary model-development and internal evaluation experiments were performed using high-resolution RGB imagery and corresponding reference masks from the open-access dataset acquired using a DJI Air 2S UAV (SZ DJI Technology Co., Ltd., Shenzhen, China). In addition, RGB imagery independently acquired with a DJI Mavic 3 Multispectral (SZ DJI Technology Co., Ltd., Shenzhen, China) was used for pilot field acquisition and external validation of the selected U-Net-ResNet50 model. Multispectral and thermal channels were not used for model training or quantitative segmentation evaluation in the present study.
In the proposed field-deployment workflow, UAV imagery may be acquired using RGB or multispectral sensors together with positioning metadata required for subsequent geospatial processing. In the present study, this acquisition workflow was pilot-tested using the DJI Mavic 3 Multispectral, while the primary model-development experiments remained based on the open DJI Air 2S RGB dataset [
9,
17,
18].
For future field deployment, the acquired UAV imagery can undergo radiometric and geometric correction, photogrammetric alignment, and orthomosaic generation. In the experimental part of this study, the RGB images and corresponding reference masks available in the open dataset were used directly and divided into equal 512 × 512-pixel patches for subsequent model development. The specific parameters of photogrammetric processing, the size of the fragments, and the normalization method are presented in
Section 2.5.
Each source RGB image in the open dataset was accompanied by a corresponding pixel-wise reference mask. During tiling, identical spatial cropping was applied to both the RGB image and its reference mask, after which the resulting image–mask patches were assigned to the training, validation, and test subsets. This division takes into account the spatial origin of the images to ensure that adjacent fragments from the same area do not end up in both the training and test sets simultaneously. This reduces the risk of data leakage and ensures a more objective assessment of the model’s ability to handle new areas [
23,
25].
During the deep learning phase, the models are trained to perform pixel-level recognition of classes defined in the annotation protocol. The models produce pixel-level segmentation masks that reflect the spatial distribution of annotated vegetation patches and serve as the basis for further spatial analysis and assessment of potential problem areas. The models are compared using identical sample structures and consistent training conditions. IoU, the Dice coefficient, Precision, Recall, and the F1-score are used for evaluation, which allows for separate consideration of localization accuracy and errors related to missed or incorrectly detected pixels [
6,
22,
23].
In the final experimental stage, the segmentation results are combined with RGB-based vegetation indicators to identify candidate RGB heterogeneity zones. These zones are converted into image-space vector objects for which pixel-based coordinates, geometric boundaries, and relative areas are calculated. The resulting indicators are subsequently used by the decision-support module to prioritize image regions for further field inspection. In future georeferenced field deployment, these vector objects can be transformed into geographic coordinates and integrated with GIS platforms.
Thus,
Figure 2 illustrates the overall framework of the proposed intelligent crop-monitoring system. The following sections describe the experimental dataset, proposed UAV hardware configuration and flight protocol, preprocessing procedures, dataset preparation, deep learning models, training strategy, spatial interpretation, and decision-support procedures.
2.2. Experimental UAV Data Set
To experimentally validate the proposed intelligent monitoring system, we used the open dataset “High-Resolution RGB Images and Corresponding Masks of Agricultural Fields,” available in the open repository Zenodo (Record 12607112) [
42]. The use of an open dataset ensures the reproducibility of experiments, the possibility of independent verification of the obtained results, and a fair comparison of the performance of different deep learning models under identical conditions.
The dataset contains high-resolution RGB aerial images of agricultural fields captured using a DJI Air 2S unmanned aerial vehicle, together with corresponding reference masks. According to the repository description, the imagery covers several agricultural categories, including orchards, olive groves, green wheat, and vineyards. Because verified image-level crop labels were not available, no crop-specific subset was selected, and the dataset was treated as a mixed agricultural vegetation dataset. Following an automated data-integrity and image–mask pairing audit, 326 unambiguous valid RGB image–reference mask pairs were retained and used in the subsequent stages of the study.
The Zenodo repository description reports 325 source images; however, the reproducibility audit of the downloaded archive identified 326 unambiguous decodable RGB–mask pairs. Therefore, the retained count reported here reflects the actual set of image–mask pairs used in the final computational experiment.
Examples of raw aerial images, reference masks, and their combinations are shown in
Figure 3. These examples demonstrate varying spatial distributions of vegetation cover, heterogeneity in crop density, and the complexity of boundaries between vegetation and background objects. It is precisely these characteristics that make the dataset suitable for evaluating the effectiveness of modern semantic segmentation methods.
The main characteristics of the dataset used are presented in
Table 1. All source materials are provided in TIFF format, which allows the original quality of the aerial images to be preserved without loss, and the availability of pixel-level reference masks enables the use of supervised learning models for semantic segmentation.
Analysis of the audited dataset showed that the source UAV RGB images have high spatial resolution and represent heterogeneous agricultural vegetation scenes. To standardize the model input, the 326 retained RGB image–mask pairs were automatically divided into 21,324 non-overlapping patches of 512 × 512 pixels. Based on the foreground fraction in the corresponding reference masks, 19,999 patches were classified as foreground-dominant (>50% vegetation pixels), whereas 1325 patches were classified as background-dominant (≤50% vegetation pixels).
Thus, the resulting patch dataset contains a broad range of vegetation-cover conditions and was used for subsequent model development, repeated grouped evaluation, and comparative semantic segmentation analysis (
Table 1).
The source RGB data are individual high-resolution UAV images provided as raw drone outputs rather than orthomosaics. According to the repository documentation, the corresponding masks were generated using orthomosaic processing software and were not reported as manually delineated expert annotations. Therefore, throughout this study, they are referred to as reference masks rather than definitive ground-truth masks. The automated masks may contain residual uncertainties related to vegetation-boundary delineation, omission or commission errors, and local image–mask correspondence. The reproducibility audit verified file integrity, decodability, image–mask pairing, and dimensional consistency but did not constitute an independent expert re-annotation of the masks.
When creating the experimental dataset, special attention was paid to avoiding spatial information leakage between samples. To achieve this, the division into training, validation, and test sets was not performed randomly for individual fragments but was based on the original aerial photographs. This approach ensures that adjacent sections of the same field do not end up in different samples simultaneously, thereby ensuring a more objective assessment of the generalization ability of deep learning models.
2.3. DJI Mavic 3 Multispectral Platform and Pilot Field Validation
The DJI Mavic 3 Multispectral (Mavic 3M) was used for pilot field validation of the hardware component of the proposed crop-monitoring framework and to assess the feasibility of transferring the developed computational workflow to real UAV data acquisition conditions. The platform integrates an RGB camera, multispectral sensors, and RTK positioning capabilities, enabling the acquisition of spatially referenced field imagery.
The study comprised two complementary experimental components. The open-access RGB dataset acquired using a DJI Air 2S UAV was used for model training, internal validation, architecture selection, and repeated held-out testing. Separately, the DJI Mavic 3 Multispectral platform was used for pilot field data acquisition and independent external validation of the selected U-Net-ResNet50 model. The Mavic 3M imagery was not used for model training, hyperparameter selection, or internal model comparison. Instead, independently annotated RGB image patches acquired over a wheat field were used exclusively to evaluate the transferability of the trained model to a different UAV platform and real field conditions without additional retraining or fine-tuning.
The architecture of the proposed hardware-software system is shown in
Figure 4. It includes a UAV platform, RGB and multispectral cameras, an RTK module, a Raspberry Pi 4 single-board computer, a local data storage system, a ground control station, and a module for subsequent photogrammetric and intelligent data processing.
In the proposed hardware architecture, high-precision spatial georeferencing can be provided by the integrated RTK positioning system. The use of RTK correction data enables centimeter-level positioning and spatial synchronization of the acquired images with geographic coordinates. Such georeferencing is required for the subsequent generation of spatially referenced products and the integration of semantic segmentation results with geographic information systems.
The proposed architecture includes a Raspberry Pi 4 Model B single-board computer with 8 GB of RAM (Raspberry Pi Ltd., Cambridge, UK) as an auxiliary onboard computing unit. It can be used for telemetry recording, preliminary data processing, temporary data storage, and communication with external sensor modules. The specific implementation of interfaces between the onboard computer, UAV sensors, and positioning system depends on the selected hardware configuration and will be investigated during future field deployment of the system.
A 128 GB microSD memory card is proposed for local storage of telemetry and auxiliary monitoring data. The Raspberry Pi 4 supports USB 3.0, USB 2.0, Wi-Fi IEEE 802.11ac, Bluetooth 5.0, and GPIO interfaces, which provide flexibility for communication with peripheral devices and external sensor modules within the proposed monitoring architecture.
Flight mission planning and automated UAV control are to be carried out using DJI Pilot 2 (version v2.5.1.15, SZ DJI Technology Co., Ltd., Shenzhen, China) software, which allows users to create flight routes, set aerial photography parameters, monitor mission execution, and automatically synchronize telemetry data with photographic data. The general technical specifications of the hardware and software configuration proposed for future field deployment are presented in
Table 2.
2.5. Deep Learning Dataset and Model Development
2.5.1. Dataset Preparation
After preprocessing, the original RGB images were automatically divided into non-overlapping patches measuring 512 × 512 pixels, which were used as the basic units for training semantic segmentation models. This approach preserved sufficient spatial context for vegetation-cover analysis while enabling efficient use of computational resources during model training.
Experimental validation was performed using the open-access dataset “High-Resolution RGB Images and Corresponding Masks of Agricultural Fields”, which contains high-resolution individual UAV RGB images and corresponding software-generated reference masks [
42]. Because verified image-level crop labels were unavailable, the dataset was treated as a mixed agricultural vegetation dataset rather than a winter-wheat-specific subset. The 326 retained image–mask pairs yielded 21,324 non-overlapping patches, including 19,999 foreground-dominant and 1325 background-dominant patches.
For the open Zenodo dataset, the supplied reference masks were used directly for model development and internal evaluation; no additional manual pixel-wise re-annotation of the Zenodo masks was performed. The independent Mavic 3M field-validation subset was annotated separately to obtain vegetation reference masks for external model evaluation. The models were trained using a two-class semantic segmentation scheme comprising the following classes:
- (1)
Background/non-target—background and non-target objects;
- (2)
Annotated vegetation region—pixels belonging to the vegetation regions represented by the reference masks.
The total number of generated segments was defined as:
where
Ntrain,
Nval and
Ntest are the number of fragments in the training, validation, and test sets, respectively.
To reduce source-image-level information leakage, a source-image-grouped holdout strategy was used. All patches originating from the same source UAV image were assigned exclusively to the training, validation, or test subset within each repeat. This grouping prevents patches from the same source image from appearing in different subsets; however, it does not guarantee full geospatial independence because spatial overlap between different source UAV images could not be excluded. Verified field identifiers and geospatial metadata sufficient to reconstruct non-overlapping field-level blocks were not available for the open dataset. Therefore, the adopted procedure is referred to throughout the manuscript as source-image-grouped splitting rather than spatially independent splitting (
Table 4).
In each repeated experiment, all 21,324 retained patches were used, with no patch-level subsampling. The grouping procedure ensured that patches derived from the same source image were assigned exclusively to one subset within a given repeat.
2.5.2. Deep Learning Models
Three deep learning architectures were evaluated for semantic segmentation: U-Net-ResNet50, DeepLabV3-ResNet50, and FCN-ResNet50. These models were selected due to their different architectural designs and widespread use in the analysis of high-resolution aerial images.
The FCN architecture implements a fully convolutional approach to pixel-wise image classification without using fully connected layers. U-Net uses an encoder–decoder structure with skip connections, which combine high-level semantic features with detailed spatial information. DeepLabV3 employs atrous convolutions and multiscale context aggregation to capture vegetation patterns at multiple spatial scales.
The source RGB images were divided into 512 × 512-pixel patches and resized to 256 × 256 pixels for model training. To ensure a fair architectural comparison, all three evaluated models used a ResNet-50 encoder/backbone initialized with ImageNet1K V2 pretrained weights. U-Net was implemented as a ResNet50-based encoder–decoder architecture with skip connections, while DeepLabV3-ResNet50 and FCN-ResNet50 used their standard architecture-specific segmentation heads. Thus, the compared models differed primarily in their decoder/head design rather than in backbone initialization or transfer-learning strategy.
The models were trained using a composite loss function combining weighted cross-entropy and Dice loss:
where
LCE—categorical cross-entropy, and
LDice—Dice loss.
Equal weighting was used to balance pixel-wise classification accuracy and spatial overlap during model optimization.
The cross-entropy loss function was defined using the formula:
where
N is the number of pixels;
C is the number of semantic classes, with
C = 2;
yic is the reference-label indicator for pixel
i and class
c;
pic is the corresponding probability predicted by the model after the softmax transformation; and
ε is a small constant introduced for numerical stability.
The Dice Loss function was calculated using the following expression:
where
N is the number of pixels;
C is the number of classes;
pic is the predicted probability for pixel
i and class
c;
yic is the corresponding one-hot encoded reference-label value; and
ε = 10
−6 is used to ensure numerical stability.
The main configurations of the evaluated models are summarized in
Table 5. All architectures used the same 256 × 256 input resolution, ResNet-50 backbone, ImageNet1K V2 backbone initialization, combined Cross-Entropy and Dice loss, and common optimization protocol. Architecture-specific decoder/head components were randomly initialized. This standardization removes the transfer-learning asymmetry present in the previous experimental design and provides a more controlled basis for comparing the segmentation architectures (
Table 5).
2.5.3. Training Strategy
All models were trained under identical experimental conditions using the PyTorch library. To ensure a fair comparison of architectures, we used the same input image size, a common loss function, the same optimization algorithm, and a unified evaluation strategy. All models were trained within the same experimental evaluation framework using the PyTorch library. The same input resolution, loss formulation, optimization algorithm, dataset-splitting strategy, and evaluation metrics were applied. For all three architectures, the ResNet-50 backbone was initialized using the same ImageNet1K V2 pretrained weights, whereas the architecture-specific decoder or segmentation head was randomly initialized.
Parameter optimization was performed using the AdamW algorithm, which combines adaptive learning rate adjustment with weight regularization. The best model was selected based on the maximum Intersection over Union (IoU) value obtained on the validation set:
where
e* denotes the epoch corresponding to the highest validation IoU;
E is the total number of completed training epochs; and
is the validation IoU obtained after epoch
e. The model weights corresponding to
e* were retained as the final checkpoint.
Early stopping was controlled by validation vegetation IoU. Training was required to continue for at least 10 epochs and was allowed to proceed for a maximum of 30 epochs. After the minimum epoch requirement was satisfied, training was terminated only when validation IoU failed to improve for five consecutive epochs. The checkpoint corresponding to the highest validation IoU was retained for evaluation. Thus, an early best epoch did not imply immediate termination of training. The main parameters of the training process are presented in
Table 6.
The comparative evaluation was performed in full-data mode without patch subsampling. To assess the robustness of the results, three repeated source-image-grouped holdout experiments were conducted using random seeds 42, 123, and 2026. Each repeat used all 21,324 retained image patches. For seed 42, the training, validation, and held-out test subsets contained 14,872, 3203, and 3249 patches, respectively; for seed 123, they contained 14,910, 3127, and 3287 patches; and for seed 2026, they contained 14,950, 3251, and 3123 patches. All patches originating from the same source image were assigned to a single subset within each repeat, and no source-image group overlap occurred between the training, validation, and test subsets. Model selection was based on the highest mean validation vegetation IoU across the three repeated experiments, whereas the test subsets remained held out until final evaluation. Performance was summarized using the mean and standard deviation across the three repeats, and uncertainty was additionally assessed using 1000-replicate group-bootstrap 95% confidence intervals.
Data augmentation was applied only to the training subset and was performed synchronously for each RGB image and its corresponding reference mask. Horizontal and vertical flips were independently applied with probabilities of 0.5. Random rotations were selected uniformly from 0°, 90°, 180°, and 270°. In addition, an RGB-only brightness/contrast-like perturbation was applied with a probability of 0.4 using a multiplicative factor uniformly sampled from 0.85 to 1.15 and an additive intensity shift from −10 to +10. Validation and test images were not augmented. After resizing to 256 × 256 pixels, RGB values were normalized using the ImageNet mean and standard deviation.
The use of a unified training strategy and identical optimization parameters ensured a fair comparison of the architectures under study and allowed for an objective assessment of their ability to perform semantic segmentation of high-resolution UAV images.
External field validation was additionally performed using RGB imagery acquired with the DJI Mavic 3 Multispectral over a wheat field during the pilot experiment described in
Section 2.4. A representative set of non-overlapping 512 × 512-pixel image patches was selected from the field imagery and manually annotated to obtain independent vegetation reference masks. The selected U-Net-ResNet50 model, trained exclusively on the open DJI Air 2S dataset, was applied to the Mavic 3M image patches without additional fine-tuning or parameter adjustment. Segmentation performance was evaluated using Precision, Recall, vegetation IoU, and Dice coefficient calculated from globally aggregated pixel-level confusion counts. This procedure was used to assess the transferability of the trained model to independently acquired UAV imagery collected using a different platform and under real field conditions.
2.6. RGB-Based Spatial Interpretation and Decision-Support Algorithm
Following semantic segmentation, an RGB-based spatial analysis procedure was applied to the vegetation regions identified by the best-performing semantic segmentation model. The purpose of this stage was not to diagnose physiological crop stress, but to detect spatial heterogeneity in vegetation appearance and identify areas requiring subsequent field verification. Therefore, the detected regions are consistently referred to as candidate RGB heterogeneity zones. Such zones may reflect canopy gaps, soil exposure, shadows, senescence, crop-density differences, or other variations in RGB appearance and should not be interpreted as evidence of a specific physiological disorder.
For each analyzed image patch, the red (R), green (G), and blue (B) channels were converted from the original 8-bit representation to floating-point values in the range [0,1]. Three visible-band vegetation indices were then calculated: the Visible Atmospherically Resistant Index (VARI), Excess Green Index (ExG), and Green Leaf Index (GLI).
The VARI index was calculated as:
where
Ri,
Gi and
Bi are the normalized red, green, and blue channel values of pixel
i, respectively;
ε = 10
−6 is a small constant introduced to prevent division by zero.
The Excess Green Index was calculated as:
where
represents the relative predominance of the green component for pixel
i.
The Green Leaf Index was determined as:
The calculated VARI, ExG, and GLI values were restricted to the interval [−1,1]. To avoid forcing each image patch into the same relative index range, patch-specific normalization was not used in the final analysis. Instead, fixed normalization bounds were estimated exclusively from vegetation pixels in the training subset of the reference split. For each RGB vegetation index, the 5th and 95th percentiles of the training data were calculated once and subsequently applied unchanged to all held-out image patches:
where
Ik denotes the value of VARI, ExG, or GLI for pixel
p;
and
are the 5th and 95th percentiles of index
k, estimated exclusively from reference vegetation pixels in the training subset; and the clipping operation restricts the normalized value to [0,1].
The same normalization bounds were applied to every validation and held-out test patch. These percentiles serve only as fixed RGB scaling constants and should not be interpreted as physiologically validated thresholds.
The normalized vegetation indices were integrated into a single RGB Appearance Score. The relative contributions of VARI, ExG, and GLI were set to 0.40, 0.35, and 0.25, respectively:
where
,
and
are the robustly normalized values of the corresponding RGB vegetation indices.
Candidate RGB heterogeneity pixels were identified only within the vegetation mask predicted by the semantic segmentation model. A pixel was assigned to the candidate RGB heterogeneity class when its RGB Appearance Score was below 0.30:
where
Mi is the binary candidate RGB heterogeneity mask;
Mcrop denotes the predicted vegetation mask.
The threshold of 0.30 was heuristically specified for the proof-of-concept RGB screening procedure and should not be interpreted as a physiologically validated threshold for a specific type of crop stress.
To suppress isolated noise, the initial candidate RGB heterogeneity mask was processed using an 8-connected component analysis. Connected components with an area smaller than 30 pixels were removed. The remaining connected components were considered individual RGB-based candidate RGB heterogeneity zones.
For each image patch
j, the relative proportion of candidate RGB heterogeneity pixels within the predicted vegetation region was determined as:
where
Nstress,j is the number of pixels retained in the candidate RGB heterogeneity mask, and
Ncrop,j is the total number of pixels classified as vegetation in image patch
j.
For presentation as a percentage, Sj was multiplied by 100%.
The average RGB-based condition of the vegetation within each image patch was characterized by the Mean RGB Appearance Score:
where
represents the Mean RGB Appearance Score for image patch
j.
After filtering, each retained connected component was represented as an image-space vector object characterized by its pixel area, centroid, contour geometry, and bounding box. Patch-level coordinates were transformed to the coordinate system of the corresponding source image using the known patch offsets. The resulting coordinates are expressed exclusively in source-image pixels and are not geographic coordinates. No affine transformation to a projected coordinate reference system was applied to the Zenodo RGB dataset because the required georeferencing metadata were not available. Therefore, the resulting polygons should be interpreted as localized computer-vision vector objects rather than GIS features.
The derived RGB indicators were subsequently integrated into a decision-support procedure for prioritizing image patches requiring further field inspection. For each analyzed image patch
j, an integrated Priority Score was calculated as:
where
Pj ∈ [0,1] is the integrated monitoring priority score;
Sj is the normalized share of candidate RGB heterogeneity pixels;
is the Mean RGB Appearance Score;
Zj is the number of candidate RGB heterogeneity zones detected in image patch
j; and
Zmax is the maximum number of candidate RGB heterogeneity zones observed among all analyzed image patches.
Thus, the Priority Score simultaneously accounts for the relative extent of RGB-based vegetation heterogeneity, the RGB Appearance Score deficit, and the spatial fragmentation of the detected candidate RGB heterogeneity zones.
Based on the calculated Priority Score, image patches were assigned to one of three monitoring-priority classes:
Low-priority patches were assigned to routine UAV monitoring, medium-priority patches were recommended for targeted field inspection and verification of crop condition, and high-priority patches were assigned the highest priority for ground verification and subsequent site-specific agronomic assessment.
The baseline parameters of the RGB-based decision-support procedure were defined as follows: VARI, ExG, and GLI weights of 0.40, 0.35, and 0.25, respectively; an RGB heterogeneity threshold of 0.30; Priority Score weights of 0.50, 0.30, and 0.20; and Low–Medium–High priority boundaries of 0.20 and 0.45. These parameters were treated as heuristic baseline settings rather than agronomically calibrated constants. Their robustness was therefore evaluated through a dedicated sensitivity and ablation analysis described below.
Sensitivity analysis was performed on a deterministic subset of 200 held-out image patches. One-at-a-time perturbations included RGB heterogeneity thresholds of 0.20, 0.25, 0.35, and 0.40; minimum connected-component areas of 60 and 240 source-image pixels; single-index VARI, ExG, and GLI formulations; equal RGB-index weights; removal of individual Priority Score components; alternative priority-class boundaries of 0.15/0.40 and 0.25/0.50; and patch-local normalization as an ablation comparator. In addition, 40 joint perturbation scenarios were generated by simultaneously varying the normalized RGB-index and Priority Score weights by ±20%, the heterogeneity threshold within 0.24–0.36, and the priority boundaries within 0.16–0.24 and 0.40–0.50.
Robustness was assessed using Spearman rank correlation of Priority Scores relative to the baseline, Jaccard similarity of the top-20 ranked patches, the fraction of patches changing priority class, the mean RGB heterogeneity share, the number of detected zones, and the number of high-priority patches. The analysis was used solely to assess robustness of the heuristic formulation and should not be interpreted as empirical agronomic calibration.