2.1. Image-Based Olive Disease Recognition
Deep learning has substantially advanced automatic olive-disease recognition by replacing manually designed descriptors with learned visual representations. Among the earliest studies devoted specifically to olive-leaf diseases, Uğuz and Uysal developed a convolutional neural network for distinguishing healthy leaves, olive peacock spot, and damage caused by
Aculus olearius [
14]. Their results demonstrated the effectiveness of deep feature extraction for olive-disease classification and established an important benchmark for subsequent research. Nevertheless, the task was formulated as conventional image classification, without explicitly representing disease progression, quantifying symptom severity, or incorporating environmental factors associated with fungal development.
To improve the interpretability of image-based diagnosis, Uğuz subsequently introduced a Single Shot Detector framework for localizing peacock-spot lesions on olive leaves [
15]. In contrast to whole-image classification, lesion detection identifies the spatial positions of symptomatic regions and can therefore provide useful visual evidence for the predicted diagnosis. However, object localization does not necessarily quantify the proportion of affected leaf tissue or distinguish among fine-grained disease-severity stages. Moreover, the model remains dependent exclusively on visual information and cannot account for environmental conditions that may influence symptom development.
Bruno et al. proposed an adaptive ensemble of two EfficientNet-B0 models to improve prediction robustness while maintaining computational efficiency suitable for practical agricultural applications [
16]. Their findings showed that combining complementary CNN predictions can provide more stable classification than relying on an individual network. However, the framework remained focused on categorical image recognition and did not incorporate soil, meteorological, vegetation, or moisture-related variables.
A different balance between predictive performance and computational cost was investigated by Ksibi et al., who developed the MobiRes-Net architecture by combining characteristics of MobileNet and ResNet50 [
17]. The proposed lightweight model demonstrated the feasibility of deploying olive-disease recognition systems on resource-constrained platforms. Despite this practical advantage, the diagnostic formulation was still single-modal and classification-oriented, without continuous severity estimation or explicit modeling of environmental context.
Recent studies have also extended olive-disease monitoring from individual leaves to orchard-scale imagery. Sarantakos et al. combined unmanned aerial vehicle imagery with convolutional neural networks and Vision Transformers to support large-scale monitoring of olive orchards [
18]. UAV-based acquisition can increase spatial coverage and reduce the effort required for manual field inspection. Nevertheless, aerial imagery primarily captures visual canopy-level manifestations and does not, by itself, explicitly represent the soil and atmospheric conditions associated with pathogen development.
Collectively, these studies confirm that CNNs, lightweight architectures, ensembles, object detectors, UAV imaging, and Vision Transformers can provide accurate and operationally useful olive-disease recognition. However, most existing approaches remain centered on visual diagnosis and produce either categorical disease labels or lesion locations. Comparatively less attention has been given to the joint prediction of fine-grained disease stages, lesion coverage, and leaf-yellowing percentage, particularly within a framework that combines visual evidence with environmental descriptors.
Table 1 summarizes the main differences between representative image-based olive-disease studies and the framework developed in this work.
2.2. Multimodal Artificial Intelligence for Plant Disease Diagnosis
Plant-disease development is influenced by interacting visual, biological, meteorological, and edaphic factors. A single image may capture visible symptoms at the time of acquisition, but it does not necessarily represent the environmental conditions that promoted infection, symptom expansion, or plant stress. Multimodal artificial intelligence addresses this limitation by combining images with complementary information such as environmental measurements, soil properties, temporal sensor observations, textual symptom descriptions, multispectral data, thermal imagery, or remote-sensing products.
Lee et al. proposed a multimodal crop-disease diagnosis framework that combines RGB images with environmental measurements acquired in smart-farm environments [
19]. Their approach used meteorological observations, including temperature, humidity, and dew point, collected at short temporal intervals before image acquisition. Visual features were extracted using a convolutional neural network, whereas the temporal environmental measurements were modeled using a Long Short-Term Memory network. The fusion of the two modalities improved disease-recognition performance relative to an image-only formulation, demonstrating that environmental measurements can provide complementary diagnostic information.
The present study shares the general principle that visual symptoms should be interpreted together with environmental context; however, its data formulation and prediction objectives differ from those of Lee et al. Their framework relied on locally acquired measurements temporally synchronized with individual images, while the present study constructs environmentally enriched image records using regional and seasonal descriptors obtained from public repositories. These repositories include NASA Power for atmospheric variables, SoilGrids and WoSIS for soil-related properties, Copernicus/ERA5 products for climatic and moisture-related information, and Sentinel-2 products for vegetation-related indicators. Accordingly, the environmental variables used in the present work should be interpreted as contextual regional–seasonal covariates rather than as direct sensor readings taken simultaneously from the exact tree represented in each image.
The two studies also differ in predictive scope. Lee et al. primarily addressed general crop-disease classification, whereas the present framework targets olive peacock spot and simultaneously predicts a seven-stage disease category, lesion coverage percentage, and leaf-yellowing percentage. Furthermore, the present work includes an explicit workflow for environmental-variable extraction, feature engineering, image–environment matching, feature selection, duplicate control, and prevention of information leakage during dataset partitioning.
Wu et al. introduced PlantIF, a multimodal semantic interactive fusion framework that integrates plant images with textual semantic information through graph-based reasoning [
20]. In PlantIF, the non-image modality represents semantic knowledge related to disease names, botanical characteristics, and symptom descriptions. This approach improves the interaction between visual observations and conceptual disease knowledge. However, textual semantics and numerical environmental covariates serve fundamentally different purposes. PlantIF emphasizes semantic reasoning between images and disease descriptions, whereas the present study investigates whether meteorological, soil, vegetation, moisture, and stress-related descriptors can complement visual features in olive peacock spot assessment.
Albahli proposed AgriFusionNet, a lightweight multimodal architecture integrating RGB or multispectral imagery with environmental and Internet-of-Things sensor information for plant-disease recognition [
21]. The model employed an EfficientNetV2-B4-based design and demonstrated that heterogeneous sensing modalities can be combined while maintaining comparatively efficient computation. This work further supports the relevance of multimodal learning in precision agriculture. Nevertheless, its objective was general plant-disease recognition rather than olive-specific, fine-grained, multi-task severity assessment.
Other multimodal agricultural studies have explored combinations of RGB, hyperspectral, multispectral, thermal, remote-sensing, and sensor-based information. These studies collectively demonstrate that heterogeneous modalities may describe complementary aspects of crop condition. At the same time, their applicability depends strongly on the spatial and temporal correspondence among modalities. Directly synchronized measurements provide strong sample-level correspondence but require dedicated field infrastructure. Public geospatial and climatic repositories offer greater accessibility and reproducibility, although their variables generally represent contextual conditions at regional, grid, or seasonal scales rather than exact tree-level measurements.
The proposed framework belongs to the latter category. It does not claim that the environmental values constitute synchronized measurements from each photographed leaf. Instead, it evaluates a proof-of-concept strategy in which real-source environmental descriptors are matched to olive-leaf images according to available regional and temporal information. This formulation is intended to investigate the potential value of environmental context while maintaining a transparent distinction between contextual matching and direct sensor synchronization.
Table 2 compares the proposed framework with representative multimodal plant-disease approaches.
2.3. Critical Synthesis, Research Gap, and Study Motivation
The reviewed literature reveals substantial progress in automated plant- and olive-disease diagnosis. Deep CNNs have improved visual feature extraction, transfer learning has reduced the need for training models entirely from scratch, ensemble methods have strengthened predictive stability, lightweight architectures have supported deployment on constrained devices, and UAV- and Transformer-based systems have extended disease monitoring to larger spatial scales. Multimodal plant-disease studies have further demonstrated that non-image information can complement visual evidence when the additional modalities are appropriately selected and matched.
Despite this progress, three interconnected limitations remain particularly relevant to olive peacock spot assessment.
First, representative olive-disease systems are predominantly image-based. Such models can learn discriminative patterns associated with visible lesions, discoloration, and texture changes, but they do not explicitly include the environmental conditions associated with disease development. This omission is important because fungal infection and symptom progression are influenced by interacting atmospheric, soil, moisture, and vegetation conditions. Environmental information should not replace image evidence, but it may provide complementary contextual signals that improve the representation of disease state.
Second, most previous olive-disease studies formulate diagnosis as binary classification, multi-class disease recognition, or lesion detection. These outputs are useful for identifying the presence or type of disease, but they provide limited information about progressive symptom intensity. In practical disease management, distinguishing among multiple levels of infection may be more informative than producing a single infected/non-infected label. Moreover, continuous estimates of lesion coverage and leaf yellowing can provide quantitative information that is not fully represented by a categorical disease stage alone.
Third, relatively few studies combine fine-grained classification and continuous symptom estimation within one multimodal, multi-task framework. Disease stage, lesion coverage, and yellowing percentage are related but non-identical indicators. Joint learning can allow the shared representation to capture common disease characteristics while preserving task-specific outputs. However, such a formulation requires careful control of label consistency, image duplication, feature-selection leakage, and dataset partitioning.
A further methodological gap concerns the use of publicly accessible environmental repositories. Smart-farm systems can provide synchronized sensor measurements, but such infrastructures are not available for many existing plant-image datasets or production environments. Public sources such as NASA Power, SoilGrids, WoSIS, Copernicus/ERA5, and Sentinel-2 provide an alternative source of real environmental information. Nevertheless, matching these products to images must be described cautiously because the resulting variables represent contextual regional–seasonal conditions rather than exact measurements from the photographed tree. The scientific value of this approach therefore depends on transparent matching rules, leakage prevention, and an explicit acknowledgement of its spatial and temporal limitations.
Motivated by these gaps, the present study develops an environmentally enriched multimodal deep-learning framework for olive peacock spot assessment. The framework combines RGB olive-leaf images with selected meteorological, soil, moisture, vegetation, and environmental-stress descriptors through feature-level fusion between a ResNet50 image branch and a multilayer perceptron numerical branch. The fused representation is optimized within a multi-task architecture to produce three complementary outputs:
- 1.
Classification of olive peacock spot into seven ordered disease-severity stages;
- 2.
Regression-based estimation of lesion coverage percentage; and
- 3.
Regression-based estimation of leaf-yellowing percentage.
The contribution of the study is therefore not limited to replacing one image-classification backbone with another. Instead, it investigates a broader diagnostic formulation that combines fine-grained staging, continuous symptom quantification, and environmental contextualization. In addition, the study establishes a reproducible workflow for constructing an environmentally enriched real-source fused dataset, selecting numerical features using training data only, controlling near-duplicate images before dataset partitioning, and comparing the complete fusion model with appropriate single-modal and reduced-feature baselines.
The proposed framework should be interpreted as a proof of concept rather than as evidence of universal field generalization. Its regional–seasonal environmental matching provides contextual information but does not substitute for synchronized tree-level sensing. Consequently, external evaluation across independent orchards, seasons, cultivars, imaging devices, and environmental conditions remains necessary. Future synchronized data collection will also be important for determining the extent to which tree-level sensor measurements improve upon regional and seasonal environmental descriptors.
Within these stated boundaries, the study addresses a comparatively underexplored intersection of olive-disease recognition, environmental data fusion, fine-grained severity staging, and multi-task quantitative symptom estimation.