1. Introduction
Growing concerns regarding global food security, coupled with the increasing frequency of climate-induced agricultural disruptions, have intensified the need for reliable, operational crop monitoring systems [
1,
2]. Satellite-based crop mapping is considered an essential tool for agricultural management, policy formulation, and food security assessment, with numerous national and international programs relying on timely and accurate crop classification data [
3]. The unprecedented data availability from the European Space Agency’s Sentinel constellation, particularly the synergistic combination of Sentinel-1’s all-weather radar capabilities and Sentinel-2’s high-resolution optical imagery, has significantly enhanced our capacity to monitor agricultural landscapes at relevant spatio-temporal scales [
4,
5].
While deep learning and ensemble methods have achieved high classification accuracies within single agricultural seasons, a critical knowledge gap remains: how to maintain classification performance when models trained on one year are applied to subsequent years without retraining. This ‘temporal domain shift’ problem, in which phenological timing, weather conditions, and crop development patterns vary inter-annually, leads to substantial accuracy degradation in operational settings. Current approaches either require annual retraining with newly collected field data (which is costly and often infeasible) or rely on domain adaptation techniques that have been primarily developed for spatial transfer rather than temporal generalization. In remote sensing, domain adaptation has been successfully applied to reduce discrepancies between different sensors, geographical regions, or acquisition conditions [
6]. However, adapting across years remains challenging because the shift is not simply a matter of different marginal distributions; the conditional distribution of crop phenology given the time of year changes due to climate variability. Recent efforts have explored temporal shift estimation [
7] and phenology alignment [
8] for crop mapping, yet a systematic approach that selects stable features across years rather than aligning distributions after the fact has been lacking. Our study directly addresses this gap by proposing a stability-driven framework that explicitly optimizes for inter-annual feature consistency, thereby reducing the need for year-specific model retraining.
Despite these technological advances, operational crop mapping faces a persistent and critical challenge: classification models meticulously optimized for one growing season often demonstrate substantially degraded performance when applied to subsequent years [
9]. This “temporal generalization gap” represents a fundamental obstacle to deploying sustainable agricultural monitoring systems, particularly in regions characterized by high inter-annual climate variability [
10]. The problem is especially acute in Mediterranean agricultural systems like Morocco’s, where rainfall variability, temperature fluctuations, and shifting phenological patterns create dynamic growing conditions that challenge static classification approaches [
11,
12].
In the last decade, the remote sensing community has made substantial progress in developing sophisticated machine learning algorithms for crop classification, with ensemble methods emerging as particularly effective for handling complex, multi-sensor data [
13,
14,
15]. Random Forest classifiers have become a widely adopted standard due to their robustness and built-in feature importance metrics, while Voting and Stacking ensembles further improve generalization by leveraging diversity across multiple algorithms [
16]. Concurrently, researchers have identified numerous spectral, temporal, and polarimetric features that contribute to crop discrimination, including vegetation indices, texture measures, and phenological metrics [
17]. However, current feature selection practices remain predominantly static, either identifying “optimal” features from single-year studies [
18] or aggregating multi-year data in ways that mask important annual variations [
19]. This creates a fundamental methodological disconnect: models are optimized for historical performance rather than future robustness in operational scenarios [
20]. The critical question “Which features are most important?” should be complemented with “How stable is this importance across different environmental conditions and growing seasons?” [
21]. Feature stability, the consistency of importance rankings across temporal samples and environmental contexts, has gained considerable attention in other machine learning domains [
22] but remains conspicuously under-explored in operational remote sensing applications. While several studies have incidentally noted inter-annual variations in feature importance [
23], none have systematically analyzed this phenomenon as a central research question or developed practical frameworks for leveraging stability analysis in the design of automated, production-ready systems [
24].
This gap is critical because feature importance instability likely reflects fundamental ecological and agronomic processes [
25]. Crop responses to inter-annual environmental variations, including water availability, temperature regimes, and management practices, manifest in modified spectral signatures and phenological patterns [
26]. Consequently, the relative importance of different remote sensing features for crop discrimination naturally fluctuates with environmental conditions [
27]. Understanding these dynamics is not merely an academic exercise but a practical necessity for building resilient monitoring systems capable of automated, multi-annual deployment.
Recent advances in ensemble learning provide a promising pathway forward. By leveraging consensus across multiple models, Voting Ensemble approaches can provide more robust importance estimates than single classifiers, filtering out algorithm-specific noise and revealing the underlying phenological signal [
28]. Furthermore, novel composite metrics that explicitly balance predictive power with temporal stability, the Reliability Index (RI) and Automatic Selection Score (AuSS) can transform stability analysis from a diagnostic tool into a prescriptive framework for operational system design.
This study addresses these critical gaps by proposing and validating a stability-driven framework for automated crop classification. We conduct a comprehensive investigation of feature importance stability across six agricultural years (2018, 2019, 2020, 2023, 2024 and 2025) in Morocco’s diverse agricultural landscapes, analyzing 156 multi-sensor features across 12 monthly composites, derived from 13 spectral indices (9 optical and 4 radar). We integrate statistical stability metrics, hierarchical clustering, and novel composite indices, the Reliability Index (RI) and Automatic Selection Score (AuSS), to move beyond observation toward prescriptive system design. This work is guided by three fundamental research questions: (1) To what extent do feature importances remain stable across growing seasons, and what phenological and sensor-specific patterns characterize this stability? (2) What is the inherent trade-off between feature discriminative power and temporal stability, and how can it be quantified to inform feature portfolio construction? (3) How does a stability-optimized feature portfolio influence cross-year generalization performance and model certainty, and what are the implications for operational automation?
Our work makes several distinct contributions to the field of agricultural remote sensing. Methodologically, we introduce a metrics-driven stability framework that integrates novel indices (RI, AuSS) with clustering and Pareto analysis to filter noise and identify robust feature portfolios. Empirically, we provide the first systematic phenological mapping of feature stability using multi-model agreement and quantify the efficiency of a pared-down, stability-optimized feature set that reduces the feature space to 6 indices (VH, VV, NDVI, NDRE, GCVI, RVI) capturing 57.2% of cumulative predictive importance. Practically, we demonstrate that this framework enables the design of classifiers that maintain high, consistent accuracy (87.4% accuracy, 87.2% F1-score) with a generalization gap reduced from 18.4% to 10.5%, providing a direct pathway to climate-resilient, automated crop mapping. Ultimately, our findings challenge conventional feature selection paradigms and offer a new operational pathway toward robust, reliable crop mapping in our era of accelerating environmental change. In the context of crop monitoring, an operational system should ideally maintain consistent accuracy across years without retraining, process large areas efficiently, and require minimal ground data.
3. Results
3.1. Phenological Patterns of Feature Stability
Stability, defined as the inverse of the coefficient of variation (1/CV), is computed at the feature level across years. As a result, all monthly instances of a given spectral index share the same stability value, since stability reflects inter-annual variability rather than intra-annual dynamics. Consequently, grouping by phenological stages (sowing, growth, maturation, harvest) yields identical stability distributions, while substantial differences emerge between indices.
Figure 3 illustrates these differences through a boxplot grouping indices by their physical category. The Phenology group (NDRE) exhibits the highest stability, with a nearly degenerate distribution reflecting both its strong temporal consistency and the limited number of features in this category. Water Stress indices (NDMI, MSI) and Canopy Structure indices (EVI, SAVI) follow, showing relatively high and consistent stability. In contrast, Radar and Vegetation Vigor indices (NDVI, GNDVI, GCVI) display lower median stability and greater dispersion, with the latter showing the widest spread, indicating higher sensitivity to inter-annual climatic variability.
Table 12 further quantifies this hierarchy by ranking the 13 indices according to their stability. The red-edge index NDRE emerges as the most stable feature (5.179), followed by water stress indices NDMI (4.251) and MSI (3.294). At the lower end, vegetation vigor indices such as GNDVI (2.133) and NDVI (3.461) exhibit higher variability, reflecting their sensitivity to environmental fluctuations. Radar-based features (VH, VV, Ratio, RVI) occupy an intermediate range (2.866–3.654), supporting their role as relatively stable structural indicators across years.
3.2. The Fundamental Trade-Off: Importance vs. Stability
The relationship between predictive importance (mean importance across years) and temporal stability is illustrated in
Figure 4. The scatter plot is divided by the median importance (344.25) and median stability (3.3423), defining four distinct behavioral quadrants:
Dominant Stable (high importance, high stability): BSI, GCVI, NDRE, NDVI, RVI. These indices form the ideal core for an automated system, combining strong discriminative power with consistent inter-annual performance.
Performant Volatile (high importance, low stability): VH, VV. They provide excellent predictive capability, but are highly sensitive to climatic variability, requiring conditional use.
Stable Minor (low importance, high stability): NDMI, SAVI. They serve as structural anchors, offering reliable but modest contributions.
Noise/Unstable (low importance, low stability): EVI, GNDVI, MSI, Ratio. These indices contribute little and vary unpredictably, making them clear candidates for pruning.
This four-quadrant typology, derived from Voting Ensemble consensus, provides a framework for feature selection: retain Dominant Stable indices as the foundational portfolio, employ Performant Volatile indices adaptively, keep Stable Minor indices as baseline anchors, and eliminate the Noise group to reduce overfitting and improve generalization.
3.3. Confusion Matrix of the Stable Portfolio
To complement the stability-driven analysis,
Table 13 presents the confusion matrix (row-wise normalized percentages) for the Stable Portfolio model, averaged over the three test years (2019, 2020, 2024). Values indicate the percentage of true class *c* (rows) predicted as each class (columns).
The most frequent misclassifications occur between Soft Wheat and Durum Wheat (≈11–14%), reflecting their spectral-phenological similarity. Tree crops are the best discriminated class (94.5% accuracy), while “Other Crops” (including vegetables, fallow, and weeds) show moderate confusion with cereal classes (≈9% total). A detailed multi-model classification benchmark, including per-class F1-scores, ROC curves, and year-wise confusion matrices, is provided in a companion study focusing on classifier optimization [
36].
3.4. Top Stable Pillars for Operational Automation
To operationalize this trade-off, we ranked all features by the Reliability Index (RI), defined as the ratio of mean importance to the coefficient of variation, defined as follows:
This metric directly identifies features that balance predictive power with temporal consistency.
Table 14 presents the top-ranked features according to RI. The red-edge index NDRE ranks first, confirming its strong combination of high importance and low variability. Radar features (VH, VV, RVI) also occupy prominent positions, highlighting their relatively consistent behavior across varying environmental conditions. Vegetation indices such as NDVI and GCVI appear among the top-ranked features, indicating that certain optical signals retain both discriminative power and acceptable temporal stability. These results suggest that early-stage spectral responses may be more robust than commonly assumed. Overall, the RI-based ranking confirms that combining importance and stability provides a more reliable basis for feature selection than considering either criterion independently.
3.5. Sensor Complementarity: Synergy of Optical and Radar
The correlation matrix of the 13 indices (
Figure 5) reveals two distinct patterns. First, strong intra-family correlations indicate redundancy among features derived from similar physical processes. For example, MSI and NDMI exhibit a high correlation (r = 0.88), as do RVI and Ratio (r = 0.88), suggesting that one feature from each pair may be sufficient. Second, cross-correlations between radar and optical indices remain consistently low (|ρ| < 0.3), confirming that the two sensor families provide complementary information. Optical indices primarily capture biochemical properties such as chlorophyll content and leaf area index, whereas radar backscatter is sensitive to structural and dielectric properties of the canopy.
To further investigate sensor complementarity, a multi-criteria robustness comparison was performed using five criteria: Stability, Importance, Availability, Timeliness, and Weather Independence (
Figure 6). Stability and Importance were obtained directly from the quantitative feature-analysis framework, where Stability corresponds to the mean Stability Index and Importance corresponds to the mean feature importance across years. The remaining criteria represent operational characteristics of the sensor families. Availability reflects the consistency of usable observations throughout the growing season, Timeliness reflects the ability to support early-season crop discrimination, and Weather Independence reflects resilience to cloud contamination and adverse atmospheric conditions. To facilitate comparison among heterogeneous criteria, all metrics were normalized to a common 0–1 scale. While Stability and Importance are derived from quantitative analysis, the operational criteria provide complementary contextual information intended to summarize the practical advantages and limitations of radar and optical observations in agricultural monitoring.
After normalization to a common scale, radar and optical sensor families exhibit distinct yet complementary robustness patterns. Radar features achieve higher scores for Availability, Timeliness, and Weather Independence, reflecting their capacity to provide consistent observations regardless of cloud cover and support crop monitoring from the earliest stages of the growing season. Optical features exhibit slightly higher Stability and contribute detailed spectral and phenological information that is particularly valuable during key crop development stages. These results suggest that neither sensor family alone fully captures the diversity of information required for robust crop classification. Instead, their integration provides a balanced framework in which radar data ensures operational continuity and resilience, while optical observations contribute complementary phenological and spectral detail under favorable acquisition conditions.
3.6. Temporal Trajectory of Predictive Information
Feature importance is not static but varies throughout the agricultural season, reflecting changes in crop development and environmental conditions. To capture this dynamic behavior,
Figure 7 presents the temporal evolution of mean consensus importance for radar and optical features, with shaded regions indicating inter-annual variability (±1 standard deviation).
Radar features exhibit consistently high importance across the entire season, with noticeable increases during key transition periods such as early growth, mid-season, and harvest. This suggests that radar provides a stable and continuous source of predictive information, largely independent of seasonal constraints. In contrast, optical features display a more variable trajectory, with importance increasing during periods of active vegetation development and declining during early and late stages of the season. This behavior reflects the sensitivity of optical signals to vegetation dynamics and their dependence on favorable acquisition conditions. These contrasting temporal patterns highlight the complementary roles of radar and optical data, with radar ensuring consistent baseline performance and optical features contributing enhanced discriminative power during specific phenological stages.
3.7. Pareto Optimization: Pruning Noise, Preserving Signal
The consensus importance derived from the Voting Ensemble across six years enables the identification of diminishing returns in feature accumulation. To determine an optimal feature subset for automated classification, all 156 features were ranked according to their Automatic Selection Score (AuSS), and the cumulative proportion of total predictive importance was computed as a function of the number of selected features.
The resulting Pareto curve (
Figure 8) reveals a clear diminishing return behavior, where the initial features contribute disproportionately to the overall predictive signal. A distinct inflection point is observed at approximately six features, which together account for about 57% of the cumulative importance. These six features correspond to spectral indices (VH, VV, NDVI, NDRE, GCVI, RVI). For each selected index, all 12 monthly composites (September to August) are retained, resulting in a Stable Portfolio of 72 monthly features (6 indices × 12 months). The cumulative importance percentage (57.2%) refers to the index-level importance—i.e., the sum of the importances of the 12 monthly composites for each index, normalized by the total importance of all 156 monthly features. Thus, the reduction in the number of monthly features is also 54% (from 156 to 72). Beyond this point, the marginal contribution of additional features decreases substantially, indicating increasing redundancy within the feature space. This suggests that a relatively small subset of features captures the majority of stable and informative signal, while the remaining features contribute limited additional value. These findings highlight the effectiveness of the AuSS-based ranking in identifying a compact and efficient feature subset, supporting the development of streamlined and generalizable automated classification systems.
3.8. Automatic Selection Score: Gateway for Transferability
The Automatic Selection Score (AuSS) is defined as follows:
combines predictive importance with a logarithmic stability penalty.
Figure 9 presents AuSS values by month for both sensor families, with the global automation threshold (mean AuSS = 527.0) indicated by a red dashed line.
Radar features maintain consistently high AuSS values throughout the agricultural season, frequently exceeding the threshold, with notable peaks observed in March (~700), mid-season (~800), and December (>1100). This indicates that radar-derived features combine strong predictive power with temporal robustness, making them suitable for automated use across most periods. In contrast, optical features exhibit a more variable pattern, remaining below the threshold for most of the season and exceeding it only during a limited period between September and November, with a peak around 700. This reflects their dependence on vegetation dynamics and acquisition conditions. These results further highlight the complementary roles of radar and optical data: radar provides a stable baseline for automated feature selection, while optical features contribute selectively during periods of increased discriminative power. The threshold therefore serves as a practical criterion for identifying features that are both informative and reliable across time.
The automation threshold (dashed red line) is set at the global mean AuSS across all months and sensor types (527.0). Features that consistently exceed this threshold are considered suitable for automated, zero-intervention use.
3.9. Feature Certainty and Entropy
Beyond stability and predictive importance, the clarity of the signal provided by each feature is critical for reducing model uncertainty in operational settings. To assess this aspect, we computed the information entropy of each index based on its normalized importance distribution across years. Lower entropy values indicate more consistent and less ambiguous behavior, while higher values reflect greater variability and uncertainty.
Table 15 summarizes the entropy values alongside stability and reliability metrics. The entropy values for all indices fall within a very narrow range (2.39–2.47) corresponding to a difference of only 0.08. This small range indicates that all features provide essentially the same level of informational consistency. No statistically meaningful difference is observed between indices (e.g., between GNDVI and NDRE). Therefore, we conclude that the Stable Portfolio maintains consistent prediction entropy across all selected indices; the observed variations are too small to be meaningfully interpreted. This consistency, rather than differences between indices, demonstrates that stability-driven selection reduces model uncertainty regardless of which index enters the portfolio.
Overall, entropy provides a complementary perspective on feature robustness, highlighting the trade-off between consistency and discriminative power in automated classification systems.
3.10. Voting Ensemble Performance with Stable Portfolio
Using the Stable Portfolio (6 indices) as input, we trained a Voting Ensemble (Random Forest, Extra Trees, XGBoost, and LightGBM) on 3773 training samples and evaluated it on 953 independent test samples. The results are summarized in
Table 16. The model achieved an overall test accuracy of 87.4%, with a macro-averaged F1-score of 87.2%. Class-wise performance is consistently high across all crop types, with F1-scores ranging from 0.86 to 0.90, indicating balanced classification performance. The overfitting gap between training accuracy (94.8%) and cross-validation accuracy (88.6%) is limited to 6.2%, suggesting that the stability-driven feature selection effectively reduces sensitivity to year-specific variability.
3.11. Cross-Year Generalization and Statistical Significance
To evaluate temporal transferability, we compared the Stable Portfolio (72-feature set derived from the 6-index core: VH, VV, NDVI, NDRE, GCVI, RVI, each with 12 monthly values) against the full 156-feature set using the nine train-test combinations (training on one year, testing on all others).
Table 17 reports the improvements achieved by the Stable Portfolio.
The generalization gap narrows from 18.4% to 10.5%, a 43% relative reduction, directly addressing the “temporal generalization gap” that continues to challenge operational remote sensing. The worst-case accuracy improves by 13.5%, providing a critical safety margin against anomalously challenging years.
Statistical significance was assessed with a Friedman test followed by Nemenyi post hoc comparisons (
Table 18). The Friedman test revealed significant differences among subsets (χ
2 = 24.36,
p < 0.001). The Stable Portfolio significantly outperforms both the Full Feature Set (
p < 0.01) and the Volatile Set (
p < 0.001), validating the effectiveness of stability-based selection. Non-significant differences with the Top-RI and Top-AuSS sets confirm that these metrics are effective selection criteria.
Together, these results demonstrate that the stability-driven framework identifies features capturing stable phenological signals rather than year-specific noise, enabling robust, climate-resilient crop classification.
3.12. Spatial Application and Stable Portfolio Validation
To evaluate the spatial transferability of the proposed framework, the Voting Ensemble classifier was applied to generate high-resolution (10 m) annual crop maps over four representative irrigated perimeters (
Figure 10). The maps demonstrate the model’s ability to produce spatially coherent crop patterns across different agricultural landscapes.
In parallel, we assessed the impact of using a Stable Portfolio of features selected based on inter-annual stability (low coefficient of variation) as the sole input feature space. The Voting Ensemble trained on this Stable Portfolio achieved a test accuracy of 87.4% with a macro-averaged F1-score of 87.2% across the five crop classes (
Table 15). Per-class F1-scores were 90.0% for Soft Wheat, 88.0% for Durum Wheat, 86.0% for Barley, 92% for Trees, and 90.0% for Other Crops, with an overfitting gap below 10%. Statistical significance of performance differences between feature subsets was assessed using paired t-tests and the non-parametric Friedman test with post hoc Nemenyi analysis (α = 0.05), confirming that the stability-driven feature selection does not compromise accuracy while enhancing model parsimony.
The spatial maps in
Figure 10 are intended as a visual illustration of the Stable Portfolio’s outputs, not as a rigorous spatial validation. Detailed spatial accuracy assessment, including per-perimeter confusion matrices and area-level validation against official statistics, is provided in our companion paper focusing on classification benchmarking [
43].
4. Discussion
4.1. The Importance-Stability Trade-Off: A Fundamental Constraint for Automation
This study demonstrates that feature importance in crop classification is fundamentally dynamic, and that the conventional paradigm of identifying a single “optimal” feature set from one or two years of data must be replaced by a stability-aware framework. Our analysis across six agricultural years (2018, 2019, 2020, 2023, 2024 and 2025) reveals a consistent and quantifiable trade-off: the most predictive features, notably optical indices during peak vegetation vigor (NDVI, GCVI) and radar backscatter (VH, VV), exhibit the highest inter-annual volatility, while the most stable features (NDRE, NDMI, MSI) offer only moderate predictive power [
44,
45]. This finding challenges the common assumption that a feature’s importance in a single year is a reliable indicator of its utility for long-term operational systems [
46].
The four-quadrant behavioral typology (
Figure 4) goes beyond simple ranking by offering a strategic decision framework: Dominant Stable indices (BSI, GCVI, NDRE, NDVI, RVI) become the core of an automated system; Performant Volatile indices (VH, VV) are used conditionally; Stable Minor indices (NDMI, SAVI) serve as structural anchors; and Noise/Unstable indices (EVI, GNDVI, MSI, Ratio) are pruned. This typology aligns with ecological theory that vegetation responses to climate are inherently dynamic and that effective remote sensing indicators must balance sensitivity to biophysical changes with resilience to inter-annual noise [
47,
48]. By aggregating importance estimates from four distinct tree-based algorithms (RF, ET, XGBoost, LightGBM), the Voting Ensemble consensus ensures that this typology reflects genuine phenological signals rather than algorithm-specific artifacts [
49].
Critically, the typology reveals that no single feature excels in both criteria. This is not a limitation of our dataset but a fundamental property of remote sensing of vegetation: features that are highly sensitive to biophysical changes (e.g., NDVI at peak greenness) are inevitably also sensitive to the climatic drivers that cause inter-annual variability. Recognizing this trade-off forces a shift from optimizing for peak performance to designing for robustness, a concept well established in other domains (e.g., finance, engineering) but rarely applied in agricultural remote sensing.
4.2. Sensor Complementarity as the Key to Year-Round Automation
A key contribution of this study is the quantitative demonstration that Sentinel-1 radar and Sentinel-2 optical data are not redundant but fundamentally complementary [
50]. The consistently low cross-correlation between the two families (|ρ| < 0.3,
Figure 3) confirms that they capture orthogonal aspects of crop canopies: radar responds to dielectric properties and canopy geometry, while optical indices track chlorophyll content and leaf area index [
51]. This complementarity is often assumed but rarely quantified with multi-year data.
More importantly, the temporal trajectory (
Figure 5) and the AuSS by month (
Figure 9) reveal when each sensor should dominate an automated pipeline. Radar maintains high importance and high AuSS throughout the year, with peaks in March, July, and December, periods corresponding to early growth, mid-season stress detection, and harvest. Optical indices, in contrast, become critical only during September–November, the post-harvest window. This pattern provides a direct operational rule: during sowing, growth, and harvest, rely on radar’s stability; during the narrow phenological window where optical indices exceed the automation threshold, integrate them for their high discriminative power. This phenologically adaptive strategy is a departure from static sensor-weighting approaches [
52,
53] and is essential for achieving year-round reliability.
The robustness fingerprint (
Figure 6) further quantifies this complementarity: radar excels in stability, availability, timeliness, and weather independence, while optical indices dominate only in predictive importance. This profile reinforces the concept of a robust baseline (radar) combined with a high-resolution but conditionally applied sensor (optical)—a design that is both theoretically grounded and practically feasible [
54,
55]. In the context of operational monitoring, such a design reduces the risk of system failure during periods of cloud cover or atmospheric disturbances, a common problem in Mediterranean climates.
4.3. The Stable Portfolio: A Pareto-Optimal Foundation for Automation
Pareto optimization (
Figure 8) shows that 6 indices—VH, VV, NDVI, NDRE, GCVI, RVI—capture 57.2% of cumulative importance, while the remaining 7 indices contribute a marginal additional signal. This 54% reduction in the number of indices (from 13 to 6) demonstrates that conventional full-feature approaches carry substantial redundancy and overfitting risk [
56,
57]. The Stable Portfolio identified by this method is not merely a subset; it is a balanced ensemble that includes three radar indices (VH, VV, RVI) and three optical indices (NDVI, NDRE, GCVI), reflecting the sensor complementarity discussed above.
Unlike conventional feature selection that maximizes within-year accuracy [
58], our Automatic Selection Score (AuSS) explicitly balances predictive importance with a logarithmic stability penalty, ensuring that selected features are both powerful and reliable across diverse climatic years. The Reliability Index (RI) provides an even more direct measure of this trade-off and ranks NDRE, VH, and VV as the top three pillars (
Table 13). The entropy analysis (
Figure 8) adds an additional layer: low-entropy features (NDRE, NDMI, SAVI) provide unambiguous signatures that ground the model’s decisions, reducing vulnerability to edge cases and atmospheric perturbations [
59]. Together, RI, AuSS, and entropy form a multi-dimensional screening process that yields a portfolio optimized for operational deployment. Similar dimensionality reductions have been successfully applied in large-area crop mapping [
60].
Importantly, the Stable Portfolio is not merely a list of indices; it is a dynamic concept. The inclusion of both radar and optical indices, and the explicit use of monthly composites, means that the portfolio inherently captures the phenological shifts in feature importance. This is a departure from static feature sets that are applied uniformly across the entire season.
4.4. The Automated Model: Voting Ensemble with Stable Portfolio
The Voting Ensemble trained on the Stable Portfolio achieved 87.4% test accuracy and a macro-averaged F1-score of 87.2% across five crop classes (
Table 15). Critically, the per-class performance is balanced (F1 ranging from 0.86 for Barley to 0.90 for Soft Wheat, Trees, and Other Crops), indicating that the Stable Portfolio captures discriminative information for all classes without favoring spectrally distinct ones [
52]. The slight underperformance of Barley reflects its higher phenological plasticity and wider cultivation range, factors that increase intra-class spectral variability, a known challenge in Mediterranean environments [
61].
The primary misclassifications occurred between spectrally similar classes (Durum Wheat/Soft Wheat, Trees/Other Crops), a pattern consistent with the inherent spectral similarity of these categories [
62]. This suggests that further improvements may require additional data sources (e.g., time-series phenological metrics) rather than more feature engineering within the current sensor suite. It also highlights that the Stable Portfolio, while robust, does not eliminate all confusion; rather, it focuses on features that are stable and generally discriminative, accepting that some spectral overlap is unavoidable without additional information.
Most importantly, the overfitting gap between training (94.8%) and cross-validation (88.6%) accuracy is only 6.2%, suggesting that stability-driven feature selection effectively eliminates year-specific noise [
63]. This low gap is a direct consequence of using the Stable Portfolio and demonstrates that the model is not memorizing climatic anomalies but learning persistent phenological relationships.
4.5. Solving the Temporal Generalization Gap
The cross-year validation results (
Table 17) provide the strongest evidence for the operational value of our framework. The Stable Portfolio achieves 84.3% cross-year accuracy compared to 76.8% for the full feature set, representing a 7.5% absolute improvement. More importantly, the generalization gap narrows from 18.4% to 10.5%, a 43% relative reduction. This directly addresses the “temporal generalization gap” that has long been recognized as a fundamental obstacle in operational remote sensing [
64,
65].
Models trained on the full feature set capture year-specific patterns that fail to transfer, as reflected in their high same-year accuracy (95.2%) but poor cross-year performance. In contrast, the Stable Portfolio, by focusing on features with proven multi-year stability, captures underlying phenological signals that persist across climatic variations [
66]. The worst-case accuracy improvement from 68.2% to 81.7% is particularly significant for operational systems, where reliability under challenging conditions often matters more than peak performance. This 13.5% gain provides a critical safety margin against anomalous years, reducing the risk of catastrophic failure when deploying automated systems in new environments [
67].
Statistical significance testing (
Table 17) confirms that these improvements are not due to chance. The Stable Portfolio significantly outperforms the Full Feature Set (
p < 0.01) and the Volatile Set (
p < 0.001), while the non-significant differences with the Top-RI and Top-AuSS sets validate that these metrics are effective selection criteria. The marginal significance of the Stable Set (
p < 0.05) suggests that pure stability without importance weighting is insufficient, a key insight that justifies the composite nature of RI and AuSS.
4.6. A Pathway to Less-Intervention Operational Mapping
This study provides a complete, replicable pathway from raw satellite data to a fully automated, climate-resilient classification system. The framework begins with multi-year consensus importance derived from a Voting Ensemble, which filters out algorithm-specific noise and reveals stable phenological signals [
68]. Stability metrics (CV, Stability Index) quantify inter-annual consistency, while the novel composite indices RI and AuSS balance importance with stability, providing actionable selection criteria [
69,
70]. Pareto optimization identifies the minimal feature set (6 indices, 57.2% of cumulative importance), and the final Voting Ensemble trained on this Stable Portfolio achieves high accuracy with a minimal generalization gap.
The resulting system significantly reduces the manual intervention for feature selection or model retraining across years [
71]. The Stable Portfolio was identified once using historical data (2018, 2019, 2020, 2023, 2024 and 2025) and transfers effectively to new years without modification. This represents a fundamental shift from the current practice of annual model retraining and feature re-optimization [
72,
73], and it is precisely the kind of approach needed to scale satellite-based crop monitoring to operational levels. The potential cost savings in reduced ground data requirements and computational overhead could be substantial, and the increased reliability makes the system suitable for integration into national monitoring programs [
74].
Our framework was developed and validated on five irrigation zones with a Mediterranean winter cropping system (sowing September–November). Morocco also has Atlantic zones (cooler, wetter, earlier sowing) and pre-Saharan zones (arid, summer cropping). While we did not test these zones, the framework can be adapted by: (1) shifting monthly composites to match local phenology or (2) retraining stability metrics with 2–3 years of local data. Future work will validate this transferability.
Moreover, the framework is transferable. While the specific Stable Portfolio was derived for Morocco’s cereal-dominated landscapes, the methodology, multi-year ensemble consensus, stability metrics, and Pareto optimization can be applied to any region and any crop type. The demonstrated patterns of sensor complementarity (radar stable year-round, optical peak in autumn) are likely to hold in other Mediterranean and semi-arid environments, though regional calibration would be beneficial.
4.7. Limitations and Future Research Directions
While this study provides comprehensive insights into feature stability, several limitations should be acknowledged. Although the six-year analysis period is substantial, it represents a specific climatic window in Morocco; the gap between 2020 and 2023 means that certain inter-annual variability patterns may not be fully captured. As shown by Forkel et al. [
75], even gaps of 2–3 years can affect breakpoint detection performance, particularly for gradual trend changes, suggesting our stability estimates may be conservative. Extending this analysis to longer time series and different agro-ecological zones would validate the generalizability of the Stable Portfolio and the stability metrics [
76].
Second, the focus on cereal crops in Morocco’s Mediterranean climate means the specific feature rankings may not transfer directly to other cropping systems or climatic regimes. However, the methodological framework, Voting Ensemble consensus, stability metrics, and Pareto optimization, is designed to be transferable, and similar patterns of sensor complementarity likely apply across agricultural systems [
77,
78]. Future work should test the framework in regions with different crop types (e.g., root crops, orchards) and climatic conditions (e.g., humid tropics, temperate zones).
Third, we used only raw radar backscatter (VV, VH) and derived indices (VH/VV, RVI). Future work should investigate whether other SAR metrics (e.g., polarimetric decompositions, coherence) offer different stability-importance trade-offs [
78]. Similarly, the inclusion of additional optical indices (e.g., those sensitive to canopy water content or lignin) could further refine the portfolio.
Fourth, the Voting Ensemble, while robust, requires more computational resources than single classifiers. For very large-scale operational deployment, investigating lighter models (e.g., distilled versions of the ensemble, or simpler models like logistic regression with the Stable Portfolio) that maintain the generalization would be valuable.
Future research directions include:
Integration of meteorological data to understand the environmental drivers of feature volatility and develop conditional models that adapt to predicted climate anomalies. For example, in years forecasted to be dry, the model could automatically increase the weight of water stress indices.
Temporal transfer learning approaches that fine-tune the Stable Portfolio with minimal new data each year rather than retraining from scratch. This could further reduce operational costs. Also, Iterative decomposition methods such as BFAST could inform this process by identifying when seasonal or trend components shift significantly [
79].
Extension to other crop types and regions to validate the universality of the importance-stability trade-off and the effectiveness of the AuSS metric.
Deep learning architectures that can automatically learn stable feature representations, potentially identifying patterns not captured by hand-crafted indices. Such architectures could directly output stability-weighted features, bypassing the manual index engineering step.
Operational deployment studies that test the Stable Portfolio-based Voting Ensemble in real-time mapping scenarios, measuring not just accuracy but also computational efficiency and reliability under operational constraints (e.g., cloud cover, data latency).
5. Conclusions
This study demonstrates that feature importance in crop classification is fundamentally dynamic, and that stability-aware feature selection is not merely diagnostic but essential for operational automation. Through a comprehensive analysis of 156 multi-sensor features across six agricultural years in Morocco, using Voting Ensemble consensus to derive robust importance estimates, we have established several key findings with both scientific and practical significance. While the framework demonstrates strong performance across five irrigation zones, full national operational capability would require validation in additional agro-ecological zones (e.g., Atlantic and pre-Saharan regions) and sensitivity analysis of the automation threshold across these zones. Nevertheless, the methodology is transferable, and the Stable Portfolio provides a robust foundation for scaling to national monitoring with modest local calibration
First, we quantified the fundamental importance-stability trade-off that governs feature utility in multi-annual crop mapping. The four-quadrant behavioral typology comprising Dominant Stable, High-Performing Volatile, Stable Minor, and Noise provides a strategic framework for understanding this trade-off and making informed feature selection decisions. No feature excels in both criteria, confirming that operational systems must balance rather than optimize.
Second, we demonstrated the complementary roles of Sentinel-1 radar and Sentinel-2 optical data across the phenological cycle. Radar provides exceptional stability throughout the year, with peaks in early growth, mid-season stress detection, and harvest, serving as the structural backbone of automated systems. Optical vegetation indices dominate during the post-harvest period (September–November), providing irreplaceable discriminative power despite their higher volatility. This complementarity enables year-round automation through adaptive, phenology-aware sensor weighting.
Third, we introduced novel composite indices—the Reliability Index (RI) and Automatic Selection Score (AuSS)—that explicitly balance predictive importance with temporal stability. These metrics provide actionable criteria for feature selection, identifying the pillars of automation that perform consistently across diverse climatic conditions.
Fourth, we applied Pareto optimization to identify a Stable Portfolio of six indices—VH, VV, NDVI, NDRE, GCVI, and RVI—which captures 57.2% of cumulative predictive importance while filtering out inter-annual noise. This 54% reduction in the number of indices demonstrates that conventional full-feature approaches carry substantial redundancy and overfitting risk.
Fifth and most significantly, we validated that the Voting Ensemble trained on this Stable Portfolio achieves operational readiness. With 87.4% test accuracy, 87.2% macro F1-score, and a generalization gap of only 10.5% (compared to 18.4% for the full feature set), this model is capable of year-to-year transfer with minimal intervention. The 13.5% improvement in worst-case accuracy provides a critical safety margin against anomalous years and directly helps address the temporal generalization gap that has long hindered operational remote sensing.
This work has several practical implications, operational systems can adopt the Stable Portfolio of six indices as a foundation, eliminating the need for annual feature re-optimization. They can implement the Voting Ensemble combining Random Forest, Extra Trees, XGBoost, and LightGBM as the classifier, leveraging multi-model consensus for robust predictions. They can use the Automatic Selection Score threshold of 527 as a guideline for incorporating new features, ensuring that additions meet stability criteria. They can apply the cross-year validation protocol as a standard practice, recognizing that same-year accuracy overestimates operational performance. And they can leverage sensor complementarity by trusting radar during transition phases and optical indices during their peak windows, adapting weights phenologically.
From a broader perspective, this work challenges the conventional paradigm of static feature selection in remote sensing. The features that appear most important in single-year studies are often the most volatile, optimizing for historical performance rather than future robustness. By shifting focus from importance alone to the importance-stability trade-off, and by providing the quantitative tools of the Reliability Index and Automatic Selection Score to manage this trade-off, we enable a new generation of automated classification systems designed for resilience rather than peak performance.
The framework developed here, encompassing Voting Ensemble consensus importance, stability metrics, Pareto-optimized portfolio, and final ensemble validation is transferable to other regions, crop types, and sensor combinations. As climate variability increases and the demand for timely agricultural information grows, such stability-aware approaches will become essential for building monitoring systems that remain reliable in the face of environmental change. Future work should extend this analysis to longer time series, integrate meteorological drivers of volatility, and explore deep learning architectures that can automatically learn stable feature representations. Fully automated, climate-resilient crop mapping at scale is now an achievable goal, guided by the principles and tools established in this study.