Next Article in Journal
Phytotoxicity Response of Tomato to Simulated Clomazone Residues as Affected by Soil Type and Soil Moisture
Previous Article in Journal
Organosilicon-Mediated Cadmium Uptake, Transport, and Detoxification in Rice: Comparison Between Low-Cd and Conventional Rice Varieties
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Spatially Varying Climate Importance in United States Agricultural Insurance Claim Severity

by
Erich Seamon
School of Earth and Environmental Sciences, Baylor University, Waco, TX 76798, USA
Agriculture 2026, 16(18), 1978; https://doi.org/10.3390/agriculture16181978
Submission received: 5 August 2026 / Revised: 8 September 2026 / Accepted: 11 September 2026 / Published: 16 September 2026
(This article belongs to the Special Issue Agricultural Applications of Meteorology and Climate Forecasting)

Abstract

Climate relationships with agricultural insurance claims are unlikely to be spatially uniform because crops, hazards, and insurance experience vary geographically. To identify geographic variation in climate predictors of claim severity, we fit six geographically weighted random forest models to monthly county observations across the conterminous United States from 1994 through 2022, pairing corn, soybeans, or wheat with drought or excess moisture/precipitation/rain. Model-specific coverage ranged from 1722 to 1934 counties, with the outcome variable being logarithmically transformed positive loss per claim, adjusted to constant 2022 dollars. Models included monthly precipitation, maximum temperature, minimum temperature, vapor-pressure deficit, actual evapotranspiration, and soil moisture, together with seasonal controls and year; each local forest used an adaptive neighborhood of the 10 nearest eligible counties. Total positive climate importance was greatest in the Great Plains and Midwest but varied by commodity and cause. Maximum temperature ranked first in five of the six models, while actual evapotranspiration ranked first for soybean excess moisture claims. Corn and soybean drought models emphasized maximum temperature and vapor pressure deficit, while excess moisture models were more heterogeneous. These findings demonstrate that a single national profile of climate importance cannot adequately represent local relationships between climate and claim severity.

1. Introduction

Agricultural production is often exposed simultaneously to chronic climatic conditions and acute weather extremes [1]. High temperatures can reduce yields nonlinearly once crop-specific thresholds are exceeded, while water deficits, atmospheric demand, and excess precipitation can impair crop growth through distinct physiological and field-management pathways [2,3,4,5]. At broader scales, anthropogenic climate change has generally slowed agricultural productivity growth, increasing concern about how changing hazards will alter production risk and the institutions designed to manage it [6,7].
The U.S. Federal Crop Insurance Program is a central component of agricultural risk management [8]. Because indemnity records identify commodities, locations, timing, and reported causes of loss, insurance data also provide an extensive archive for examining weather-related agricultural damages [9,10]. Previous work has linked warming to increased yield risk and program costs and has shown that insurance incentives may influence adaptation to extreme heat [9,11,12,13,14]. Cause-of-loss records further show that drought and excess moisture are among the most consequential weather-related categories in the program [10,15]. Corn, soybeans, and wheat are major U.S. field crops with documented sensitivities to high temperature and water-related stress [2,3,4,5], providing a useful basis for evaluating whether relationships between climate and claim severity exhibit consistent spatial structure across cropping systems. These studies establish that climate matters for insured agricultural losses, but most empirical analyses estimate average relationships over large geographic domains or impose a common functional form across locations.
A spatially uniform relationship is unlikely to describe the full U.S. agricultural system. Crop calendars, cultivars, soils, irrigation, management, hazard climatology, and insurance participation vary geographically: a climate variable that is informative for claim severity in one production region may add little information elsewhere, and the locally important predictors of drought claims may differ from those associated with excess-moisture claims. National average coefficients or global machine-learning importance scores can obscure this fine-grained nonstationarity.
Geographically weighted methods address spatial nonstationarity by estimating models within locally defined subsets of the data [16]. Geographically weighted regression allows relationships between predictors and an outcome to vary across space, but it generally requires a prespecified functional form and may be difficult to interpret when relationships are nonlinear or strongly interactive [17]. Random forests provide a complementary approach by representing nonlinear relationships and interactions through ensembles of regression trees without requiring a predetermined response surface [18]. Geographically weighted random forests combine spatial localization with the flexibility of random forests by fitting separate ensembles within neighborhoods surrounding focal locations [19]. The resulting local models can be used to map geographic variation in predictor importance and model behavior. Although this approach does not provide conventional hypothesis tests or causal estimates, it is well suited to a national county panel in which the climate variables associated with agricultural insurance claim severity are expected to vary geographically. Recent applications of geographically weighted random forests have similarly used local models to characterize spatial heterogeneity in predictor importance, including large-area population-health analyses [20]; the present study extends this framework to agricultural insurance data and climate claim severity relationships.
This study asks: (1) where is the combined local usefulness of climate predictors greatest for explaining positive agricultural insurance claim severity; (2) which climate predictor is locally dominant; and, (3) how do these patterns differ among corn, soybeans, and wheat and between drought and excess-moisture damage causes? To address these questions, we analyze county–month observations from 1994–2022 using six county-centered geographically weighted random forest models, emphasizing spatial heterogeneity in conditional claim severity by evaluating the size of positive payments per claim among county–months in which at least one positive insurance payment occurred. This approach does not estimate claim probability, total expected loss, or exposure-normalized insurance risk.

2. Materials and Methods

2.1. Study Design and Analytical Units

The study area was the conterminous United States, with insurance observations organized as a county–month–year panel for 1994–2022. Separate models were fitted for each of the six combinations formed by crossing three commodities (corn, soybeans, and wheat) with two damage causes (drought and excess moisture/precipitation/rain). Insurance records were derived from publicly available U.S. Department of Agriculture Risk Management Agency (RMA) cause-of-loss and Summary of Business resources [21,22]. Insurance payments were adjusted for inflation using the monthly U.S. city average Consumer Price Index for All Urban Consumers (CPI-U), not seasonally adjusted, and expressed in constant 2022 dollars.
The analysis was limited to county–months with at least one positive insurance payment for the specified commodity and damage cause. Observations with non-positive losses per claim were excluded because the analysis was designed to model the severity of positive insurance payments. The results therefore characterize spatial variation in conditional claim severity only for county–months in which a positive payment occurred and do not estimate claim occurrence, unconditional expected loss, or exposure-normalized agricultural insurance risk. The number of counties included in each model varies because commodity production, reported damage causes, and positive insurance payments are not distributed uniformly across the United States. The overall analytical workflow, from insurance and climate data preparation through county-centered geographically weighted random forest modeling, is summarized in Figure 1.

2.2. Claim-Severity Outcome and Inflation Adjustment

The primary response was positive real insurance loss per claim, expressed in constant 2022 U.S. dollars. Nominal losses were adjusted using the monthly Consumer Price Index for All Urban Consumers, U.S. city average, all items, not seasonally adjusted (Bureau of Labor Statistics series CUUR0000SA0) [23]. The modeled response was
Y c m t = log 1 + real _ loss _ per _ claim _ 2022 c m t ,
where c, m, and t index county, month, and year. The logarithmic transformation reduced the leverage of the strongly right-skewed loss distribution while retaining small positive values. In the Supplemental Materials and associated datasets, this transformed response is denoted as log_real_loss_per_claim; this variable is exactly the log ( 1 + x ) transformation defined above.
Dividing loss by claim count adjusts the outcome for the number of claims represented in a county–month, but does not normalize by insured acreage, liability, premium, policy count, coverage level, or other measures of insurance exposure. The outcome therefore represents payment severity per claim, not a loss rate or total agricultural vulnerability. All references to claim severity in this study refer to conditional positive payment severity among county–months in which at least one indemnity was recorded, rather than unconditional agricultural loss or exposure-normalized insurance risk.

2.3. Climate Predictors and Temporal Controls

Each model included six monthly climate predictors: precipitation (ppt), maximum temperature (tmax), minimum temperature (tmin), vapor-pressure deficit (vpd), actual evapotranspiration (aet), and soil moisture (soil). The soil variable represents monthly soil-moisture conditions from TerraClimate and should not be interpreted as a soil taxonomic class. Because the study spans the conterminous United States and multiple crop-production regions, categorical soil properties were not included as predictors in the present analysis. Together, these variables represent water supply, thermal conditions, atmospheric moisture demand, realized land-surface water use, and soil-water availability. Monthly climate and climatic water-balance data were obtained from TerraClimate through the University of Idaho’s Northwest Knowledge Network THREDDS data service [24]. The gridded data were spatially subset to the conterminous United States and aggregated to county polygons using coverage-fraction-weighted means across raster cells intersecting each county. This procedure produced one observation for each county, year, month, and climate variable.
The RMA temporal field used in the analysis is month of loss, which identifies the month in which the insured loss occurred, rather than the month in which a claim was filed or an indemnity was paid [21]. County-month climate values for 1994–2022 were joined to the insurance records by county FIPS code, year of loss, and month of loss. Monthly climate variables were used to align environmental conditions directly with the reported month of loss and to provide a consistent temporal specification across commodities, damage causes, and production regions. We did not impose lagged, cumulative, or growing-season windows because the relevant antecedent period may differ by commodity, crop phenology, production region, and damage mechanism. The resulting models should therefore be interpreted as evaluating associations between climate conditions during the reported loss month and conditional claim severity; potential effects of antecedent or cumulative climate conditions are not represented explicitly.
Because calendar year had high permutation importance in the primary models, robustness was evaluated by refitting all six final GWRF models without the year term, while retaining the six climate predictors and the seasonal sine and cosine controls. All other model settings were held constant. The resulting climate importance surfaces and model-wide predictor rankings were compared with the primary year-controlled models (see Supplementary Materials, Tables S34–S36).
Temporal stability was also evaluated by refitting the six final GWRF models separately for 1994–2008 and 2009–2022 using the same 500-tree, k = 10 specification, including the year and seasonal controls. County-level total positive climate importance, dominant-predictor agreement, and model-wide predictor rankings were compared for counties represented in both periods. Comparisons were repeated after restricting to counties with at least 50 and 100 local observations in both periods (Supplementary Materials, Tables S46–S48 and Figure S67).
To account for recurring seasonal structure, calendar month was represented using sine and cosine terms:
month _ sin = sin 2 π m 12 , month _ cos = cos 2 π m 12 ,
where m denotes calendar month. Calendar year was included as a continuous temporal control to absorb broad long-term changes in program structure, technology, management, reporting, and climate that were not otherwise represented by the monthly predictors.

2.4. County-Centered Geographically Weighted Random Forest

Random forests estimate an ensemble of regression trees from resampled data and randomly selected predictor subsets, allowing flexible representation of nonlinear and interacting relationships [18]. Similar geographically localized random-forest approaches have been used to evaluate spatial heterogeneity in predictive relationships across large geographic domains [20]. In the county-centered geographically weighted random forest design, one local random forest was centered on every eligible focal county. The focal county was represented by a reproducible anchor row used only to supply coordinates. The full monthly panel for the focal county and neighboring counties remained available to fit the local model.
Neighborhoods were defined using a location-based adaptive bandwidth. For each focal county, the primary model selected the 10 nearest unique eligible county locations, including the focal county. The bandwidth therefore refers to counties rather than panel rows. An adaptive bandwidth of k = 10 unique county locations was selected to preserve strong geographic localization while retaining sufficient observations for local model fitting. A broader k = 25 specification was evaluated as a sensitivity analysis to assess the effect of increased spatial smoothing. This design avoided zero-distance and duplicated-location problems that arise when repeated monthly observations are treated as independent spatial locations. Within each adaptive neighborhood, observations were weighted by distance from the focal county using a bi-square spatial kernel. These weights were supplied to the local ranger model as case weights, so observations from counties nearer the focal county contributed more strongly to fitting the local forest. Because spatial distance was defined at the county-location level, all monthly observations from the same county received the same spatial weight within a given focal model.
For observation i relative to focal county c, the spatial weight was
w i c = 1 d i c b c 2 2 , d i c < b c , 0 , d i c b c .
where d i c is the distance between the county associated with observation i and focal county c, and b c is the adaptive local bandwidth. The kernel assigned its greatest weight to observations from the focal county and decreased to zero at the local bandwidth.
Because each selected county contributed all eligible monthly observations, local training sample sizes varied with historical support in the focal neighborhood. To identify local fits with limited training support, we used 50 local observations as a pragmatic low-support threshold. The primary analysis retained all eligible focal counties, but counties with fewer than 50 local observations were explicitly flagged (see Supplementary Materials, Tables S5–S9). Sensitivity was evaluated by recalculating aggregate climate importance summaries and predictor rankings after restricting results to counties with at least 50 and at least 100 local observations. These thresholds were used as diagnostics of result stability rather than as criteria for refitting the primary GWRF models.
The final county-centered random forests used 500 trees, mtry = 3, and a minimum node size of 5. Variable importance was calculated using permutation importance, and the bi-square spatial kernel weights described above were supplied to ranger as case weights. The final model specification was:
Y ppt + tmax + tmin + vpd + aet + soil + month _ sin + month _ cos + year .
Tree-count stability was evaluated by comparing results from 50, 250, and 500-tree ensembles. Comparisons included county-level total positive climate importance, climate-predictor rankings, and dominant-predictor agreement, including stratification by local training support. The 500-tree ensemble was retained as the final production specification (see Supplementary Materials, Tables S28 and S29). The principal output was county-specific permutation importance from each county-centered local random forest. Importance was calculated using the unscaled permutation measure implemented in ranger [25]. For each predictor, values were permuted among out-of-bag observations, and importance was quantified as the average increase in out-of-bag squared prediction error relative to the unpermuted predictions. Larger positive values indicate greater predictive reliance on a variable within the fitted local model, while values near zero indicate little contribution to out-of-bag predictive performance. Negative values can occur when permutation does not degrade, or slightly improves, predictive performance, and should not be interpreted as evidence of a protective effect. Permutation importance measures predictive reliance rather than effect direction; it thus does not indicate whether higher predictor values increase or decrease claim severity and should not be interpreted causally.
A second run using an adaptive bandwidth of 25 unique county locations was generated as a sensitivity analysis. While the k = 10 run is the primary analysis because it preserves greater geographic localization, the k = 25 results were used to compare with the primary models using county-level correlations in total positive climate importance, dominant predictor agreement, model-wide predictor-rank correlations, and adjacent-county roughness.

2.5. Climate Importance Metrics and Map Construction

One direct county-level model output and two derived county-level summaries were used. The direct output was the signed permutation importance, I c j , for climate predictor j at focal county c. Variable-specific maps display these values, including positive, near-zero, and negative importance values.
The first derived summary was total positive climate importance for focal county c:
C c = j K max ( I c j , 0 ) ,
where K contains the six climate predictors. This index summarizes their combined positive importance without allowing negative values to offset positive values.
The second derived summary was the dominant local climate predictor,
D c = arg max j K I c j ,
provided that at least one climate predictor had positive importance. Counties with no positive climate importance were assigned a separate category. Dominance identifies the predictor with the largest positive importance within a county; it does not imply that its importance differs significantly from that of the second ranked predictor. To assess uncertainty in dominant-predictor classification, we calculated the relative importance margin between the two highest-ranked positive climate predictors and mapped locations with small rank margins (Supplementary Materials, Tables S44 and S45 and Figures S65 and S66). Model-wide means and medians were calculated across eligible focal counties. Because importance scales depend on the fitted response and local data structure, comparisons of raw importance magnitudes across separately fitted models are purely descriptive. Within-model rankings and mapped spatial variation therefore provide the primary evidence. Predictor correlation was evaluated within the final local GWRF neighborhoods, and a grouped held-out permutation analysis was used to assess the robustness of broader climate-process signals to correlated predictors (see Supplementary Materials, Tables S39–S43 and Figures S63 and S64).
Panel construction, spatial processing, GWRF estimation, summary statistics, and figure production were implemented in R 4.5.2. Local random forests were fitted using version 0.1.1 of the gwrf package [26], which was developed by the author to support geographically weighted random forests with adaptive neighborhoods defined by unique spatial locations, and is publicly available through the Comprehensive R Archive Network (CRAN). The package uses ranger (version 0.18.1) as the random forest engine [25]. This implementation allowed repeated county-month observations to be included without treating repeated rows from the same county as distinct neighboring locations. Core data management, spatial processing, and visualization packages included arrow, dplyr, sf, terra, and ggplot2. Analysis code, package-version information, and supporting metadata are provided in the Data Availability Statement.

2.6. Model Performance and Validation

Predictive performance was evaluated using 10-fold cross-validation grouped by county. All monthly observations from a held-out county were excluded from model training, preventing repeated observations from the same spatial location from appearing in both the training and test sets. The same county-fold assignments were used to compare the geographically weighted random forest with an otherwise equivalent global random forest. Both approaches used the same response and predictors and were fitted with 500 trees, mtry = 3, and a minimum node size of 5. For GWRF, predictions for each held-out county were generated from a spatially weighted local model based on the 10 nearest eligible training counties using the same bi-square kernel as the primary analysis. Predictive performance was summarized using root mean squared error (RMSE), mean absolute error (MAE), and out-of-sample R 2 .
Local out-of-bag (OOB) performance was also obtained from each final 500-tree county-centered forest, with local OOB RMSE and R 2 characterizing predictive performance within the fitted local neighborhood. Because monthly observations from the same neighboring counties remain within these local training datasets, OOB diagnostics were interpreted as measures of internal local-model fit rather than as substitutes for the county-grouped cross-validation. Complete validation results are provided in the Supplementary Materials (Tables S30–S33 and Figure S62).

3. Results

3.1. Panel Coverage and Local Training Support

The six models included 74,527–129,966 eligible county–month–year observations accumulated over the full 1994–2022 study period, representing 1722–1934 unique focal counties (Table 1). Each observation corresponded to one county in one month and year for a specified commodity and damage cause, conditional on a positive insurance payment. Corn and soybean observations were concentrated across the central and eastern United States, whereas wheat had a broader western and northern footprint. Mean local training support, defined as the number of eligible panel observations available to an individual county-centered forest from its adaptive county neighborhood, ranged from 426.4 to 759.1 observations across all models. Although average local support was substantially larger, some peripheral focal counties (particularly in the wheat–drought model) had limited training data, with the minimum local neighborhood containing 14 panel observations. These differences reflect commodity production geography and the frequency of eligible positive-payment observations.
Low-support model fits were uncommon in five of the six models, with under 2% of focal counties falling below the 50-observation threshold in the corn and soybean models and in the wheat excess-moisture model. Wheat drought was the exception, with 15.6% of focal counties below the 50-observation threshold. Additionally, sensitivity analyses using minimum-support thresholds of 50 and 100 observations showed that the model-wide highest-ranked climate predictor was unchanged at the 50-observation threshold in five of six models. In wheat drought, minimum temperature replaced maximum temperature as the highest-ranked mean predictor. The full support distributions and threshold-sensitivity results are provided in the Supplementary Materials (Tables S5–S9).
Figure 2 summarizes the analytical workflow. For each focal county, the adaptive neighborhood was defined using unique county locations, while all eligible monthly observations from the selected counties were retained for local model fitting. The focal anchor observation supplied the coordinates used to center the local model and was not interpreted as a county-average prediction or residual.
Tree-count sensitivity indicated strong stability of the principal climate importance patterns. Comparing the original 50-tree results with the final 500-tree models, county-level total positive climate importance remained highly correlated across all six models (Pearson r = 0.970–0.988; Spearman ρ = 0.961–0.988), and the model-wide highest-ranked climate predictor was unchanged in all six models. Convergence was stronger between 250 and 500 trees, with median Pearson correlations of 0.974–0.994 across local-support classes and dominant-predictor agreement of 81.1–88.3%. These results supported the use of 500 trees for the final analysis (Supplementary Materials, Tables S28 and S29).

3.2. Model-Wide Climate Importance Rankings

Permutation importance values represent unscaled increases in out-of-bag squared prediction error and are therefore not restricted to values between zero and one; larger positive values indicate greater predictive reliance within a fitted model. When averaged across focal counties, maximum temperature had the highest mean permutation importance in five of the six models (Table 2). Its prominence was greatest in the corn and soybean drought models, with mean importance values of 0.767 and 0.788, respectively. In these two models, vapor pressure deficit ranked second (0.711 for corn and 0.617 for soybeans), followed by minimum temperature. The only exception to the overall maximum-temperature ranking was soybean excess moisture, with actual evapotranspiration ranked first (0.517), followed by maximum and minimum temperature. Corn excess moisture likewise showed relatively little separation among maximum temperature, actual evapotranspiration, and minimum temperature. The wheat models had lower mean climate importance values overall and less differentiation among the leading thermal and atmospheric-demand predictors. Strong local correlations were evident among maximum temperature, minimum temperature, and vapor pressure deficit, while grouped held-out permutation analyses showed consistent predictive contributions from broader thermal-demand and moisture/water-balance predictor groups across the six models (Supplementary Materials, Tables S39–S43).
Year had the highest mean permutation importance among the temporal controls in all six models. Mean year importance ranged from 0.777 for wheat excess moisture to 2.282 for soybean excess moisture, with particularly high values in the corn- and soybean-excess-moisture models. This pattern indicates that broad temporal structure contributed substantially to local prediction. Because the year term may capture a mixture of technological change, policy and reporting shifts, changes in exposure and coverage, and long-term climatic variation, it was treated as a control variable rather than as a climate predictor.
A sensitivity analysis removing the calendar-year control showed that the principal spatial climate importance patterns were retained. County-level total positive climate importance remained strongly correlated with the primary models across all six commodity–damage-cause combinations (Pearson r = 0.918–0.958), and model-wide climate-predictor rankings remained similar ( ρ = 0.829–1.000). Removing year increased the magnitude of climate importance but changed the highest-ranked climate predictor in only two of six models (see Supplementary Materials, Tables S34–S36).

3.3. Total Positive Climate Importance

Total positive climate importance (the sum of positive local permutation-importance values across the six climate predictors) varied substantially among counties in every model (Figure 3). In the corn and soybean models, counties with the highest values of this measure were concentrated primarily across the central and northern Great Plains and the western Corn Belt, especially in parts of the Dakotas, Nebraska, Kansas, Iowa, and Minnesota. This pattern was most spatially continuous in the corn-drought and soybean-drought models. The corresponding excess-moisture models retained a central Great Plains and Midwest concentration but displayed more localized hotspots and less continuous expansion into eastern and southern portions of their model regions.
Wheat displayed a different geographic pattern. In the drought model, elevated total positive climate importance extended farther west through portions of the northern Plains, southern High Plains, and interior West, while values were generally lower across many central-eastern counties. The wheat excess-moisture model showed elevated values across parts of the northern tier and central-to-southern Plains, together with scattered hotspots elsewhere. Collectively, the six maps demonstrate that total positive climate importance varied spatially rather than forming a geographically uniform pattern.

3.4. Dominant Local Climate Predictors

The identity of the climate predictor with the greatest positive local permutation importance varied across space, although several coherent regional patterns were evident (Figure 4). Maximum temperature was the dominant predictor across broad portions of the corn- and soybean-drought domains, particularly from the central Great Plains through much of the Midwest. Vapor pressure deficit and other climate variables were dominant along portions of the western and northern margins of these domains and in smaller local clusters. In the corn and soybean-excess-moisture models, actual evapotranspiration was dominant across a broad Midwestern corridor extending approximately from Nebraska through Iowa, Illinois, Indiana, and Ohio. Outside this corridor, the excess-moisture maps showed greater local variation among maximum temperature, minimum temperature, precipitation, soil moisture, vapor pressure deficit, and actual evapotranspiration.
The wheat drought model displayed a different and more fragmented geographic pattern. Vapor pressure deficit was dominant across much of the northern portion of the wheat domain, particularly through parts of the northern Great Plains, whereas minimum temperature was prominent in the southern High Plains, including the Texas Panhandle and adjacent areas. Elsewhere, the dominant predictor changed frequently among neighboring counties, and some counties had no positive permutation importance for any of the six climate predictors. The wheat excess-moisture model was the most geographically heterogeneous, with frequent local changes in the highest-ranked predictor and no single climate variable dominating across a broad region.
These categorical maps represent the predictor with the highest positive permutation importance within each county. They should therefore be interpreted as local rankings rather than evidence that the mapped predictor controlled claim severity or was the sole driver of losses. Where two or more predictors had similar importance values, relatively small differences could determine the mapped category. 27.5–45.2% of counties had a relative importance margin of 10% or less between the two highest-ranked predictors across the six models (Supplementary Materials, Tables S44 and S45 and Figures S65 and S66).

3.5. Spatial Variation in Maximum Temperature Importance

Because maximum temperature had the highest mean permutation importance in five of the six models, its county-level importance was examined in greater detail as a representative variable-specific result, with Figure 5 illustrating why model-wide rankings alone are incomplete. In the corn and soybean-drought models, maximum temperature showed broadly positive importance across the Great Plains, Midwest, and parts of the eastern production domains, with localized areas of particularly high importance embedded within these regions. The corresponding excess-moisture models also showed positive maximum-temperature importance across much of the central agricultural region, but the spatial pattern was weaker and more variable, with scattered near-zero and negative values.
In the wheat drought model, positive maximum-temperature importance extended across portions of the northern Plains, southern High Plains, and interior West. The wheat excess-moisture model showed a weaker and more discontinuous spatial pattern, although localized areas of elevated importance remained. Positive values indicate that permuting maximum-temperature values increased out-of-bag prediction error and therefore reduced local predictive performance. They do not demonstrate that higher temperatures caused larger insurance payments. Determining effect direction and response shape would require a separate model-interpretation analysis beyond the scope of this study.

3.6. Bandwidth Sensitivity

The broader adaptive neighborhood of k = 25 unique county locations produced results that were generally consistent with the primary k = 10 analysis, although the exact identity of the dominant local predictor was more sensitive to bandwidth choice (Table 3). Across the six models representing combinations of commodity and damage cause, Pearson correlations in county-level total positive climate importance between the two bandwidths ranged from 0.813 to 0.902, and Spearman rank correlations ranged from 0.771 to 0.904. Model-wide correlations in the rankings of the six climate predictors ranged from 0.943 to 1.000, and the top-ranked predictor was unchanged in five of the six models.
The only change in the top-ranked model-wide predictor occurred in the wheat drought model, for which maximum and minimum temperature were nearly indistinguishable under the primary bandwidth. Maximum temperature ranked first under k = 10 , with a mean permutation importance of 0.371, compared with 0.369 for minimum temperature. Under k = 25 , minimum temperature ranked slightly above maximum temperature. This change therefore reflected a near-tie rather than a substantial reordering of predictor importance.
County-level agreement in the dominant climate predictor ranged from 39.6% to 65.6%, indicating that the precise categorical assignment of the highest-ranked predictor was more bandwidth sensitive than total climate importance or model-wide predictor rankings. The broader k = 25 neighborhoods substantially increased local training support and reduced standardized adjacent-county roughness by 5.5% to 19.3% across the six models. Thus, the broader bandwidth produced smoother importance surfaces and altered some local predictor classifications, but it did not materially change the broader conclusions concerning geographic variation in climate importance and model-wide predictor rankings.

3.7. Temporal Sub-Period Sensitivity

Refitting the six final GWRF models separately for 1994–2008 and 2009–2022 showed that broad climate-importance structure persisted across periods, although local patterns were not temporally fixed (Figure 6). Among counties represented in both periods, county-level total positive climate importance showed moderate to strong correspondence, with Pearson correlations of 0.508–0.741 and Spearman correlations of 0.499–0.745 across the six models. Model-wide climate-predictor rankings were more stable, with Spearman correlations of 0.829–0.943, and the highest-ranked predictor was unchanged in four of six models. Agreement in the locally dominant predictor was lower, ranging from 20.6% to 34.6%, consistent with the near-tie sensitivity documented above. The same broad pattern remained after restricting comparisons to counties with at least 50 or 100 local observations in both periods (Supplementary Materials, Tables S46–S48 and Figure S67).

3.8. Model Performance and Validation

County-grouped cross-validation included 624,570 held-out county–month observations across the six models. GWRF produced lower pooled RMSE and MAE and higher out-of-sample R 2 than the corresponding global random forest in all six commodity and damage cause combinations. Across the 60 model-by-fold comparisons, GWRF had lower RMSE in 55 folds and lower MAE in 58 folds, indicating that spatial localization generally improved predictive performance when complete county histories were held out. County-level performance was more heterogeneous, however, showing that localization did not improve prediction uniformly at every location (Supplementary Materials, Tables S30 and S31). Local OOB diagnostics provided complementary measures of performance within the fitted neighborhoods. Median local OOB R 2 ranged from 0.138 to 0.388 across models, with positive local OOB R 2 in 74.0–95.9% of focal counties. Lower OOB performance was most evident in the wheat models and in neighborhoods with limited local training support. Because neighboring county observations remain within these training datasets, the county-grouped cross-validation provides the more stringent assessment of out-of-sample performance (Supplementary Materials, Tables S32 and S33 and Figure S62).

4. Discussion

4.1. Climate-Claim Relationships Are Spatially Nonstationary

A single nationwide climate–claim severity relationship does not adequately characterize the spatial structure observed in agricultural insurance data. Within the same commodity and reported damage cause, counties differed in both the magnitude of positive climate importance and in which climate predictor ranked highest locally. These results indicate that national or model-wide summaries can mask substantial geographic heterogeneity in the climate signals associated with positive claim severity.
The bandwidth sensitivity analysis supports this interpretation (Table 3). County-level total positive climate importance remained strongly correlated between the k = 10 and k = 25 models, and model-wide predictor rankings were nearly unchanged. In contrast, agreement in the identity of the dominant local predictor was lower, and the broader k = 25 neighborhoods produced smoother spatial surfaces. Broad geographic patterns were therefore more robust than the precise county-level classification of the highest-ranked predictor. The dominant-predictor maps should be interpreted as spatially varying local rankings rather than as fixed boundaries between mechanistically distinct regions.
This result is consistent with the rationale for geographically weighted modeling. Observations assigned the same national outcome category may arise from different local combinations of hazard exposure, cropping systems, soils, water availability, management, and insurance structure. Geographical random forests make this heterogeneity explicit by fitting a separate nonlinear ensemble around each focal location [19]. In this study, the method does not simply identify where insurance payments were large. It identifies where climate variables contributed more or less to locally predicting the severity of positive payments that had already occurred.

4.2. Drought Claims Emphasize Heat and Atmospheric Demand

Maximum temperature and vapor pressure deficit were the two highest mean-ranked climate variables in both corn and soybean drought models. Their maps also showed broad areas of positive importance across the central production belt. This pattern is physically plausible because high temperature and atmospheric moisture demand can intensify crop water stress and amplify the effects of precipitation deficits [27,28,29,30]. Previous yield and insurance studies have documented nonlinear heat damages, increased yield risk under warming, and potential insurance-related disincentives to adaptation to extreme heat [2,11,13]. The GWRF results extend that literature by showing that predictive usefulness varies considerably within the national crop domain.
The analysis does not determine whether temperature or vapor pressure deficit is the proximate cause of higher payment severity, nor can correlated climate predictors be interpreted independently. Maximum temperature, minimum temperature, vapor pressure deficit, evapotranspiration, and soil moisture are physically linked. Permutation may distribute predictive credit among correlated variables differently across local forests [31,32]. The appropriate inference is that positive drought-claim severity contains a strong and geographically variable hydrothermal signal, not that any single climate variable has an isolated causal effect.

4.3. Excess Moisture Claims Reflect Multiple Hydroclimatic Pathways

The excess-moisture models produced more heterogeneous dominant-predictor mosaics than the corn and soybean drought models. Actual evapotranspiration ranked first for soybean excess moisture and nearly tied with maximum temperature for corn excess moisture. Precipitation itself was not the highest mean-ranked variable in either model. This does not imply that precipitation was unimportant to claims attributed to excess moisture. Monthly precipitation totals may incompletely represent saturation, drainage, antecedent soil conditions, evapotranspirative demand, field access, crop stage, and submonthly precipitation intensity. Excess precipitation can damage crops through delayed planting, waterlogging, erosion, disease, and reduced field operations [3,33,34,35].
The local importance of actual evapotranspiration, soil moisture, and temperature may reflect the broader water balance and crop-development context in which precipitation occurs rather than independent causal effects. Future directional analyses should examine interactions and conditional response shapes rather than interpreting importance maps as additive effects. Event-scale precipitation intensity and antecedent wetness metrics may improve the representation of excess moisture processes.

4.4. Wheat Differs from Corn and Soybeans

Wheat models had lower mean climate importance values and a different geographic footprint. Wheat production spans winter and spring systems, dryland and irrigated regions, and a wider range of western and northern climates than the primary corn and soybean domains. These differences likely contribute to the more fragmented dominant-predictor maps and counties with no positive climate importance in the wheat drought model.
Lower mean importance does not imply that wheat claim severity is insensitive to climate. It indicates that, given the monthly predictors, temporal controls, local sample sizes, and response definition, permuting individual climate variables caused a smaller average degradation in local predictive performance. Lower support in some wheat neighborhoods and greater production-system heterogeneity may also reduce stability. The local-support maps retained in the supplement are therefore important when interpreting geographically isolated patterns.

4.5. Long-Term Temporal Structure

The models retained substantial long-term temporal structure beyond that represented by the monthly climate predictors and seasonal controls. The continuous calendar-year covariate had the highest mean permutation importance of all predictors in every model, particularly in the corn- and soybean-excess-moisture models. However, calendar year is not a mechanistic driver: it may be absorbing changes in commodity prices and insurance coverage, technology, yields and management, program and reporting practices, the geographic composition of insured production, and long-term climate. Insurance subsidies may also influence crop choices and therefore the composition and geographic distribution of insured production [36].
Including calendar year reduces the likelihood that the climate predictors serve as proxies for all long-term change, but it does not identify the source of the remaining temporal variation. The no-year sensitivity analysis further showed that removing this temporal control increased the magnitude of climate importance but did not fundamentally change the broader geographic patterns. Because the covariate may represent multiple climatic and nonclimatic processes simultaneously, its importance should not be interpreted as evidence of any single temporal mechanism. The temporal sub-period sensitivity analysis showed that broad model-wide climate-predictor rankings remained similar between 1994–2008 and 2009–2022, while county-level importance magnitudes and dominant-predictor classifications varied more substantially (Figure 6; Supplementary Materials, Tables S46–S48 and Figure S67). The full-period analyses provided as part of the Supplemental Materials should therefore be interpreted as a long-term summary rather than as evidence that local climate-importance relationships were temporally fixed. More formal time-varying local models remain a subject for future work.

4.6. Implications for Agricultural Risk Analysis

The findings have three methodological implications. First, national agricultural-risk models should evaluate spatial nonstationarity explicitly rather than assume that a single climate-response structure applies across all production regions. Second, the most predictively useful climate variables may differ by commodity, reported damage cause, and location, suggesting that predictor selection and model interpretation should account for regional production and hazard contexts. Third, risk communication should distinguish among claim occurrence, conditional claim severity, and exposure-adjusted loss rates because these outcomes answer different scientific and actuarial questions.
The maps provide a geographic screening framework for identifying regions where more detailed process-based or actuarial analysis may be warranted. However, they are not premium-rating maps and should not be used directly to recommend insurance prices or program changes. Such applications would require explicit information on liability, insured acreage, coverage, and other measures of exposure, together with actuarial validation, decision-specific out-of-sample testing, and careful consideration of policy and producer behavior. Previous studies have shown that climate change can affect both yield risk and Federal Crop Insurance Program costs [6,12,13,14]. The present results suggest that these effects may be geographically differentiated in ways that pooled national models can obscure.

4.7. Limitations

Several aspects of the data and modeling approach limit how these results should be interpreted. The analysis is conditional on county–months with at least one positive insurance payment; the models therefore do not estimate whether a claim occurs, and cannot be interpreted as models of unconditional expected loss. Although dividing total payments by claim count adjusts the outcome for the number of claims, it does not control for insured acreage, liability, policy characteristics, coverage levels, price elections, or participation. Spatial differences may thus reflect both physical loss processes and the composition of the insurance system.
Reported damage cause is an administrative classification: a claim recorded as drought or excess moisture may involve interacting hazards, and classification or reporting practices may vary over time and space. Monthly county-level climate summaries may also obscure short-duration extremes, within-county environmental variation, crop stage timing, irrigation, soil properties, and antecedent conditions. In addition, local permutation importance is sensitive to predictor correlation, local training support, and model specification [31,32]. It measures predictive usefulness but does not provide effect direction, statistical significance, or causal attribution.
Some county-centered neighborhoods had relatively small local training samples, particularly in the wheat drought model, because adaptive neighborhoods guarantee a specified number of county locations rather than a fixed number of county–month observations. Finally, the bandwidth sensitivity analysis indicated that broad geographic patterns and model-wide predictor rankings were generally stable, but the exact county-level classification of the dominant climate predictor remained sensitive to neighborhood size. Temporal sub-period analyses likewise indicated that broad predictor rankings persisted between 1994–2008 and 2009–2022, but county-level importance magnitudes and dominant-predictor classifications were not temporally fixed.

5. Conclusions

This study evaluated whether the climate predictors associated with positive agricultural insurance claim severity vary geographically across the conterminous United States. County-centered geographically weighted random forests revealed substantial spatial heterogeneity across the six models representing combinations of commodity and damage cause. Maximum temperature was the highest mean-ranked climate predictor in five of the six models, while actual evapotranspiration ranked first for soybean excess-moisture claims. Maximum temperature and vapor pressure deficit were particularly prominent in the corn- and soybean-drought models, whereas the excess-moisture models exhibited more heterogeneous local predictor patterns.
The analysis also demonstrated that broad model-wide rankings do not fully represent local climate–claim relationships. Total positive climate importance and the identity of the highest-ranked climate predictor varied among counties within the same commodity and reported damage cause. The bandwidth sensitivity analysis showed that broad geographic patterns and model-wide predictor rankings were generally consistent between the k = 10 and k = 25 neighborhoods, although the dominant predictor assigned to individual counties was more sensitive to neighborhood size. Temporal sub-period analyses similarly showed that broad predictor rankings persisted between 1994–2008 and 2009–2022, while county-level importance patterns exhibited meaningful temporal variation. These results support the use of spatially localized models while emphasizing that mapped predictor boundaries should not be interpreted as fixed or mechanistically distinct regions.
The findings apply specifically to conditional payment severity among county–months with at least one positive insurance payment and do not estimate claim probability, unconditional expected loss, exposure-adjusted risk, or causal climate effects. Nevertheless, this work provides a reproducible geographic framework for identifying spatial variation in climate–claim relationships. Future research should integrate insured acreage, liability, coverage, and other measures of exposure; evaluate directional and interactive climate effects; and examine how these relationships change over time.

Supplementary Materials

The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/agriculture16181978/s1, the supporting information includes the complete county-centered GWRF analysis for all six models representing combinations of commodity and damage cause, with historical panel support, local training support, total positive climate importance, dominant climate predictor, variable-specific importance maps, and sensitivity analyses including temporal sub-period comparisons.

Funding

This research was supported by faculty startup funds provided by Baylor University’s College of Arts and Sciences and the School of Earth and Environmental Sciences.

Institutional Review Board Statement

Not applicable. This study used publicly available aggregate administrative and environmental data and did not involve human participants or animals.

Informed Consent Statement

Not applicable.

Data Availability Statement

The original agricultural insurance data used in this study are publicly available from the U.S. Department of Agriculture Risk Management Agency. Cause of Loss data are available at https://www.rma.usda.gov/tools-reports/summary-of-business/cause-loss (accessed on 10 September 2026), and Summary of Business data are available at https://www.rma.usda.gov/tools-reports/summary-of-business (accessed on 10 September 2026). Monthly climate and climatic water-balance data are publicly available from TerraClimate at https://www.climatologylab.org/terraclimate.html (accessed on 10 September 2026). Monthly Consumer Price Index data for series CUUR0000SA0 are publicly available from the U.S. Bureau of Labor Statistics at https://data.bls.gov/timeseries/CUUR0000SA0 (accessed on 10 September 2026). The processed analytical datasets, model outputs, supporting metadata, and code required to reproduce the analyses and figures are archived in Dryad (https://datadryad.org) at https://doi.org/10.5061/dryad.r4xgxd2vm. The gwrf R package (version 0.1.1), including source code and documentation, is publicly available through the Comprehensive R Archive Network (CRAN) at https://CRAN.R-project.org/package=gwrf (accessed on 10 September 2026); the development repository is available at https://github.com/hac-lab/gwrf (accessed on 10 September 2026).

Acknowledgments

The author gratefully acknowledges the institutional support provided by Baylor University’s College of Arts & Sciences, particularly through the School of Earth and Environmental Sciences.

Conflicts of Interest

The author declares no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
CPI-UConsumer Price Index for All Urban Consumers
GWRFGeographically weighted random forest
RMARisk Management Agency
USDAU.S. Department of Agriculture
THREDDSThematic Real-time Environmental Distributed Data Services

References

  1. Lesk, C.; Rowhani, P.; Ramankutty, N. Influence of extreme weather disasters on global crop production. Nature 2016, 529, 84–87. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Schlenker, W.; Roberts, M.J. Nonlinear temperature effects indicate severe damages to U.S. crop yields under climate change. Proc. Natl. Acad. Sci. USA 2009, 106, 15594–15598. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Rosenzweig, C.; Tubiello, F.N.; Goldberg, R.; Mills, E.; Bloomfield, J. Increased crop damage in the US from excess precipitation under climate change. Glob. Environ. Change 2002, 12, 197–202. [Google Scholar] [CrossRef] [Scilit]
  4. Schauberger, B.; Archontoulis, S.; Arneth, A.; Balkovic, J.; Ciais, P.; Deryng, D.; Elliott, J.; Folberth, C.; Khabarov, N.; Müller, C.; et al. Consistent negative response of US crops to high temperatures in observations and crop models. Nat. Commun. 2017, 8, 13931. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Zhao, C.; Liu, B.; Piao, S.; Wang, X.; Lobell, D.B.; Huang, Y.; Huang, M.; Yao, Y.; Bassu, S.; Ciais, P.; et al. Temperature increase reduces global yields of major crops in four independent estimates. Proc. Natl. Acad. Sci. USA 2017, 114, 9326–9331. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Lobell, D.B.; Schlenker, W.; Costa-Roberts, J. Climate trends and global crop production since 1980. Science 2011, 333, 616–620. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Ortiz-Bobea, A.; Ault, T.R.; Carrillo, C.M.; Chambers, R.G.; Lobell, D.B. Anthropogenic climate change has slowed global agricultural productivity growth. Nat. Clim. Change 2021, 11, 306–312. [Google Scholar] [CrossRef] [Scilit]
  8. Miranda, M.J.; Glauber, J.W. Systemic Risk, Reinsurance, and the Failure of Crop Insurance Markets. Am. J. Agric. Econ. 1997, 79, 206–215. [Google Scholar] [CrossRef] [Scilit]
  9. Seamon, E.; Gessler, P.E.; Abatzoglou, J.T.; Mote, P.W.; Lee, S.S. A climatic random forest model of agricultural insurance loss for the Northwest United States. Environ. Data Sci. 2022, 1, e29. [Google Scholar] [CrossRef] [Scilit]
  10. Seamon, E.; Gessler, P.E.; Abatzoglou, J.T.; Mote, P.W.; Lee, S.S. Climatic Damage Cause Variations of Agricultural Insurance Loss for the Pacific Northwest Region of the United States. Agriculture 2023, 13, 2214. [Google Scholar] [CrossRef] [Scilit]
  11. Annan, F.; Schlenker, W. Federal Crop Insurance and the Disincentive to Adapt to Extreme Heat. Am. Econ. Rev. 2015, 105, 262–266. [Google Scholar] [CrossRef] [Scilit]
  12. Tack, J.; Coble, K.; Barnett, B. Warming temperatures will likely induce higher premium rates and government outlays for the U.S. crop insurance program. Agric. Econ. 2018, 49, 635–647. [Google Scholar] [CrossRef] [Scilit]
  13. Perry, E.D.; Yu, J.; Tack, J. Using insurance data to quantify the multidimensional impacts of warming temperatures on yield risk. Nat. Commun. 2020, 11, 4542. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Crane-Droesch, A.; Marshall, E.; Rosch, S.; Riddle, A.; Cooper, J.; Wallander, S. Climate Change and Agricultural Risk Management Into the 21st Century; Economic Research Report Number 266; U.S. Department of Agriculture (USDA) Economic Research Service: Washington, DC, USA, 2019.
  15. Reyes, J.J.; Elias, E. Spatio-temporal variation of crop loss in the United States from 2001 to 2016. Environ. Res. Lett. 2019, 14, 074017. [Google Scholar] [CrossRef] [Scilit]
  16. Brunsdon, C.; Fotheringham, A.S.; Charlton, M.E. Geographically Weighted Regression: A Method for Exploring Spatial Nonstationarity. Geogr. Anal. 1996, 28, 281–298. [Google Scholar] [CrossRef] [Scilit]
  17. Fotheringham, A.S.; Yang, W.; Kang, W. Multiscale Geographically Weighted Regression (MGWR). Ann. Am. Assoc. Geogr. 2017, 107, 1247–1265. [Google Scholar] [CrossRef] [Scilit]
  18. Breiman, L. Random Forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef] [Scilit]
  19. Georganos, S.; Grippa, T.; Niang Gadiaga, A.; Linard, C.; Lennert, M.; Vanhuysse, S.; Mboga, N.; Wolff, E.; Kalogirou, S. Geographical random forests: A spatial extension of the random forest algorithm to address spatial heterogeneity in remote sensing and population modelling. Geocarto Int. 2021, 36, 121–136. [Google Scholar] [CrossRef] [Scilit]
  20. Seamon, E.; Ridenhour, B.; Miller, C.; Johnson-Leung, J. Predictive Spatial Modeling of Sociodemographic Risk for COVID-19 Mortality. BMC Public Health, 2026; in press. [CrossRef] [Scilit]
  21. U.S. Department of Agriculture, Risk Management Agency. Cause of Loss Historical Data Files; U.S. Department of Agriculture, Risk Management Agency: Washington, DC, USA, 2026.
  22. U.S. Department of Agriculture, Risk Management Agency. Summary of Business Historical Data Files; U.S. Department of Agriculture, Risk Management Agency: Washington, DC, USA, 2026.
  23. U.S. Bureau of Labor Statistics. Bureau of Labor Statistics Consumer Price Index (CPI-U); U.S. Bureau of Labor Statistics: Suitland, MD, USA, 2026.
  24. Abatzoglou, J.T.; Dobrowski, S.Z.; Parks, S.A.; Hegewisch, K.C. TerraClimate, a high-resolution global dataset of monthly climate and climatic water balance from 1958–2015. Sci. Data 2018, 5, 170191. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Wright, M.N.; Ziegler, A. ranger: A Fast Implementation of Random Forests for High Dimensional Data in C++ and R. J. Stat. Soft. 2017, 77, 1–17. [Google Scholar] [CrossRef] [Scilit]
  26. Seamon, E. gwrf: Geographically Weighted Random Forests; CRAN: Vienna, Austria, 2026. [Google Scholar] [CrossRef] [Scilit]
  27. Lobell, D.B.; Hammer, G.L.; McLean, G.; Messina, C.; Roberts, M.J.; Schlenker, W. The critical role of extreme heat for maize production in the United States. Nat. Clim. Change 2013, 3, 497–501. [Google Scholar] [CrossRef] [Scilit]
  28. Grossiord, C.; Buckley, T.N.; Cernusak, L.A.; Novick, K.A.; Poulter, B.; Siegwolf, R.T.W.; Sperry, J.S.; McDowell, N.G. Plant responses to rising vapor pressure deficit. New Phytol. 2020, 226, 1550–1566. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Novick, K.A.; Ficklin, D.L.; Stoy, P.C.; Williams, C.A.; Bohrer, G.; Oishi, A.; Papuga, S.A.; Blanken, P.D.; Noormets, A.; Sulman, B.N.; et al. The increasing importance of atmospheric demand for ecosystem water and carbon fluxes. Nat. Clim. Change 2016, 6, 1023–1027. [Google Scholar] [CrossRef] [Scilit]
  30. Lesk, C.; Anderson, W.; Rigden, A.; Coast, O.; Jägermeyr, J.; McDermid, S.; Davis, K.F.; Konar, M. Compound heat and moisture extreme impacts on global crop yields under climate change. Nat. Rev. Earth Environ. 2022, 3, 872–889. [Google Scholar] [CrossRef] [Scilit]
  31. Strobl, C.; Boulesteix, A.L.; Kneib, T.; Augustin, T.; Zeileis, A. Conditional variable importance for random forests. BMC Bioinform. 2008, 9, 307. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  32. Nicodemus, K.K.; Malley, J.D.; Strobl, C.; Ziegler, A. The behaviour of random forest permutation-based variable importance measures under predictor correlation. BMC Bioinform. 2010, 11, 110. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  33. Tian, L.-x.; Zhang, Y.-c.; Chen, P.-l.; Zhang, F.-f.; Li, J.; Yan, F.; Dong, Y.; Feng, B.-l. How Does the Waterlogging Regime Affect Crop Yield? A Global Meta-Analysis. Front. Plant Sci. 2021, 12, 634898. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Voesenek, L.A.C.J.; Bailey-Serres, J. Flood adaptive traits and processes: An overview. New Phytol. 2015, 206, 57–73. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  35. Setter, T.L.; Waters, I.; Sharma, S.K.; Singh, K.N.; Kulshreshtha, N.; Yaduvanshi, N.P.S.; Ram, P.C.; Singh, B.N.; Rane, J.; McDonald, G.; et al. Review of wheat improvement for waterlogging tolerance in Australia and India: The importance of anaerobiosis and element toxicities associated with different soils. Ann. Bot. 2009, 103, 221–235. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Yu, J.; Sumner, D.A. Effects of subsidized crop insurance on crop choices. Agric. Econ. 2018, 49, 533–545. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Analytical workflow for evaluating spatial heterogeneity in climate–agricultural insurance claim severity. County-month insurance records were combined with monthly climate data, and indemnity payments were adjusted to constant 2022 dollars using monthly CPI-U. Separate datasets were constructed for six commodity × damage-cause combinations. For each eligible focal county, a geographically weighted random forest was fit using the complete monthly panel from the focal county and its adaptive neighborhood of the 10 nearest unique county locations. Permutation importance represents local predictive usefulness and does not indicate effect direction or causality.
Figure 1. Analytical workflow for evaluating spatial heterogeneity in climate–agricultural insurance claim severity. County-month insurance records were combined with monthly climate data, and indemnity payments were adjusted to constant 2022 dollars using monthly CPI-U. Separate datasets were constructed for six commodity × damage-cause combinations. For each eligible focal county, a geographically weighted random forest was fit using the complete monthly panel from the focal county and its adaptive neighborhood of the 10 nearest unique county locations. Permutation importance represents local predictive usefulness and does not indicate effect direction or causality.
Agriculture 16 01978 g001
Figure 2. County-centered geographically weighted random-forest workflow. Agricultural insurance records were joined to monthly TerraClimate and temporal predictors at the county–month scale. For each focal county, an adaptive neighborhood of unique county locations was selected, and all eligible monthly observations from those counties were used to fit a local random forest. County-specific permutation importance values were subsequently summarized and mapped.
Figure 2. County-centered geographically weighted random-forest workflow. Agricultural insurance records were joined to monthly TerraClimate and temporal predictors at the county–month scale. For each focal county, an adaptive neighborhood of unique county locations was selected, and all eligible monthly observations from those counties were used to fit a local random forest. County-specific permutation importance values were subsequently summarized and mapped.
Agriculture 16 01978 g002
Figure 3. Total positive climate importance for (A) corn drought, (B) corn excess moisture/ precipitation/rain, (C) soybean drought, (D) soybean excess moisture/precipitation/rain, (E) wheat drought, and (F) wheat excess moisture/precipitation/rain. Values are the sum of positive local permutation importance across precipitation, maximum temperature, minimum temperature, vapor pressure deficit, actual evapotranspiration, and soil moisture. A common color scale is used across all six panels. Because the models were fitted separately, comparisons of raw importance magnitudes among models remain descriptive.
Figure 3. Total positive climate importance for (A) corn drought, (B) corn excess moisture/ precipitation/rain, (C) soybean drought, (D) soybean excess moisture/precipitation/rain, (E) wheat drought, and (F) wheat excess moisture/precipitation/rain. Values are the sum of positive local permutation importance across precipitation, maximum temperature, minimum temperature, vapor pressure deficit, actual evapotranspiration, and soil moisture. A common color scale is used across all six panels. Because the models were fitted separately, comparisons of raw importance magnitudes among models remain descriptive.
Agriculture 16 01978 g003
Figure 4. Dominant local climate predictor for (A) corn drought, (B) corn excess moisture/ precipitation/rain, (C) soybean drought, (D) soybean excess moisture/precipitation/rain, (E) wheat drought, and (F) wheat excess moisture/precipitation/rain. The mapped category is the climate variable with the highest positive permutation importance in each focal county. “No positive importance” indicates that all six local climate importance values were nonpositive. Dominance does not convey effect direction or causality.
Figure 4. Dominant local climate predictor for (A) corn drought, (B) corn excess moisture/ precipitation/rain, (C) soybean drought, (D) soybean excess moisture/precipitation/rain, (E) wheat drought, and (F) wheat excess moisture/precipitation/rain. The mapped category is the climate variable with the highest positive permutation importance in each focal county. “No positive importance” indicates that all six local climate importance values were nonpositive. Dominance does not convey effect direction or causality.
Agriculture 16 01978 g004
Figure 5. County-centered local permutation importance for monthly maximum temperature in (A) corn drought, (B) corn excess moisture/precipitation/rain, (C) soybean drought, (D) soybean excess moisture/precipitation/rain, (E) wheat drought, and (F) wheat excess moisture/precipitation/rain. Positive values indicate greater local predictive usefulness. Negative values do not indicate a beneficial temperature effect. A common color scale is used across all six panels. Because the models were fitted separately, comparisons of raw importance magnitudes among models remain descriptive.
Figure 5. County-centered local permutation importance for monthly maximum temperature in (A) corn drought, (B) corn excess moisture/precipitation/rain, (C) soybean drought, (D) soybean excess moisture/precipitation/rain, (E) wheat drought, and (F) wheat excess moisture/precipitation/rain. Positive values indicate greater local predictive usefulness. Negative values do not indicate a beneficial temperature effect. A common color scale is used across all six panels. Because the models were fitted separately, comparisons of raw importance magnitudes among models remain descriptive.
Agriculture 16 01978 g005
Figure 6. Temporal sub-period comparison of county-level total positive climate importance for (A) corn drought, (B) corn excess moisture/precipitation/rain, (C) soybean drought, (D) soybean excess moisture/precipitation/rain, (E) wheat drought, and (F) wheat excess moisture/precipitation/rain. Each panel compares GWRF estimates from models fitted separately for 1994–2008 and 2009–2022 for counties represented in both periods. Points represent counties, dashed lines indicate 1:1 correspondence, and annotations report Spearman rank correlation ( ρ ) and the number of matched counties. Each panel uses identical x-axis and y-axis limits within that model, although limits differ among models.
Figure 6. Temporal sub-period comparison of county-level total positive climate importance for (A) corn drought, (B) corn excess moisture/precipitation/rain, (C) soybean drought, (D) soybean excess moisture/precipitation/rain, (E) wheat drought, and (F) wheat excess moisture/precipitation/rain. Each panel compares GWRF estimates from models fitted separately for 1994–2008 and 2009–2022 for counties represented in both periods. Points represent counties, dashed lines indicate 1:1 correspondence, and annotations report Spearman rank correlation ( ρ ) and the number of matched counties. Each panel uses identical x-axis and y-axis limits within that model, although limits differ among models.
Agriculture 16 01978 g006
Table 1. County–month coverage and local training support for the six primary county-centered GWRF models. The adaptive bandwidth was 10 unique county locations, and each local forest used 500 trees. Local training observations are the county–month records used to fit each county-centered random forest; mean, minimum, and maximum values summarize support across focal counties.
Table 1. County–month coverage and local training support for the six primary county-centered GWRF models. The adaptive bandwidth was 10 unique county locations, and each local forest used 500 trees. Local training observations are the county–month records used to fit each county-centered random forest; mean, minimum, and maximum values summarize support across focal counties.
Local Training Observations
ModelCounty–MonthsCountiesMeanMinMax
Corn × Drought109,2521880595.0161553
Corn × Excess moisture121,1591934635.8191841
Soybeans × Drought108,8771722646.6251471
Soybeans × Excess moisture129,9661743759.1212018
Wheat × Drought74,5271797426.4142187
Wheat × Excess moisture80,7891921428.9351932
Table 2. Mean local permutation importance, with median in parentheses, for the six climate predictors. Rankings are most defensible within each model; raw magnitudes across separately fitted models are descriptive. Boldface identifies the predictor with the highest mean permutation importance within each model.
Table 2. Mean local permutation importance, with median in parentheses, for the six climate predictors. Rankings are most defensible within each model; raw magnitudes across separately fitted models are descriptive. Boldface identifies the predictor with the highest mean permutation importance within each model.
Modelppttmaxtminvpdaetsoil
Corn × Drought0.331 (0.331)0.767 (0.713)0.590 (0.575)0.711 (0.682)0.410 (0.409)0.373 (0.342)
Corn × Excess moisture0.355 (0.331)0.466 (0.440)0.446 (0.419)0.381 (0.371)0.461 (0.407)0.319 (0.265)
Soybeans × Drought0.317 (0.297)0.788 (0.760)0.589 (0.576)0.617 (0.599)0.411 (0.394)0.353 (0.312)
Soybeans × Excess moisture0.359 (0.346)0.485 (0.484)0.485 (0.483)0.411 (0.421)0.517 (0.499)0.305 (0.273)
Wheat × Drought0.221 (0.164)0.371 (0.293)0.369 (0.271)0.355 (0.278)0.277 (0.216)0.188 (0.143)
Wheat × Excess moisture0.252 (0.226)0.395 (0.364)0.380 (0.344)0.353 (0.315)0.363 (0.323)0.213 (0.175)
Table 3. Sensitivity of climate importance results to the adaptive neighborhood bandwidth.
Table 3. Sensitivity of climate importance results to the adaptive neighborhood bandwidth.
ModelTotal-Importance
Pearson r
Dominant-
Predictor
Agreement (%)
Predictor-Rank
ρ
Top Predictor
k = 10 k = 25
Roughness
Reduction (%)
Corn × drought0.89362.71.000Max. temp. → Max. temp.15.2
Corn × excess moisture0.88846.31.000Max. temp. → Max. temp.10.8
Soybeans × drought0.90265.61.000Max. temp. → Max. temp.18.4
Soybeans × excess moisture0.87152.11.000AET → AET5.5
Wheat × drought0.86143.50.943Max. temp. → Min. temp.19.3
Wheat × excess moisture0.81339.61.000Max. temp. → Max. temp.14.7
Notes: Total-importance Pearson r compares county-level total positive climate importance between the k = 10 and k = 25 models. Dominant-predictor agreement is the percentage of matched counties assigned the same highest-ranked positive climate predictor. Predictor-rank ρ is the Spearman correlation between the model-wide rankings of the six climate predictors. The arrow identifies the top-ranked predictor under k = 10 and k = 25 , respectively. AET denotes actual evapotranspiration. Roughness reduction is the percentage decrease in standardized adjacent-county roughness under k = 25 relative to k = 10 ; larger values indicate greater spatial smoothing.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Seamon, E. Spatially Varying Climate Importance in United States Agricultural Insurance Claim Severity. Agriculture 2026, 16, 1978. https://doi.org/10.3390/agriculture16181978

AMA Style

Seamon E. Spatially Varying Climate Importance in United States Agricultural Insurance Claim Severity. Agriculture. 2026; 16(18):1978. https://doi.org/10.3390/agriculture16181978

Chicago/Turabian Style

Seamon, Erich. 2026. "Spatially Varying Climate Importance in United States Agricultural Insurance Claim Severity" Agriculture 16, no. 18: 1978. https://doi.org/10.3390/agriculture16181978

APA Style

Seamon, E. (2026). Spatially Varying Climate Importance in United States Agricultural Insurance Claim Severity. Agriculture, 16(18), 1978. https://doi.org/10.3390/agriculture16181978

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop