Next Article in Journal
Navigating the Path to AI and Virtual Immersion: An Exploratory Study of Educational Escape Rooms with the ED-SCALE Model
Previous Article in Journal
SV-GEN: Synergizing LLM-Empowered Variable Semantics and Graph Transformers for Vulnerability Detection
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Explaining Seasonal 5G Path Loss in a Vineyard: From Empirical Models to Interpretable Machine Learning

1
Research Group E-Government, Faculty of Computer Science, University of Koblenz, 56070 Koblenz, Germany
2
Research Group Computer Networks, Faculty of Computer Science, University of Koblenz, 56070 Koblenz, Germany
*
Authors to whom correspondence should be addressed.
Future Internet 2026, 18(5), 237; https://doi.org/10.3390/fi18050237
Submission received: 18 March 2026 / Revised: 17 April 2026 / Accepted: 22 April 2026 / Published: 28 April 2026
(This article belongs to the Section Smart System Infrastructure and Applications)

Abstract

Radio network planning is critical for 5G deployments, particularly for temporary installations in rural areas where terrain and vegetation significantly impact signal propagation. While empirical path loss (PL) models characterize propagation environments through scenario-specific parameters—leading to inherently noisy predictions at individual sites—machine learning (ML) approaches can predict site-specific path loss from multiple features simultaneously. This study conducts a systematic literature review of rural path loss prediction methods and introduces a novel dataset collected via a 5G nomadic measurement platform in a vineyard environment, capturing real-world propagation characteristics. We present a comprehensive comparison of machine learning and interpretable machine learning techniques, demonstrating that vegetation dynamics (quantified through the Normalized Difference Vegetation Index, NDVI) is an important driver of path loss variability when combining data across seasonal campaigns—though not within individual campaigns, where distance dominates. Cross-campaign NDVI transfer, however, is sensitive to satellite resolution, which appears to conflate vine canopy with seasonally managed inter-row ground cover. In cross-campaign transfer, XGBoost proves substantially less susceptible to NDVI-induced degradation than Explainable Boosting Machines (EBM), and a hybrid Log-Normal Shadowing (LNS) and XGBoost model confirms that NDVI captures seasonal variability more effectively than empirical path loss parameters alone. Still, the data captured the expected seasonal trend between April and June 2025, from which our interpretable models derived useful propagation insights. Tree-based models like Random Forest and XGBoost achieved the highest prediction accuracy ( R 2 up to 0.924 on individual campaigns, 0.891 on combined data, and up to 0.945 (individual) and 0.907 (combined) with antenna pattern-corrected path loss), while explainable boosting machines achieved near-parity ( R 2 up to 0.919; 0.876 on combined data) with the advantage of interpretability. Among individual campaigns, June—with densest canopy cover—yielded the highest R 2 values. These findings provide actionable insights for optimizing temporary 5G networks in precision agriculture and other rural applications.

Graphical Abstract

1. Introduction

Smart farming (SF) involves a variety of digital technologies to increase the efficiency of autonomous farming operations. These technologies require high data rates and low latency for reliable wireless networks [1]. Farms are typically located in rural and sparsely populated areas, where SF benefits are held back by the lack of reliable internet connectivity [2]. Without universal public LTE (Long Term Evolution) or 5G coverage in such areas, SF requires tailored wireless networks to meet its broadband demand.
To address these challenges, 5G nomadic networks are considered a viable solution for SF [3], providing short-term on-demand connectivity where local conditions change frequently. Vineyards are one such setting as follows: they span a wide range of conditions—from flat to very steep, from homogeneous to complex terrain—each posing distinct challenges for wireless coverage. In steep-slope viticulture specifically, labor-intensive operations such as soil cultivation, defoliation, and spraying are prime candidates for automation, making reliable data communication infrastructure a prerequisite for economically viable farming. The present study extends one of the steep-slope vineyards from [4] across three seasonal campaigns. Each such site presents a unique microenvironment, so that signal performance may differ significantly even over short distances [1,5]. While foliage attenuation is a known concern at millimeter-wave frequencies, it has also been quantified in the sub-6 GHz range, e.g., for WiMAX at 3.5 GHz over mixed rural terrain with seasonal campaigns [6]. Our study extends this line of work to 5G NR, combining seasonal measurements with interpretable ML to decompose the vegetation effect into individual feature contributions.
The goal of this paper is to quantify the effect of the rural microenvironment—in particular seasonal vegetation—on 5G path loss using empirical and machine learning (ML) models, and to decompose that effect into individual environmental drivers using interpretable ML.
Traditionally, empirical path loss (PL) models and machine learning models have been used for PL prediction in different environments. Empirical models are transparent and computationally efficient, but they rely on simplified parametric assumptions that may not capture the variability of complex microenvironments [7]. While black-box ML models (RF, XGBoost, and MLP) can handle microenvironments well, their lack of transparency makes it hard to trust their predictions. Glass-box ML models offer an inherently interpretable alternative [8] that need not sacrifice much accuracy [9] as recent work has shown. For propagation modeling, interpretability is particularly valuable as follows: an additive model structure allows decomposing predictions into individual feature contributions, turning path loss prediction into an explanation of which environmental factors drive signal attenuation.
Prior ML-based path loss studies are typically based on single snapshot data and therefore do not account for, e.g., seasonal vegetation dynamics or their effect on model transferability across time [7,10,11]. In contrast, this paper presents a new path loss dataset from a measurement campaign in a vineyard of the Moselle Valley. We examine how seasonal variability of vegetation impacts PL predictions by introducing the Normalized Difference Vegetation Index (NDVI), a satellite-derived measure of vegetation density, as a feature for modeling path loss in environments with dynamic foliage. The main contributions of this work are as follows:
1.
A new 5G path loss dataset at 3.75 GHz from three seasonal measurement campaigns in a steep-slope vineyard, capturing the April–June vegetation transition of the year 2025.
2.
Empirical path loss characterization via log-normal shadowing (LNS) and a comparison of glass-box (Explainable Boosting Machine (EBM) and Generalized Additive Model (GAM)) and black-box (Random Forest (RF), XGBoost, and Multi-Layer Perceptron (MLP)) models, showing that EBM achieves near-parity with the best black-box model while providing inherently interpretable predictions.
3.
Introduction of NDVI as a vegetation feature for path loss prediction, improving multi-campaign models while providing evidence that cross-campaign transfer may benefit from higher satellite resolution. Cross-campaign degradation is model-dependent, with XGBoost less susceptible than EBM.
4.
A hybrid LNS and XGBoost evaluation demonstrating that NDVI captures seasonal variability more effectively than empirical path loss parameters.
5.
A seasonal path loss decomposition using EBM’s additive structure, attributing the majority of the April-to-June path loss increase to vegetation change.
The remainder of this paper is organized as follows. Section 2 reviews related work on rural path loss modeling. Section 3 describes our 5G nomadic measurement platform and the dataset collected across three seasonal campaigns. Section 4 presents the empirical path loss analysis. Section 5 introduces the machine learning models used, and Section 6 reports the prediction results along with the interpretability analysis. Section 7 concludes the paper.

2. Related Work

To contextualize our study, we first conducted a systematic literature review on 5G path loss modeling for rural environments guided by the PRISMA methodology [12]. To identify the relevant work, we conducted enriched research from the databases Scopus, Web of Science, IEEE Xplore, Springer, and ScienceDirect. To encompass a vast spectrum of pertinent papers, we defined the following search string:
path   loss   model   and   5 G   and   rural .
The initial search yielded a total of 313 papers. Since the search spanned several databases, we removed the duplicates of the papers, which reduced the number of papers to 299. The screening step involved a title search, and we selected the papers that were applicable to 5G path loss modeling in a rural environment, resulting in a total of 65 papers. The next step involved a thorough examination of the abstract and a rigorous review of all the publications to see if the findings were applicable to our area of interest and excluded path loss modeling studies involving drones and railway networks in a rural context. This scrutiny resulted in 11 papers for our study. The identified literature can be broadly classified into two categories. The first category of studies utilizes typical empirical models to predict the path loss in a rural environment, while the second category of studies utilizes advanced modeling techniques like ML and DL to predict PL.
Table 1 summarizes the findings from the first category of papers. These studies utilized empirical models to predict path loss across a wide range of frequencies (from 800 MHz to 75 GHz) in various rural scenarios. The table reveals that there is no universal empirical model that is valid for all rural scenarios. In low-frequency ranges (1–15 GHz), traditional models like WINNER II and Okumura-Hata remain relevant for path loss prediction [13,14]. In contrast, for higher frequency ranges (16–60 GHz), there is more reliance on advanced empirical models like the Close-In (CI) model [15] and FI model [16] for path loss prediction.
Moreover, the Aalto1, KAIST2 [17], and Weissberger models [18] demonstrated better adaptation to environment-specific factors like vegetation and terrain features. Similarly, for rural macro scenarios the 3GPP RMa model [1] and the FI model [16] showed better performance. Empirical models are simpler but often fail to capture local environmental variability, leading to less accurate path loss predictions [7].
Table 1. Empirical models.
Table 1. Empirical models.
PaperFrequency; ScenarioRural Environment Details; LocationPL ModelPerformance IndicatorImportant Results
[13]3.4–3.8 GHz; urban, suburban, and ruralSuburban slope and Rural village; Zurich, Ittigen, Meikrich (Switzerland)WINNER II, ECC-33 Model, SUI Terrain C Model, 3GPP RMa NLOS ModelRMSE, comparison with measured path lossWINNER II D1 rural NLOS closest match among models tested for rural (vs. 3GPP RMa NLOS); both overestimate path loss.
[15]28 GHz, 38 GHz, and 73 GHz; urban, suburban, and ruralRural area with large-scale path loss; Hong Kong (China)Close-In (CI) model, The ABG modelEstimated path loss, MAPL, Cell Radius, Cell Area, Base StationsClose-In (CI) model provided best fit for estimating cell radius and coverage in rural areas.
[14]15 GHz; urban, ruralGeneral rural scenario with tri-sector antenna; location not specifiedOkumura-Hata model is used for rural scenario, Macro cell propagation model used for urban and suburban scenarioThroughput, Fairness index, Spectral EfficiencyOkumura-Hata model suitable for rural environments; statistical results prove that the 15 GHz spectrum is available for use
[16]40 GHz; RuralRural macrocell with minimal foliage; Tanjong Karang, Selangor (Malaysia)Empirical FI (Floating Intercept) model, CI modelRMSEThe FI model is most effective and shows the lowest RMSE, the CI model is effective in Cross-Polarized Directional antennas under LOS conditions
[18]60 GHz; Greenhouse (Rural)Greenhouse (rural-like) environment; Universidad Nacional de Colombia, Bogotá campus (Colombia)3GPP Indoor Office model (InH-LOS), Weissberger modelPath loss3GPP InH-LOS model fit best in propagation parallel to furrow; Weissberger worked well for propagation perpendicular to furrow
[17]25.5, 26 GHz, 800 MHz; RuralRural Finnish Forest with dense coniferous trees, 40–700 m vegetation depth; Pornainen (Finland)FITU-R, Weissberger, COST235, KAIST1, KAIST2, MED (Aalto1), MA (Aalto2)RMSEAalto1 best overall fit; KAIST2 also suitable in high vegetation depth.
[1]28, 38, 60, 75 GHz; Urban, RuralRural macro scenarios in simulations5GCM, 3GPP, METIS, and mmMAGICPath loss3GPP RMa model best-suited for rural macro environments; 5GCM not suitable in rural scenarios.
The second group of studies explored advanced ML and DL models to overcome the limitations of empirical models. We identified only four studies in this category for rural environment. Table 2 provides summaries of studies that investigated ML/DL models for path loss prediction in rural environments. The identified studies are focused on Sub-6 GHz bands and cover a variety of rural scenarios including coastal and vegetation-specific. These studies show a paradigm shift toward an environment-aware, data-aware path loss modeling. While refs. [7,10] emphasize AI’s role in extracting site-specific features from geospatial and visual data, ref. [11] stresses the influence of physical terrain properties (vegetation and coastal surfaces) on high-frequency signals. These studies consistently report lower prediction errors for ML models compared to empirical baselines—an expected outcome, since empirical models characterize environment classes rather than individual links. Further, ref. [19] applied ANFIS (Adaptive Neuro-Fuzzy Inference System) to fringe areas in Uttarakhand, India, using 1800 MHz drive-test data. Their hybrid model achieves lower RMSE than traditional empirical models (Hata, COST-231) while offering simplicity and generalization. Beyond these two categories, we also note hybrid approaches that combine empirical and ML models. For example, ref. [20] proposed a 5G mid-band hybrid model at 3.5 GHz that combines a re-optimized COST-231 Hata baseline with Random Forest predictions via weighted averaging. Similarly, ref. [21] corrects an extended-Hata baseline with a CNN-MLP trained on satellite imagery and 3D city-map features for sub-6 GHz private 5G in urban environments. However, these hybrid studies rely on urban empirical base models and, to the best of our knowledge, no hybrid model has been evaluated for rural vegetation environments. We address this gap in Section 6.2 by evaluating a hybrid LNS and XGBoost model—where XGBoost is trained on the LNS residuals—on our vineyard dataset.
Notably, none of the identified studies employed interpretable or explainable ML models for path loss prediction in rural environments, leaving a gap in understanding which environmental factors drive model predictions. Related work by collaborating authors includes glass-box models (GAM, EBM, and NAM) applied to an urban campus dataset [9] and 5G path loss coefficients at 3.75 GHz for the vineyard site studied here and other rural sites [4], though without shadowing characterization. We extend the latter with our new O-RAN platform at a different antenna position and a complete log-normal shadowing analysis in Section 4.

3. 5G Nomadic Platform and Measurement Campaign

This section outlines the methodology employed in the study to measure 5G signal propagation in a rural environment. First, we describe the hardware and measurement campaign; then, we characterize the signal strength measurements along with the compiled dataset for path loss prediction.

3.1. Hardware and Measurement Campaign

The 5G network utilized for this study was a customized 5G nomadic platform as shown in Figure 1 and Figure 2. The presented system is O-RAN compliant and consists of the following components: an outdoor Radio unit (Benetel RAN650), an outdoor directional antenna (Alpha Wireless AW3924), a 3GPP Rel 16 compliant 5G Core (Genius core), RAN Software (Airpuls RAN based on the OpenAirInterface project), a PTP Grandmaster Fronthaul switch (μFalcon-RX), a Network Switch (MikroTik Cloud Router CRS312-4C + 8XG-RM), as well as an Edge Cloud. All equipment is housed in a ruggedized 19-inch rack along with power backup batteries. This entire setup is mounted on a mobile cart that also includes a mast for mounting the radio unit. For precise timing, an outdoor Global Navigation Satellite System (GNSS) antenna is attached to the Fronthaul switch. Internet connectivity is provided by the Starlink satellite modem. The relevant configuration parameters for the measurement campaigns are summarized in Table 3.
The 5G signal strength measurements were obtained using Keysight’s Nemo Handy solution installed on a Samsung S24 Ultra with Keysight-modified firmware. The Nemo Handy solution supports 5G NR network measurements and logs various signal strength metrics to the smartphone’s internal storage. The metric used in this study is Layer 1 filtered SS-RSRP, measured on the secondary synchronization signal (SSS) of the serving cell’s synchronization signal block (SSB) as defined in 3GPP TS 38.215. Comparative measurements with a VIAVI ONA-800 spectrum analyzer across two independent campaigns show a consistent median offset of ≈4.5 dB in per-beam SS-RSRP (interquartile range ≈ 10 dB) between Nemo Handy and the calibrated instrument, confirming comparability with calibrated RF test equipment for this measurement site.
The measurement campaign was conducted in 2025 in a steep-slope vineyard (hereinafter referred to as Arena) located in Bernkastel-Kues in the Moselle Valley, a region renowned for extreme viticulture conditions. It consists of three measurement series from 1 April, 9 May and 27 June. As shown in Figure 3, the site features steep gradients, varied elevation, varying foliage, and ferromagnetic support structures (posts and wires).
To explore the seasonal effects of propagation through the vineyard, three measurement campaigns were conducted during the growing season (April–June) in Arena while considering dynamic vegetation changes (cf. Figure 3 for foliage density and type of vine plants). Arena exhibited significant seasonal variations in foliage density, as evidenced by NDVI values ranging from 0.19 (bare soil/minimal vegetation) in early April to 0.80 (dense canopy) in late June. During peak growth in May and June, the median NDVI increased from 0.29 to 0.66, reflecting the development of dense vine canopies and intercrop vegetation.
Each campaign involved traversing vineyard rows using a calibrated Nemo Handy device, collecting georeferenced RSRP samples at consistent intervals. Measurement points were recorded at 1.5 m height to simulate robotic sensor positioning, with terrain conditions documented for each sample.

3.2. Path Loss Dataset Analysis

Subsequent post-processing of field measurements resulted in a comprehensive dataset for modeling path loss behavior in the vineyard environment. The raw measurement logs contained multiple RSRP samples at identical GNSS coordinates due to continuous recording while stationary or slow-moving; these were aggregated by averaging all measurements sharing the same coordinates. This coordinate-based averaging serves a similar purpose as the Lee method adopted in ITU-R Recommendation SM.1708-1 [22], which prescribes spatial averaging to extract the local mean signal level by suppressing small-scale (multipath) fading. While our procedure averages per GNSS coordinate rather than over a fixed spatial window, the effect is comparable, as follows: it removes fast fading and also mitigates orientation-dependent RSRP fluctuations caused by varying device attitude and hand grip. Consequently, both the empirical LNS parameters and the ML models characterize the local mean path loss—the quantity relevant for network planning—rather than instantaneous signal realizations. In total, 443,118 raw measurement rows were recorded across the three campaigns, of which 21,815 contained valid GNSS coordinates and RSRP values. After coordinate-based averaging, the resulting dataset comprises 11,359 unique measurement locations enriched with terrain profile information including longitude, latitude, elevation, clutter height, Normalized Difference Vegetation Index (NDVI), and distance between transmitter and receiver. The longitude, latitude, elevation, and RSRP were directly extracted from the Nemo Handy log files, while clutter height, distance, path loss, and NDVI values were derived through post-processing techniques. Clutter height values were derived using the Raster Calculator tool to compute the difference between the AW3D30 (DSM) and EU-DEM (DTM) rasters of the area, sampled at measurement coordinates via QGIS’s Point Sampling Tool. This method provided estimates of above-ground object heights such as vegetation and structures. The distance between each measurement point (receiver) and the base station (transmitter) was calculated using the geodesic method implemented via the geopy.distance.geodesic function in Python 3.12. The Geodesic function implements Karney’s (2013) algorithm for geodesics on the WGS84 ellipsoid [23]. The relationship between path loss and measured RSRP is given by the standard path loss equation as follows:
P L = P T X + G T X + G R X R S R P
where the following hold:
  • P L : Path Loss (dB).
  • P T X : Transmit power (dBm).
  • G T X : Transmit antenna gain (dBi).
  • G R X : Receive antenna gain (dBi).
  • R S R P : Reference Signal Received Power (dBm).
We set G R X = 0 dBi (isotropic assumption), since smartphone antenna gains vary with device orientation and hand grip, and no orientation data was recorded. This is consistent with the isotropic UE antenna assumption used in the 3GPP channel model calibration [24]. Any systematic offset is absorbed into the path loss intercept and does not affect the relative model comparisons.
Using the Effective Isotropic Radiated Power (EIRP) formulation, i.e., G T X is set to its maximum value regardless of transmitter–receiver direction, we denote the corresponding path loss as P L E I R P . Although this approach of neglecting antenna directionality is less suitable for empirical path loss evaluation, it is well-suited for machine learning applications. The direct relationship between EIRP path loss P L E I R P and RSRP values enables the ML path loss prediction model to also predict RSRP values while implicitly learning the antenna pattern from the RSRP data.
The Normalized Difference Vegetation Index (NDVI) quantifies vegetation density, serving as a critical feature for path loss modeling in this scenario. NDVI values were derived from Sentinel-2 Level-1C (L1C) satellite imagery (10 m resolution) using Band 8 (NIR) and Band 4 (Red), sampled at measurement coordinates via QGIS’s Point Sampling Tool. Imagery from 3 April, 10 May, and 30 June 2025 (with <10% cloud coverage) was selected to align temporally with measurement campaigns, assuming minimal landscape changes between acquisition and fieldwork (cf. Figure 4).
While this method captures point-specific vegetation, it omits spatial variability within the immediate surroundings. Figure 5 presents a heat map of the RSRP (Reference Signal Received Power) values collected along the vineyard measurement route, with a total of 11,359 measurement points visualized. Reference Signal Received Power measured values are presented as color-coded points to represent signal quality, ranging from strong signals in green (e.g., 61 dBm) to weak and very poor signals in red (e.g., below 115 dBm). The bottom-right region of the vineyard shows predominantly green and yellow points, indicating strong to moderate RSRP values, typically between 61 dBm and 85 dBm. This suggests favorable conditions in this area, likely due to direct line-of-sight (LOS) with the base station and minimal obstruction. In contrast, the upper-left region of the vineyard is dominated by orange and red points, with RSRP values dropping below 100 dBm, indicating severe signal attenuation or possible non-line-of-sight (NLOS) conditions. This decline in signal quality across the route aligns with increasing distance from the antenna and the presence of potential obstructions such as vegetation, terrain variations, and structural elements within the vineyard. The maps also illustrate the GNSS positioning accuracy of the smartphone receiver. Overall, the trajectories follow coherent paths along the vineyard rows with good positional consistency between successive samples, indicating that the GNSS receiver is not subject to random scatter of several meters. However, the maps reveal noticeable positional deflections in certain vineyard rows, particularly in May and June: trajectories are sometimes bent sideways, seemingly crossing into a neighboring row, which did not occur during the actual measurement. This is consistent with the growing vine canopy partially shadowing the GNSS signal in localized sections. In June, the trajectories also appear more irregular overall, which is largely attributable to dense inter-row vegetation impeding a straight walking path rather than to GNSS errors alone. Nevertheless, even where deflections are visible on the map, the positional errors remain small compared to the overall distance to the base station, which is the dominant feature for path loss prediction in both empirical models and ML models, as the EBM interpretability analysis in Section 6 will confirm. However, the effect of GNSS errors on latitude—which emerges as the most influential ML feature after distance and, when included, NDVI—is less clear and warrants further investigation, e.g., by mounting the GNSS antenna above the canopy in future campaigns.
The combined dataset of three measurement campaigns includes 11,359 unique measurement locations and 10,133 valid path loss values ( P L E I R P ). The spatial data—latitude, longitude, and elevation—provides high-resolution coverage across the vineyard, with elevation values ranging from 229 m to 262 m. The median clutter height is approximately 3.5 m, representing typical vineyard features like posts, wires, and vegetation. The distance of measurement locations from the base station varies from 0.07 m to 104.38 m, covering both near-field and far-field propagation conditions. Corresponding path loss values ( P L E I R P ) range from 113.15 dB to 169.60 dB, capturing both line-of-sight (LOS) and non-line-of-sight (NLOS) scenarios. A standard deviation of 9.36 dB in path loss ( P L E I R P ) indicates considerable signal variation due to terrain and clutter.
The statistical summaries and correlation coefficient matrices for each measurement campaign (April–June) are presented in Table 4, Table 5 and Table 6 and Table 7, Table 8 and Table 9, respectively. Statistical analysis revealed significant seasonal shifts in propagation conditions. Path loss increased monotonically (April: 141.73 dB to June: 150.85 dB), while NDVI showed non-linear progression, dropping in May (0.30) before peaking in June (0.64). Correlation analysis uncovered the following critical dynamics: distance maintained a stable relationship with path loss (Pearson correlation r 0.81 ), but NDVI’s role shifted from moderately positive in May ( r = 0.52 ) to weakly negative in June ( r = 0.09 ), suggesting mature canopies permit alternative propagation mechanisms. The near-perfect latitude-elevation collinearity ( r > 0.95 ) indicates terrain slope dominates micro-topographic effects in this vineyard environment.

4. Empirical Model Path Loss Analysis

From the literature review in Section 2, we see that only the models from [13] match the frequency range of our setting. We implement a log-normal shadowing (LNS) model that can be mathematically derived from the general WINNER II path loss framework. The WINNER II standard defines the general path loss formula as [25]:
P L ( d ) = A log 10 ( d ) + B + C log 10 f c 5.0 + X
where P L is the path loss, d is the distance between transmitter and receiver in meters, f c is the system frequency in GHz, A is the fitting parameter including the path loss exponent, B is the intercept parameter, C describes the path loss frequency dependence, and X is an optional environment-specific term (e.g., wall attenuation).
For single-frequency measurements with constant f c , the frequency term C log 10 ( f c / 5.0 ) is a constant that is absorbed into the intercept B, which then corresponds to P L ( d 0 ) in the LNS formulation. Additionally, using a reference distance d 0 and incorporating shadowing effects, this reduces to the LNS model as follows:
P L ( d ) = P L ( d 0 ) + 10 α log 10 ( d / d 0 ) + X σ ,
where the parameter α is called the path loss coefficient, and P L ( d 0 ) is the path loss value corresponding to the reference distance. Finally, X σ N ( 0 , σ 2 ) is a normally distributed shadowing variable with zero mean and standard deviation σ which typically models signal strength variations caused by static obstacles. This LNS formulation is fully compatible with the WINNER framework, where the shadowing term X σ generalizes the environment-specific term X .
The same equivalence holds for the other standard empirical path loss models used in the 5G literature. The Close-In (CI) reference distance model [26] defines P L CI ( d ) = FSPL ( f c , 1 m ) + 10 n log 10 ( d ) + X σ , where the intercept is anchored to free-space path loss at a 1 m reference distance. At a single carrier frequency, FSPL ( f c , 1 m ) is a known constant, so the CI model reduces to LNS with d 0 = 1 m and P L ( d 0 ) = FSPL ; the only difference is whether the intercept is physics-anchored or fitted. The Alpha-Beta-Gamma (ABG) model [26] is P L ABG ( d ,   f ) = 10 α log 10 ( d ) + β + 10 γ log 10 ( f c ) + X σ ; at a single frequency the term 10 γ log 10 ( f c ) is constant and absorbed into β , again yielding LNS. Similarly, the Floating Intercept (FI) model has both intercept and slope fitted, making it identical to LNS by construction. Thus, at a single carrier frequency, CI, ABG, FI, and WINNER II all collapse to the same log-distance model with Gaussian shadowing—any differences lie only in how the intercept is determined, not in the functional form. The classical Okumura-Hata [27,28] and COST 231 Hata [29] models are also log-distance in structure but prescribe frequency- and height-dependent coefficients derived from specific measurement campaigns; at fixed frequency and antenna height they again reduce to a two-parameter log-distance law, though with tabulated rather than fitted coefficients.
Moreover, while [13] suggests WINNER II D1 rural NLOS ( A = 25.1 , B = 55.4 , and C = 21.3 ) as the better-fitting of two rural models tested at this frequency range, applying its parameters to our measurements at 75–100 m yields predictions of approximately 100 dB—roughly 46 dB below the April and 58 dB below the June observations. The D1 model was designed for rural macro-cells with base station heights of 18–25 m in the underlying measurements and an application range of 20–70 m [25]. While the full formula includes a height correction term, a mast height of 4 m is far outside the model’s intended scope. Moreover, the model assumes level ground, whereas in our vineyard the terrain ranges from about 4 m below to several meters above antenna height, so that the base station is below many receiver locations on the hillside. These mismatches motivate fitting site-specific LNS parameters—which, as shown above, is the exact single-frequency form of the WINNER II framework—rather than relying on tabulated values. The same height limitation applies to the 3GPP TR 38.901 Rural Macro (RMa) model [24], which specifies a minimum base station height of 10 m (default 35 m), well above our 4 m mast. The resulting parameters α and σ characterize the propagation environment in the same tradition as path loss exponent tables for different environment types found in standard textbooks (e.g., [30]), and our seasonal values contribute vineyard-specific entries to this body of knowledge.
We determine P L ( d 0 ) by averaging the path loss for distances in an interval [ d 0 ε , d 0 + ε ] , with d 0 = 10 m and ε = 5 m.
To calculate the BS antenna’s 3D-pattern from the horizontal and vertical planes, we employ the simple addition algorithm (also denoted product algorithm) from [31]. When the path loss includes the antenna pattern, i.e., it is computed with gain values depending on the transmitter–receiver direction, we denote it by P L P A T . Figure 6 shows P L P A T across the three months both as scatter plot computed from the measurement data and as empirical model. The model parameters are listed in Table 10, while the summary of P L P A T is shown in Table 11.
As one would assume, the path loss coefficient increases with denser foliage. Notably, the shadowing standard deviation σ (Table 10) is consistently smaller than the total path loss standard deviation (Table 11); for example, 7.21 vs. 7.51 dB in April and 8.97 vs. 10.26 dB in June. This is expected, since σ captures the residual variation after removing the distance trend 10 α log 10 ( d / d 0 ) , whereas the total standard deviation includes the distance-dependent component. The relatively small difference indicates that distance alone explains only a modest fraction of the overall path loss variability, with terrain, clutter, and—in later campaigns—vegetation accounting for the remainder. Indeed, empirical models like LNS characterize a propagation environment by fitting a small set of scenario-specific parameters ( α , σ ) to an entire campaign—yielding a statistical description of the environment class rather than predictions at individual sites. This observation is consistent with [9], where three standard empirical models (Okumura-Hata, COST-231 Hata, ECC-33) have been applied in an urban environment at 1800 MHz. This motivates the multi-feature ML approach in the following section, which moves from environment characterization to site-specific prediction by incorporating these additional predictors to capture the variance that the distance-only LNS model cannot explain. Section 6.2 quantifies this gap by evaluating LNS as a predictor alongside the ML models.

5. ML Path Loss Model

This study evaluates three black-box and two glass-box ML models for path loss prediction in our scenario. For the ML models, we focus on the EIRP formulation of the path loss which is equivalent to RSRP prediction up to the maximum gain constants. Although the antenna pattern-corrected formulation ( P L P A T ) yields higher prediction accuracy, the EIRP formulation is retained as the primary target for its greater simplicity and practical applicability. Section 6.2 extends the analysis by evaluating both a hybrid LNS and XGBoost model and the effect of using P L P A T as the prediction target. The models were selected based on their strong performance in a previous campus-based study by collaborating authors [9] and their complementary characteristics in terms of predictive power and interpretability.

5.1. Black-Box ML Models

We evaluated the performance of the following three black-box models.
1.
Random Forest: RF is an ensemble learning method that constructs multiple decision trees on bootstrapped subsets of the training data and aggregates their predictions via averaging [32]. RF is robust to overfitting, handles nonlinear relationships, and provides straightforward feature importance estimates. Our RF models were optimized through 5-fold cross-validated hyperparameter tuning, exploring tree counts (200–2000), maximum depths (10–110), minimum samples per leaf (1–4), and feature selection strategies (‘sqrt’, ‘log2’). The optimal configurations varied across datasets, with tree counts ranging from 800–1400 and depths from 20–110. Notably, for larger datasets without NDVI, a maximum depth of 20 was found optimal, suggesting that shallower trees generalize better when vegetation features are absent.
2.
Extreme Gradient Boosting (XGBoost): XGBoost is an efficient implementation of gradient-boosted decision trees, which sequentially fits new trees to the residuals of previous ones to minimize a specified loss function [33]. XGBoost offers strong predictive accuracy and flexible regularization to control overfitting. XGBoost hyperparameters were tuned via 5-fold cross-validation, searching over tree counts (50–300), depths (3–10), learning rates (0.01–0.16), subsampling ratios (0.7–1.0), and regularization parameters. Optimal configurations featured 167–296 boosted trees with depths of 8–9 and learning rates of 0.023–0.072, balancing model complexity against generalization across the varied terrain and vegetation conditions.
3.
Multi-Layer Perceptron (MLP): We include MLP as a neural network baseline to ensure a broad comparison across ML model families. MLP is a feedforward artificial neural network that approximates complex, nonlinear mappings between features and target values [34]. We performed 5-fold cross-validated hyperparameter tuning over network architectures (1–3 hidden layers with 64–256 neurons), activation functions (ReLU, tanh, and logistic), the L-BFGS optimizer, learning rates (0.001–0.2), and L2 regularization (0.0001–0.1). Optimal configurations varied from compact two-layer networks (64, 32) to deeper architectures (200, 150, 100), with the L-BFGS solver consistently selected across all datasets. Tuning proved critical for MLP as follows: default configurations achieved negative R 2 on smaller datasets, while tuned models reached R 2 = 0.845–0.923.

5.2. Glass-Box ML Models

We introduce the following two intrinsically interpretable ML models:
1.
GAM (generalized additive model): GAMs are a semi-parametric extension of the generalized linear models that allow for non-parametric fittings of complex dependencies of response variables. We implement Generalized Additive Models (GAMs) as an interpretable baseline for path loss prediction using the pyGAM library. A GAM adopts a sum of arbitrary functions of variables (possibly nonlinear) that represent different features via splines, which altogether describe the magnitude and variability of the response variables [35]. The model’s additive structure ensures each feature’s impact can be isolated and analyzed, making it particularly suitable for identifying dominant propagation mechanisms in complex vineyard environments. While GAM performs well with default settings ( R 2 = 0.764–0.847), we applied 3-fold cross-validated hyperparameter tuning over smoothness penalty λ (0.1–100), number of splines per feature (10–35), spline order (2–4), and convergence parameters, sampling 100 random configurations. This yielded marginal improvements of 0.2–1.9%.
2.
Explainable Boosting Machine (EBM): EBM is a glass-box learning algorithm based on boosted and bagged shallow decision trees, which learn additive shape functions and optional pairwise interactions [36]. EBMs are highly interpretable, because the contribution of each independent variable or combination of independent variables to a final prediction can be visualized and understood by plotting. EBMs balance predictive accuracy with interpretability by providing human-readable feature contributions. We trained EBM models using the InterpretML package. EBM’s default configuration already achieves strong performance ( R 2 ≈ 0.834–0.913), demonstrating its robustness. As for GAM, we additionally performed 3-fold cross-validated hyperparameter tuning over learning rate (0.001–0.1), interaction terms (5–20), maximum leaves per tree (2–8), discretization bins (64–1024), and regularization parameters, sampling 100 random configurations. This provided similar improvements of 0.2–2.0%.

6. Model Accuracy Results

This section presents the accuracy results for PL models described in the previous section by utilizing the Arena path loss dataset. We describe the evaluation results in four parts. In the first part, we compare the accuracy of black-box and glass-box models in order to understand the effect of vegetation on their results. In the second part, we evaluate a hybrid LNS and XGBoost model that combines the empirical path loss baseline with ML-based residual prediction, and we assess the effect of using the antenna pattern-corrected path loss P L P A T as prediction target. In the third part, we present the EBM and XGBoost cross-campaign generalization to understand the seasonality effect on the model’s accuracy. In the fourth part, we present the EBM model’s interpretation results to understand how each feature contributes to the model’s predictions.

6.1. Model Accuracy Results with and Without NDVI

This section presents the prediction accuracies of the black-box ML and glass-box ML models using two different configurations of our measurement campaign dataset. In the first configuration, we took longitude, latitude, elevation, clutter height, and distance as input features and path loss as output feature. In the second configuration, we also included the NDVI values as input feature. For both configurations, prediction accuracies of the black-box ML and glass-box ML models are evaluated for individual campaigns as well as for the combined dataset. For the combined dataset, we randomly subsampled each campaign to 1495 measurements—the number of valid samples in April after removing rows with missing values—to ensure balanced representation across seasons. For training purposes, we utilized 80% data from each dataset and tested prediction accuracy on remaining 20% randomly selected data instances. Model prediction accuracies are evaluated by means of MAE (mean absolute error), RMSE (root mean square error), and R squared.
Analysis of the results shown in Table 12 and Table 13 reveals the superiority of Tree-based (RF/XGBoost) black-box models over glass-box models in terms of MAE, RMSE, and R 2 . The results reveal significant seasonal variations, with June exhibiting the best overall performance (RF R 2 = 0.924, XGBoost R 2 = 0.921) and April showing slightly lower prediction accuracy. EBM emerged as the most accurate interpretable model, achieving R 2 = 0.919 in June and competitive performance across all datasets. With proper hyperparameter tuning, MLP achieved strong results ( R 2 = 0.845–0.923), demonstrating that neural networks can effectively learn path loss patterns when appropriately configured. The combined dataset performed well ( R 2 of 0.841–0.852 for EBM and RF, respectively) with MAE values of 2–3 dB, which makes them acceptable for network planning and confirms that the selected features capture path loss variability across all three campaigns.

Role of NDVI Across Campaigns

In the second configuration, addition of vegetation data (NDVI) clearly improves results on the combined multi-campaign dataset. NDVI provides the most significant benefit when combining data across different seasonal campaigns: RF’s R 2 improves from 0.852 to 0.891 (+3.9%), and XGBoost improves from 0.851 to 0.888 (+3.7%). Similarly, EBM benefits substantially with R 2 increasing from 0.841 to 0.876 (+3.5%). For individual campaigns, however, NDVI shows minimal impact—sometimes even slightly reducing performance. The following two factors explain this observation: First, within a single campaign the vegetation state is relatively uniform, so the difference between areas with slightly higher or lower NDVI (e.g., 0.5 vs. 0.3 in April) does not translate to meaningful path loss differences when the overall canopy is sparse. Second, distance dominates path loss prediction ( r 0.81 ), leaving little residual variance for NDVI to explain. This finding indicates that NDVI’s primary value lies in distinguishing fundamentally different seasonal vegetation states rather than capturing fine-grained spatial variation within a campaign. However, as we discuss in the next section, using raw NDVI values introduces limitations for cross-campaign model transfer. Among interpretable models, EBM maintains competitive accuracy with RF on the combined dataset with NDVI ( R 2 = 0.876 vs. 0.891), with only 0.21 dB higher MAE. GAM shows the largest performance gap, achieving R 2 = 0.798 on the combined dataset with NDVI.
Although tree-based black-box models show superior performance in both configurations, EBM’s competitive accuracy (≤0.25 dB MAE difference on combined data) and native explainability make it compelling for propagation analysis in dynamic agricultural environments.

6.2. Hybrid Model: Combining LNS with XGBoost

We extend the PL analysis by adding a hybrid model that applies XGBoost to the residuals that the log-distance law of LNS leaves unexplained. We also include LNS itself in the comparison, although its main purpose is not to act as a predictor but to characterize the propagation environment. Since we explore a variety of cases, we focus on XGBoost as one of the best-performing black-box models from the previous section. The cases include LNS, XGBoost, and the hybrid model operating on both P L E I R P and P L P A T . We also re-tune XGBoost hyperparameters for P L P A T , which, however, does not have a significant effect. The LNS path loss coefficients fitted to the P L E I R P data are α = 1.78 , 2.51 , 2.97 for April, May, and June campaigns, respectively.
For individual campaigns, the hybrid model performs on par with pure XGBoost in both configurations (Table 14 and Table 15), confirming that XGBoost already captures the distance–path loss relationship from the training data; the LNS base contributes no additional information. The only notable difference appears on the combined dataset without NDVI, where the hybrid model improves the P L E I R P target R 2 from 0.851 to 0.864 ( P L P A T : 0.875 to 0.893). Here, the per-campaign path loss exponent α acts as an implicit seasonal proxy: it encodes the vegetation state that NDVI would otherwise provide. Accordingly, when NDVI is included, this advantage vanishes—the combined-dataset hybrid R 2 is 0.883, slightly underperforming pure XGBoost at 0.888, as the LNS base adds redundant information ( P L P A T : both 0.907). These results reinforce the finding from Section 4: empirical models characterize propagation environments through scenario-level parameters, whereas ML models predict site-specific path loss from multiple features simultaneously. Crucially, the hybrid model without NDVI at R 2 of 0.864 ( P L P A T : 0.893) still falls short of pure XGBoost with NDVI at 0.888 ( P L P A T : 0.907), demonstrating that satellite-derived vegetation indices—despite their resolution limitations—capture seasonal variability more effectively than empirical path loss parameters alone. Adding the LNS base to XGBoost with NDVI yields no further improvement at R 2 of 0.883 ( P L P A T : 0.907), confirming that NDVI already subsumes the seasonal information encoded in α . The P L P A T results consistently show higher R 2 but also slightly higher RMSE than P L E I R P ; all trends are preserved, and re-tuning XGBoost hyperparameters on P L P A T does not yield meaningful differences. For the LNS fit, P L P A T yields the physically more meaningful path loss exponent, as it characterizes the propagation environment without conflating it with antenna directivity. Since the conclusions are identical for both formulations, the remainder of the analysis uses P L E I R P for its greater simplicity and direct correspondence to RSRP prediction.

6.3. Cross-Campaign Performance

We evaluated EBM and XGBoost—the best-performing glass-box and one of the best-performing black-box models, respectively—trained on isolated campaign data to assess cross-campaign applicability. The results reveal sharp performance declines when models are deployed on other campaigns data, with the severity closely linked to changes in vineyard vegetation between campaigns. Table 16 and Table 17 present cross-campaign R 2 performance without and with NDVI as an input feature, respectively.
In these tables, the diagonal entries are evaluated on the full training campaign and serve as an upper-bound reference ( R 2 = 0.897 0.970 , where XGBoost shows higher diagonal values than EBM, reflecting its greater capacity to fit the training distribution). However, substantial performance drops occur when models are deployed to other campaigns. Cross-campaign generalization varies considerably: the April-trained model retains moderate performance on May (EBM R 2 = 0.570 , XGBoost R 2 = 0.574 ), while the June-trained model struggles on April (EBM R 2 = 0.160 , XGBoost R 2 = 0.127 , i.e., predictions worse than the test set mean). Both models show similar cross-campaign transfer patterns, confirming that the generalization limitation is inherent to the seasonal shift rather than model-specific.
Comparing Table 16 and Table 17 reveals a counterintuitive finding as follows: adding NDVI as a feature often degrades cross-campaign generalization. This effect is particularly pronounced for EBM, as follows: the May-trained EBM tested on June achieves R 2 = 0.465 without NDVI but R 2 = 0.523 with NDVI, whereas the corresponding XGBoost model degrades only marginally ( R 2 = 0.476 to 0.467 ). The most severe degradation occurs for the June-trained EBM tested on May, dropping from R 2 = 0.291 to R 2 = 2.151 ; XGBoost shows the same direction but not as extreme ( R 2 = 0.293 to 0.016 ). A likely explanation is that, at the 10 m Sentinel-2 resolution, each NDVI pixel captures a composite of vine canopy and inter-row ground cover whose relative contributions shift with vineyard management. Comparing Figure 3 and Figure 4 is consistent with this, as follows: in April, green ground cover appears to dominate the pixel; by May, this cover has been removed while vine leaves are only emerging, which would lower the composite NDVI despite advancing phenology. If so, cross-campaign NDVI transfer may succeed only when vegetation increases consistently across all sub-pixel components—as appears to be the case for the April-to-June transition—rather than trading one source of greenness for another. At coarse satellite resolution, this condition cannot be verified from the imagery alone and may require ground-truth inspection such as the campaign photographs in Figure 3. Simple normalization (e.g., z-score or percentile-based) would not resolve this ambiguity, since the issue is not one of scale but of what the index physically represents. Without NDVI, models rely on more stable geometric features (distance, elevation, and clutter height) that generalize better. The observation that XGBoost is less susceptible to NDVI-induced cross-campaign degradation than EBM may be attributed to XGBoost’s tree-based feature interactions, which can partially compensate for misleading NDVI values through splits on correlated geometric features, whereas EBM’s additive structure fits a dedicated NDVI shape function that is applied unconditionally.
These results align with vineyard vegetation dynamics across campaigns. Early campaigns (April) have sparse canopies and exposed soil, so propagation features are shaped primarily by ground reflection and bare-vine geometry. Models trained on these conditions fail to capture foliage-induced attenuation in later campaigns, leading to significant R 2 reductions when applied cross-seasonally. Late campaigns (June) represent advanced canopy development, producing propagation characteristics unlike earlier campaigns, explaining the severe performance drop when June-trained models are applied to April or May campaigns. This highlights the need for either normalized vegetation indices or combined datasets that contain seasonal data covering various stages of vegetation growth. Beyond vegetation dynamics, cross-campaign prediction is further complicated by slight variations in deployment parameters between campaigns—base station position, antenna height and direction, and platform levelness—which might have a large impact on the comparability of path loss values. These confounding factors make it difficult to reliably isolate the contribution of any single variable in cross-campaign transfer.

6.4. EBM Model Interpretation Results

In this section, we leverage EBM’s interpretability at both the global and local level. Global explanations reveal feature significance, marginal contributions, and feature interactions across the entire training set. Local explanations decompose individual predictions into per-feature contributions, enabling location-specific diagnostics. We begin with global feature importance. The global significance of input features is computed based on the average absolute value of each input feature’s contribution across the entire training set. Thereby, both large positive and large negative effects by individual predictions are considered important by this measure and are summed up. Figure 7 compares the global importance of features to predict PL using the combined dataset with and without the NDVI input feature. In Figure 7, bars represent the mean absolute contribution in dB, while percentages indicate each feature’s share of the total feature contribution, computed as | c j | / i | c i | where c i denotes the contribution of feature i (excluding the intercept, which represents the baseline prediction of 145.62 dB). Figure 7 shows that in both models—with and without NDVI—the EBM assigns the dominant contributions to individual main effects, while interaction terms remain low-ranked (each at most 5% of the total). In this hillside vineyard, the spatial features are physically intertwined: latitude, elevation, and distance to the base station all vary along the slope (cf. Table 7, Table 8 and Table 9). NDVI, as a satellite-derived vegetation measure, is the one feature that captures an independent physical quantity. When included, distance (3.30 dB, 20.4%) and NDVI (2.29 dB, 14.1%) together account for over a third of the total feature contribution. Distance’s contribution remains virtually unchanged between the two models, indicating that NDVI captures genuinely additional variance rather than redistributing distance’s explanatory power.

6.4.1. Marginal Effects

Figure 8 presents the marginal effects of distance, elevation, latitude and NDVI. Each marginal effect plot maps a feature’s value to its contribution to the predicted path loss. The histogram at the bottom of each sub-figure describes the density of the data distribution for that feature. The overall effect of distance from the transmitter is shown in the upper left-hand panel in Figure 8a. The effect of distance in the vineyard reveals a linear characteristic: from 0–45 m, distance has a negative effect on predictions, while from 45–100 m it has a strong positive effect, and then levels off with some fluctuation from that point onward. The effect of elevation is shown in the upper right-hand panel in Figure 8b. At the lowest elevations (229–234 m), the model predicts substantially higher path loss (up to + 8 dB). Scores cross zero near 235 m and remain slightly negative through the mid-range (235–247 m, down to ∼ 3 dB). Above 248 m the shape function is non-monotonic: scores briefly turn positive again (248–255 m, up to + 2 dB) before dropping negative at the highest elevations (256–262 m, down to 3.5 dB). The effect of latitude is shown in the lower left-hand panel in Figure 8c. At lower latitude values (49.9129–49.9130) latitude has a negative effect and the model predicts lower path loss. For middle latitude values (∼49.9131–49.9135), scores hover close to 0 dB and latitude has little effect on the predicted path loss. For higher latitude values (>49.9136) scores increase sharply to + 18 dB and thus the model predicts higher path loss. The marginal effect of NDVI is illustrated in the lower right-hand panel of Figure 8 (Figure 8d). The plot reveals a non-linear relationship between NDVI values and their contribution to the predicted signal strength. NDVI values below 0.3 exhibit slightly negative scores (∼ 2 dB), suggesting lower path loss when vegetation is sparse. Mid-range NDVI (0.3–0.5) values remain moderately negative (down to ∼ 4 dB); these scores indicate that the model predicts somewhat lower path loss at moderate vegetation density. High NDVI (0.6–0.8) values move to clearly positive scores (up to + 7 dB), suggesting increased path loss where vegetation density is high.

6.4.2. Decomposing the Seasonal Path Loss Increase

The empirical data show a clear increase in mean path loss ( P L E I R P ) from April (141.73 dB) to June (150.85 dB). To understand which features drive this +9.13 dB increase, we leverage EBM’s additive structure to decompose the change into individual feature contributions. The April-to-June transition is also the only cross-campaign direction where adding NDVI improves prediction accuracy (Table 16 and Table 17: R 2 increases from 0.068 to 0.280). As discussed in Section 6.3, this is consistent with vegetation increasing across all sub-pixel components during this transition. We therefore focus on this transition, excluding May due to its indecisive NDVI values (cf. Figure 4).
We identified 629 location pairs measured in both campaigns (42% of April locations; spatial matching tolerance: 1.5 m; mean matching distance: 0.67 m). These pairs are distributed across most of the vineyard rows covered by the April campaign, ensuring that the decomposition is not biased toward a particular subregion and that observed differences reflect temporal rather than spatial variation. For these matched locations, mean path loss increased from 143.52 dB to 150.91 dB (+7.39 dB).
Figure 9 shows the decomposition results for models with and without NDVI. Without NDVI (Figure 9a), the model attributes the path loss increase primarily to elevation (+1.85 dB, 25.0%) and various elevation interaction terms, spreading the explanation across correlated terrain features. With NDVI (Figure 9b), vegetation change emerges as the dominant factor, contributing +4.26 dB (57.6%) of the total increase—consistent with the vegetation transition from sparse early-season foliage (mean NDVI: 0.29) to dense summer canopy (mean NDVI: 0.64). Elevation contributes +1.52 dB (20.6%), likely reflecting systematic differences in measurement patterns between campaigns. Notably, distance contributes only 0.02 dB ( 0.2 %), confirming that the path loss increase is not an artifact of different measurement distances but a genuine seasonal effect driven primarily by vegetation growth.

6.4.3. High-NDVI-Increase Locations: Aggregate and Local Explanations

While global explanations reveal overall feature importance, local explanations can illuminate how NDVI affects predictions at specific locations. Focusing again on the April-to-June transition, we selected the 50 locations with the greatest NDVI increase from the 629 matched pairs identified above (mean NDVI: 0.34 to 0.71). Should the NDVI increase truly reflect vegetation growth rather than soil reflectance changes (cf. the resolution discussion in Section 6.3), NDVI would have to emerge as a significant contributor to the path loss prediction. EBM’s local explanations let us directly test whether it does.
Figure 10 uses the EBM trained on the combined dataset (same model as Figure 7b and Figure 9b). Figure 10a shows the mean absolute contributions averaged over these 50 locations, consistent with the global importance metric used above. NDVI emerges as the dominant feature at 3.78 dB (23.8% of total). The remaining main features contribute less as follows: distance at 1.82 dB (11.5%), clutter height at 1.77 dB (11.2%), and longitude at 1.53 dB (9.6%).
Figure 10b illustrates EBM’s interpretability by decomposing individual predictions into exact feature contributions—a capability absent in black-box models. To show the range of NDVI’s role across these 50 locations, we select the two samples where NDVI’s contribution share (as defined above for Figure 7) is smallest and largest, respectively.
At the location with the smallest NDVI effect (top panel, 7.2%), NDVI increased substantially (from 0.37 to 0.76) yet contributes only 1.26 dB to the prediction. The negative sign indicates that, at this particular combination of feature values, the learned NDVI shape function slightly reduces predicted path loss, which can be explained by NDVI’s contribution shifting from its main effect into interaction terms: the clutter height and NDVI (18.1%) and longitude and NDVI (12.6%) interactions together exceed the standalone NDVI term threefold. Even at this sample—selected for having the smallest NDVI main effect—vegetation-related terms collectively account for 38% of the prediction. Overall, distance ( 3.40 dB, 19.4%) and the clutter height and NDVI interaction ( 3.18 dB, 18.1%) dominate. At the location with the largest NDVI effect (bottom panel, 40.9%), NDVI dominates at + 4.72 dB, with the next-largest contributor (longitude) at 1.97 dB (17.1%). This contrast demonstrates that, even among locations with similar NDVI increases, feature attributions vary substantially depending on other environmental factors—precisely the kind of location-specific diagnostic that enables network planners to identify which factors drive attenuation at each measurement point.
Among the 50 high-NDVI-increase locations, NDVI dominates in 62% of cases; across the entire June campaign, NDVI dominates in 46% of samples (1098 of 2392) while distance dominates in only 28% (679 of 2392). Since soil color alone does not affect path loss, the significant NDVI contribution at these locations confirms that the seasonal NDVI increase here reflects genuine vegetation growth—justifying the high-vegetation interpretation for this subset.

7. Conclusions

This study presents a comprehensive measurement campaign investigating 5G propagation characteristics in the 3.75 GHz band, with particular emphasis on seasonal vegetation effects in rural environments. Through systematic RSRP measurements conducted across three distinct periods (April, May, and June 2025), we captured the complete seasonal transition from minimal to advanced vegetation density.
Our machine learning-based path loss prediction analysis yielded several key findings. First, we confirmed that Random Forest and XGBoost models achieve superior prediction accuracy ( R 2 up to 0.891 on combined data; up to 0.907 with antenna pattern-corrected path loss), while interpretable glass-box models maintain competitive performance with the added benefit of explainability, with EBM achieving R 2 up to 0.876 on combined data. Second, our analysis using both empirical path loss models as well as ML models quantified the significant impact of vegetation density variations on path loss. Using EBM’s additive structure, we decomposed the +7.39 dB path loss increase observed from April to June across 629 locations matched across the April and June campaigns: NDVI accounts for 57.6% of this increase (+4.26 dB), while distance contributes virtually nothing ( 0.02 dB), confirming that the seasonal degradation is driven primarily by vegetation growth rather than measurement artifacts. This is consistent with the empirical LNS model analysis, where the small difference between shadowing variation σ and total path loss standard deviation similarly indicated that distance is not the dominant factor in path loss variability. Third, we demonstrated a nuanced role for the Normalized Difference Vegetation Index (NDVI) in path loss prediction. Within individual campaigns, NDVI provides no predictive benefit—distance dominates (Pearson r 0.81 ) and within-campaign vegetation differences are too subtle to affect propagation. However, when combining data across multiple seasonal campaigns, NDVI substantially improves performance ( R 2 improvements of 3–4 percentage points), as it distinguishes fundamentally different vegetation states. Glass-box models maintain near-parity with black-box approaches (EBM R 2 = 0.876 vs. RF R 2 = 0.891 ), demonstrating that interpretability need not come at a significant cost in accuracy. Importantly, our cross-campaign transfer analysis revealed that raw NDVI values can degrade model transferability between seasons, as the same absolute NDVI value may represent different vegetation states across campaigns (cf. Figure 3 and Figure 4). This degradation is model-dependent: XGBoost is substantially less susceptible to NDVI-induced cross-campaign degradation than EBM, which we attribute to tree-based feature interactions that can compensate for misleading NDVI values through splits on correlated geometric features, whereas EBM’s additive structure applies its NDVI shape function unconditionally. Notably, the April-to-June transition is the only cross-campaign direction where NDVI improves prediction accuracy ( R 2 increases by approximately 0.2 for EBM and 0.08 for XGBoost), and our EBM decomposition—which focused on this transition—accordingly identifies vegetation change as the dominant driver of seasonal path loss increase. Yet for other campaign directions, NDVI provides no benefit or even degrades performance, suggesting that NDVI-based cross-campaign transfer requires higher-resolution vegetation data. A hybrid LNS and XGBoost model—where XGBoost is trained on the residuals of a per-campaign log-normal shadowing fit—matches pure XGBoost on individual campaigns but improves combined-dataset accuracy when NDVI is absent, as the per-campaign path loss exponent acts as a seasonal proxy. When NDVI is available, this advantage disappears as follows: pure XGBoost with NDVI ( R 2 of 0.888) outperforms the hybrid model without NDVI ( R 2 of 0.864), confirming that satellite-derived vegetation indices capture seasonal variability more effectively than empirical path loss parameters alone.
For rural 5G deployments, these results imply that coverage predictions based on a single-season measurement may underestimate path loss by up to about 7–9 dB as vegetation density increases.
Since these results are derived from a single vineyard site at 3.75 GHz, future work should validate whether the methodological findings transfer to other rural sites, frequencies, and crop types—in particular, (i) the cross-campaign vs. within-campaign role of vegetation indices, (ii) the sensitivity to satellite resolution, and (iii) the suitability of EBM for propagation analysis besides site-specific quantities such as path loss coefficients and the order of lower-ranked features.
A second priority is improving cross-seasonal prediction capabilities. We attribute the transfer limitation to the available satellite resolution, at which NDVI conflates vine canopy with inter-row ground cover, whose relative contributions change non-monotonically due to vineyard management—so that NDVI may only transfer reliably between campaigns when all sub-pixel vegetation components increase consistently. Resolving this ambiguity will likely require switching from the freely available Sentinel-2 imagery (10 m resolution) to commercial satellite data (1–3 m) that can resolve individual vine rows, or complementing satellite data with ground-truth photography.

Author Contributions

Conceptualization, A.I.J., D.S., D.M., H.F., and M.A.W.; methodology, D.S., A.I.J., and H.F.; validation, D.S., A.I.J., and H.F.; investigation, D.S., A.I.J., D.M., H.F., and M.A.W.; resources, M.A.W. and H.F.; data curation, D.S. and D.M.; writing—original draft preparation, A.I.J. and D.S.; writing—review and editing, H.F., M.A.W., D.M., and D.S.; visualization, A.I.J., D.S., and D.M.; supervision, H.F. and M.A.W.; project administration, M.A.W.; funding acquisition, M.A.W. and H.F. All authors have read and agreed to the published version of the manuscript.

Funding

This research was conducted within the NoLa project, which was funded within the framework of the InnoNT program of the German Federal Ministry for Digital and State Modernization (BMDS), grant 19OI23015A.

Data Availability Statement

The dataset collected for this study is openly available on Zenodo at https://doi.org/10.5281/zenodo.19744782.

Acknowledgments

The authors would like to thank the members of the cooperating partners from aeroDCS GmbH, Dienstleistungszentrum Ländlicher Raum (Mosel), MRK Media AG, Plantivo GmbH, and the University of Applied Sciences Koblenz for the fruitful collaboration in the NoLa project, without which this research would not have been possible. The authors would also like to thank the anonymous reviewers for their constructive feedback, in particular for suggesting a hybrid approach, which we implemented as the LNS+XGBoost analysis in Section 6.2. The authors have reviewed and edited the output and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

Abbreviations

3GPP3rd Generation Partnership Project
ABGAlpha-Beta-Gamma (model)
ANFIS   Adaptive Neuro-Fuzzy Inference System
BBBlack-Box (model)
CIClose-In (model)
COSTEuropean Cooperation in Science and Technology
CUCentral Unit
DLDeep Learning
DSMDigital Surface Model
DTMDigital Terrain Model
DUDistributed Unit
EBMExplainable Boosting Machine
ECCElectronic Communications Committee
EIRPEffective Isotropic Radiated Power
FIFloating Intercept (model)
FNNFeedforward Neural Network
GAMGeneralized Additive Model
GBGlass-Box (model)
InHIndoor Hotspot
L1CLevel-1C (Sentinel-2 processing level)
LOSLine-Of-Sight
LNSLog-Normal Shadowing
LSTMLong Short-Term Memory
MAEMean Absolute Error
MAPEMean Absolute Percentage Error
MAPLMaximum Allowable Path Loss
MLMachine Learning
MLPMulti-Layer Perceptron
NDVINormalized Difference Vegetation Index
NIRNear Infrared
NLOSNon-Line-of-Sight
NRNew Radio (5G NR)
O-RANOpen Radio Access Network
PCCPearson Correlation Coefficient
PLPath Loss
PLEIRPPath Loss in EIRP formulation
PLPATPath Loss with Antenna Pattern Correction
PRISMAPreferred Reporting Items for Systematic Reviews and Meta-Analyses
PTPPrecision Time Protocol
QGISQuantum Geographic Information System
RANRadio Access Network
ReLURectified Linear Unit
RFRandom Forest
RMaRural Macro (3GPP model)
RMSERoot Mean Square Error
RNNRecurrent Neural Network
RSRPReference Signal Received Power
SFSmart Farming
SS-RSRPSynchronization Signal Reference Signal Received Power
SSBSynchronization Signal Block
SSSSecondary Synchronization Signal
SUIStanford University Interim
WGS84World Geodetic System 1984
XGBoostExtreme Gradient Boosting

References

  1. Sudhamani, C.; Roslee, M.; Chuan, L.L.; Waseem, A.; Osman, A.F.; Jusoh, M.H. Performance Analysis of a Millimeter Wave Communication System in Urban Micro, Urban Macro, and Rural Macro Environments. Energies 2023, 16, 5358. [Google Scholar] [CrossRef]
  2. Heikkilä, M.; Suomalainen, J.; Saukko, O.; Kippola, T.; Lähetkangas, K.; Koskela, P.; Kalliovaara, J.; Haapala, H.; Pirttiniemi, J.; Yastrebova, A.; et al. Unmanned Agricultural Tractors in Private Mobile Networks. Network 2022, 2, 1–20. [Google Scholar] [CrossRef]
  3. Lindenschmitt, D.; Veith, B.; Gundall, M.; Alam, K.; Daurembekova, A.; Habibi, M.A.; Han, B.; Krummacker, D.; Rosemann, P.; Schotten, H.D. Nomadic non-public networks for 6G: Use cases and key performance indicators. In Proceedings of the 2024 IEEE Conference on Standards for Communications and Networking (CSCN); IEEE: New York, NY, USA, 2024; pp. 243–248. [Google Scholar]
  4. Saeed, I.A.; Abdullah, A.; Schneider, D.; Reinelt, M.; Pannek, S.; Farnschlaeder, T.; Frey, H.; Kiess, W.; Wimmer, M.A. A Measurement Study on 5G Performance in Steep Vineyards. In Proceedings of the 2025 IEEE 26th International Symposium on a World of Wireless, Mobile and Multimedia Networks (WoWMoM); IEEE: New York, NY, USA, 2025; pp. 202–211. [Google Scholar] [CrossRef]
  5. Krause, A.; Anwar, W.; Martinez, A.B.; Stachorra, D.; Fettweis, G.; Franchi, N. Network Planning and Coverage Optimization for Mobile Campus Networks. In Proceedings of the 2021 IEEE 4th 5G World Forum, 5GWF; Institute of Electrical and Electronics Engineers Inc.: New York, NY, USA, 2021; pp. 305–310. [Google Scholar] [CrossRef]
  6. Chee, K.L.; Torrico, S.A.; Kürner, T. Foliage Attenuation Over Mixed Terrains in Rural Areas for Broadband Wireless Access at 3.5 GHz. IEEE Trans. Antennas Propag. 2011, 59, 2698–2706. [Google Scholar] [CrossRef]
  7. Egi, Y.; Eyceyurt, E. Classified 3D mapping and deep learning-aided signal power estimation architecture for the deployment of wireless communication systems. EURASIP J. Wirel. Commun. Netw. 2022, 2022, 107. [Google Scholar] [CrossRef]
  8. Rudin, C. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat. Mach. Intell. 2019, 1, 206–215. [Google Scholar] [CrossRef] [PubMed]
  9. Khalili, H.; Frey, H.; Wimmer, M.A. Balancing Prediction Accuracy and Explanation Power of Path Loss Modeling in a University Campus Environment via Explainable AI. Future Internet 2025, 17, 155. [Google Scholar] [CrossRef]
  10. Hayashi, T.; Ichige, K. A Deep-Learning Method for Path Loss Prediction Using Geospatial Information and Path Profiles. IEEE Trans. Antennas Propag. 2023, 71, 7523–7537. [Google Scholar] [CrossRef]
  11. Kayaalp, K.; Metlek, S.; Genc, A.; Dogan, H.; Basyigit, İ.B. Prediction of path loss in coastal and vegetative environments with deep learning at 5G sub-6 GHz. Wirel. Netw. 2023, 29, 2471–2480. [Google Scholar] [CrossRef]
  12. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [PubMed]
  13. Schumacher, A.; Merz, R.; Burg, A. 3.5 GHz Coverage Assessment with a 5G Testbed. In Proceedings of the 89th Vehicular Technology Conference (VTC2019-Spring); IEEE: New York, NY, USA, 2019. [Google Scholar]
  14. Qamar, F.; Abbas, T.; Hindia, M.N.; Dimyati, K.B.; Noordin, K.A.B.; Ahmed, I. Characterization of MIMO Propagation Channel at 15 GHz for the 5G Spectrum. In Proceedings of the 13th Malaysia International Conference on Communications; Institute of Electrical and Electronics Engineers (IEEE): New York, NY, USA, 2017; pp. 265–270. [Google Scholar]
  15. Kaitatzis, C.; Boursianis, A.; Goudos, S.K.; Dallas, P.I. A Preliminary Coverage Study in Millimeter Wave Bands for 5G Communication Networks. In Proceedings of the MOCAST: 6th International Conference on Modern Circuits and Systems Technologies; IEEE: New York, NY, USA, 2017; p. 326. [Google Scholar]
  16. Nordin, M.A.M.; Ramli, H.A.M. Performance analysis of 5G path loss models for rural macrocell environment. IIUM Eng. J. 2020, 21, 85–99. [Google Scholar] [CrossRef]
  17. Saba, N.; Mela, L.; Ruttik, K.; Salo, J.; Jantti, R. Millimeter-Wave Radio Link Analysis for 5G FWA by Combining Measurements and Geospatial Data. In Proceedings of the 2022 IEEE Future Networks World Forum (FNWF); Institute of Electrical and Electronics Engineers Inc.: New York, NY, USA, 2022; pp. 433–438. [Google Scholar] [CrossRef]
  18. Furnieles, C.J.; Castro, J.A.; Lopez, D.S.; Arevalo, J.E.; Araque, J.L. Path Loss Measurements in the 60 GHz Frequency Band in a Greenhouse. In Proceedings of the LACAP 2024—1st Latin American Conference on Antennas and Propagation, Conference Proceedings; Institute of Electrical and Electronics Engineers Inc.: New York, NY, USA, 2024. [Google Scholar] [CrossRef]
  19. Gupta, A.; Ghanshala, K.; Joshi, R.C. Path loss predictions for fringe areas using adaptive neuro-fuzzy inference system. Int. J. Syst. Assur. Eng. Manag. 2022, 13, 866–879. [Google Scholar] [CrossRef]
  20. Shaibu, F.E.; Onwuka, E.N.; Salawu, N.; Oyewobi, S.S. A Novel Hybrid Path Loss Prediction Model for 5G Midband Networks Using Empirical, Machine Learning, and Feature Prioritization Techniques. Int. J. Antennas Propag. 2025, 2025, 3277479. [Google Scholar] [CrossRef]
  21. Matsumoto, T.; Fujii, T. ML-Assisted Empirical Modeling with Blockage-Aware Attenuation for Sub-6 GHz Private 5G Using 3D City Maps and Satellite Imagery. In Proceedings of the 2026 International Conference on Artificial Intelligence in Information and Communication (ICAIIC), Tokyo, Japan, 24–27 February 2026. [Google Scholar] [CrossRef]
  22. ITU-R. Field-Strength Measurements Along a Route with Geographical Coordinate Registrations; Recommendation SM.1708-1; International Telecommunication Union: Geneva, Switzerland, 2011. [Google Scholar]
  23. Karney, C.F. Algorithms for geodesics. J. Geod. 2013, 87, 43–55. [Google Scholar] [CrossRef]
  24. 3GPP. Study on Channel Model for Frequencies from 0.5 to 100 GHz; Technical Report TR 38.901 V19.1.0, 3rd Generation Partnership Project; 3GPP: Sophia Antipolis, France, 2025; Release 19. [Google Scholar]
  25. Kyösti, P.; Meinilä, J.; Hentilä, L.; Zhao, X.; Jämsä, T.; Schneider, C.; Narandžić, M.; Milojević, M.; Hong, A.; Ylitalo, J.; et al. WINNER II Channel Models; Technical Report D1.1.2 V1.2, IST-4-027756; WINNER II: Munich, Germany, 2007; Updated February 2008. [Google Scholar]
  26. Sun, S.; Rappaport, T.S.; Thomas, T.A.; Ghosh, A.; Nguyen, H.C.; Kovács, I.Z.; Rodriguez, I.; Koymen, O.; Partyka, A. Investigation of Prediction Accuracy, Sensitivity, and Parameter Stability of Large-Scale Propagation Path Loss Models for 5G Wireless Communications. IEEE Trans. Veh. Technol. 2016, 65, 2843–2860. [Google Scholar] [CrossRef]
  27. Okumura, Y.; Ohmori, E.; Kawano, T.; Fukuda, K. Field Strength and Its Variability in VHF and UHF Land-Mobile Radio Service. Rev. Electr. Commun. Lab. 1968, 16, 825–873. [Google Scholar]
  28. Hata, M. Empirical Formula for Propagation Loss in Land Mobile Radio Services. IEEE Trans. Veh. Technol. 1980, VT-29, 317–325. [Google Scholar] [CrossRef]
  29. COST Action 231. Digital Mobile Radio Towards Future Generation Systems; Final Report; Technical Report EUR 18957; Publications Office of the European Union: Luxembourg, 1999; ISBN 92-828-5416-7. [Google Scholar]
  30. Rappaport, T.S. Wireless Communications: Principles and Practice, 2nd ed.; Cambridge University Press: Cambridge, UK, 2024. [Google Scholar] [CrossRef]
  31. Petrita, T.; Ignea, A. A new method for interpolation of 3D antenna pattern from 2D plane patterns. In Proceedings of the 2012 10th International Symposium on Electronics and Telecommunications; IEEE: New York, NY, USA, 2012; pp. 393–396. [Google Scholar] [CrossRef]
  32. Breiman, L. Random forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef]
  33. Chen, T.; Guestrin, C. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd ACM Sigkdd International Conference on Knowledge Discovery and Data Mining; ACM: New York, NY, USA, 2016; pp. 785–794. [Google Scholar]
  34. Singh, H.; Gupta, S.; Dhawan, C.; Mishra, A. Path loss prediction in smart campus environment: Machine learning-based approaches. In Proceedings of the 2020 IEEE 91st Vehicular Technology Conference (VTC2020-Spring); IEEE: New York, NY, USA, 2020; pp. 1–5. [Google Scholar]
  35. Molnar, C. Interpretable Machine Learning: A Guide for Making Black Box Models Explainable; Leanpub: Victoria, BC, Canada, 2022. [Google Scholar]
  36. Nori, H.; Jenkins, S.; Koch, P.; Caruana, R. Interpretml: A unified framework for machine learning interpretability. arXiv 2019, arXiv:1909.09223. [Google Scholar] [CrossRef]
Figure 1. Architecture of the nomadic platform. Arrows denote physical/logical connections between components and labels indicate the interface or protocol type.
Figure 1. Architecture of the nomadic platform. Arrows denote physical/logical connections between components and labels indicate the interface or protocol type.
Futureinternet 18 00237 g001
Figure 2. Nomadic platform deployed in Arena.
Figure 2. Nomadic platform deployed in Arena.
Futureinternet 18 00237 g002
Figure 3. Seasonal changes in vineyard vegetation across measurement campaigns: (a) April showing minimal vegetation, (b) May during growing season, and (c) June at advanced vegetation density.
Figure 3. Seasonal changes in vineyard vegetation across measurement campaigns: (a) April showing minimal vegetation, (b) May during growing season, and (c) June at advanced vegetation density.
Futureinternet 18 00237 g003
Figure 4. NDVI values (L1C) for measurement campaigns: April, May, and June 2025. The plots cover the Arena vineyard and adjacent terrain. The dashed boxes bound the measurement data.
Figure 4. NDVI values (L1C) for measurement campaigns: April, May, and June 2025. The plots cover the Arena vineyard and adjacent terrain. The dashed boxes bound the measurement data.
Futureinternet 18 00237 g004
Figure 5. Routes of path loss measurement campaigns: April, May, and June 2025. The maps have been generated using OpenStreetMap and show measurement points and RSRP values across the Arena vineyard as well as the antenna location (red triangle).
Figure 5. Routes of path loss measurement campaigns: April, May, and June 2025. The maps have been generated using OpenStreetMap and show measurement points and RSRP values across the Arena vineyard as well as the antenna location (red triangle).
Futureinternet 18 00237 g005
Figure 6. Path loss P L P A T calculated from empirical RSRP measurements and antenna pattern via (1) as well as LNS model predictions for April, May, and June 2025. The increasing path loss exponent ( α = 2.27 , 3.28 , 4.23 ) reflects progressively denser vegetation across the seasonal campaigns.
Figure 6. Path loss P L P A T calculated from empirical RSRP measurements and antenna pattern via (1) as well as LNS model predictions for April, May, and June 2025. The increasing path loss exponent ( α = 2.27 , 3.28 , 4.23 ) reflects progressively denser vegetation across the seasonal campaigns.
Futureinternet 18 00237 g006
Figure 7. Global feature importance for EBMs trained (a) without NDVI and (b) with NDVI. Bars show mean absolute contribution in dB; percentages indicate each feature’s share of the total feature contribution (excluding intercept).
Figure 7. Global feature importance for EBMs trained (a) without NDVI and (b) with NDVI. Bars show mean absolute contribution in dB; percentages indicate each feature’s share of the total feature contribution (excluding intercept).
Futureinternet 18 00237 g007
Figure 8. Marginal effect plots for (a) distance, (b) elevation, (c) latitude, and (d) NDVI. The y-axis score represents each feature’s additive contribution to predicted path loss in dB.
Figure 8. Marginal effect plots for (a) distance, (b) elevation, (c) latitude, and (d) NDVI. The y-axis score represents each feature’s additive contribution to predicted path loss in dB.
Futureinternet 18 00237 g008
Figure 9. EBM decomposition of the +7.39 dB path loss increase from April to June. Analysis based on 629 matched location pairs (42% of April locations; mean spatial matching distance: 0.67 m). Bars show absolute contribution in dB; percentages indicate each feature’s share of the total path loss change. (a) EBM trained without NDVI attributes the increase primarily to elevation and its interaction terms. (b) EBM trained with NDVI included reveals vegetation change as the dominant factor (57.6%), concentrating the explanation on the actual driver of seasonal path loss variation.
Figure 9. EBM decomposition of the +7.39 dB path loss increase from April to June. Analysis based on 629 matched location pairs (42% of April locations; mean spatial matching distance: 0.67 m). Bars show absolute contribution in dB; percentages indicate each feature’s share of the total path loss change. (a) EBM trained without NDVI attributes the increase primarily to elevation and its interaction terms. (b) EBM trained with NDVI included reveals vegetation change as the dominant factor (57.6%), concentrating the explanation on the actual driver of seasonal path loss variation.
Futureinternet 18 00237 g009
Figure 10. EBM aggregate and local explanations for high-NDVI-increase locations. (a) Aggregate view: mean absolute contributions averaged over the 50 locations with greatest NDVI increase from April to June (mean NDVI: 0.34 to 0.71). (b) Local view: two individual predictions illustrating the range of NDVI’s role. Both share the same intercept (145.62 dB). Top: smallest NDVI effect (7.2% of total; predicted PL: 139.20 dB). Bottom: largest NDVI effect (40.9%; predicted PL: 149.60 dB). Signed dB values indicate direction (positive = increases path loss).
Figure 10. EBM aggregate and local explanations for high-NDVI-increase locations. (a) Aggregate view: mean absolute contributions averaged over the 50 locations with greatest NDVI increase from April to June (mean NDVI: 0.34 to 0.71). (b) Local view: two individual predictions illustrating the range of NDVI’s role. Both share the same intercept (145.62 dB). Top: smallest NDVI effect (7.2% of total; predicted PL: 139.20 dB). Bottom: largest NDVI effect (40.9%; predicted PL: 149.60 dB). Signed dB values indicate direction (positive = increases path loss).
Futureinternet 18 00237 g010
Table 2. Advanced models.
Table 2. Advanced models.
PaperFrequency & ScenarioEnvironment & LocationModelKey Findings
[10]800 MHz, 2 GHz; urban, suburban, ruralBuildings and terrain; rural Tokyo (Japan)DCNN + FNNEstimation accuracy improves with building occupancy images; highest accuracy in rural scenario
[11]3.5, 3.8, 4.2 GHz; Urban, RuralCoastal terrains and vegetation areas; Isparta and Burdur (Turkey)LSTM, RNNRNN predicts better than LSTM; Path loss higher for coastal than vegetation areas
[7]850 MHz; RuralIrregular terrains, buildings, and vegetation; Melbourne, FL (USA)DL and Computer Vision systemProposed SPPL model outperformed empirical models after 33 m distance
[19]900–2300 MHz; Urban, Suburban, RuralHills, valleys, roads; Dehradun, IndiaANFIS vs. empirical modelsANFIS achieves RMSE 11.20 vs. 82.50 (ECC-33 model)
Table 3. Configuration parameters for the measurement campaigns.
Table 3. Configuration parameters for the measurement campaigns.
SiteArena
BS location (latitude, longitude)49.912922150, 7.05047455
5G NR Frequency3.75 GHz
SSB Frequency3.748 GHz
Transmit (TX) power per antenna port37 ± 2.5 dBm
Number of antenna ports4
Transmit (TX) antenna gain (dBi)13.3 dBi (max)
Receive (RX) antenna gain (dBi)assuming 0 dBi
Bandwidth100 MHz
TX Height (m)4
RX Height (m)1.5
Table 4. Statistical summary of April measurement data.
Table 4. Statistical summary of April measurement data.
LatitudeLongitudeElevation (m)Distance (m)Clutter Height (m)NDVIPLEIRP (dB)
mean49.913257.05010247.12049.2755.5560.427141.725
median49.913247.05023247.00054.2903.5000.436143.300
stddev0.000260.000368.22327.4644.0380.0677.902
min49.912867.04941233.0000.215−0.9000.191116.250
max49.913707.05055262.00098.60612.9000.572159.000
Table 5. Statistical Summary of May Measurement Data.
Table 5. Statistical Summary of May Measurement Data.
LatitudeLongitudeElevation (m)Distance (m)Clutter Height (m)NDVIPLEIRP (dB)
mean49.913297.05007244.33952.2715.5240.296143.796
median49.913287.05006244.00053.9743.5000.292144.975
stddev0.000230.000317.66622.5834.3920.0538.577
min49.912857.04936229.0000.000−0.9000.192113.150
max49.913737.05057259.000103.79512.9000.555166.400
Table 6. Statistical summary of June measurement data.
Table 6. Statistical summary of June measurement data.
LatitudeLongitudeElevation (m)Distance (m)Clutter Height (m)NDVIPLEIRP (dB)
mean49.913317.05000244.86358.3405.8950.648150.854
median49.913337.05000245.00062.1124.7000.663153.150
stddev0.000250.000328.47522.6754.3670.0879.771
min49.912857.04931229.0000.603−0.9000.330116.500
max49.913747.05060259.000104.38112.9000.800169.600
Table 7. Correlation coefficient matrix for April measurement data.
Table 7. Correlation coefficient matrix for April measurement data.
LatitudeLongitudeElevation (m)Distance (m)Clutter Height (m)NDVIPLEIRP (dB)
Latitude1.000−0.0220.9540.7680.6850.2650.692
Longitude−0.0221.0000.255−0.6040.101−0.456−0.296
Elevation (m)0.9540.2551.0000.5720.6860.1110.555
Distance (m)0.768−0.6040.5721.0000.4740.4630.781
Clutter height (m)0.6850.1010.6860.4741.0000.1730.477
NDVI0.265−0.4560.1110.4630.1731.0000.286
Path Loss (dB)0.692−0.2960.5550.7810.4770.2861.000
Table 8. Correlation coefficient matrix for May measurement data.
Table 8. Correlation coefficient matrix for May measurement data.
LatitudeLongitudeElevation (m)Distance (m)Clutter Height (m)NDVIPLEIRP (dB)
Latitude1.0000.1180.9430.7640.7490.3950.535
Longitude0.1181.0000.395−0.4850.220−0.372−0.463
Elevation (m)0.9430.3951.0000.5490.7350.2480.346
Distance (m)0.764−0.4850.5491.0000.5130.6370.820
Clutter height (m)0.7490.2200.7350.5131.0000.3120.354
NDVI0.395−0.3720.2480.6370.3121.0000.525
Path Loss (dB)0.535−0.4630.3460.8200.3540.5251.000
Table 9. Correlation coefficient matrix for June measurement data.
Table 9. Correlation coefficient matrix for June measurement data.
LatitudeLongitudeElevation (m)Distance (m)Clutter Height (m)NDVIPLEIRP (dB)
Latitude1.0000.1730.9580.7460.7040.2120.514
Longitude0.1731.0000.423−0.4590.2440.156−0.497
Elevation (m)0.9580.4231.0000.5540.7230.2520.317
Distance (m)0.746−0.4590.5541.0000.467−0.0620.836
Clutter height (m)0.7040.2440.7230.4671.0000.3160.264
NDVI0.2120.1560.252−0.0620.3161.000−0.089
Path Loss (dB)0.514−0.4970.3170.8360.264−0.0891.000
Table 10. Monthly empirical path loss model parameters derived from measurements after antenna pattern correction. The path loss coefficient α characterizes signal attenuation with distance, while the shadowing variation σ represents environmental scattering effects. Under ideal model conditions, the expectation E [ X σ ] should equal zero; significant deviations indicate systematic model bias or insufficient model complexity.
Table 10. Monthly empirical path loss model parameters derived from measurements after antenna pattern correction. The path loss coefficient α characterizes signal attenuation with distance, while the shadowing variation σ represents environmental scattering effects. Under ideal model conditions, the expectation E [ X σ ] should equal zero; significant deviations indicate systematic model bias or insufficient model complexity.
MonthPath Loss Coefficient α Shadowing Variation σ (dB)Expected Shadowing E [ X σ ] (dB)
April2.277.210.64
May3.288.210.80
June4.238.970.77
Table 11. Path loss P L P A T (dB) by month. As opposed to P L E I R P in the other tables, here, the path loss includes the antenna pattern.
Table 11. Path loss P L P A T (dB) by month. As opposed to P L E I R P in the other tables, here, the path loss includes the antenna pattern.
AprilMayJune
mean141.71143.79150.63
median143.10144.67153.05
stddev7.518.5910.26
min119.99110.12111.97
max158.53166.20170.24
Table 12. Comprehensive model performance comparison (path loss prediction). Features: latitude, longitude, elevation, distance, and clutter height.
Table 12. Comprehensive model performance comparison (path loss prediction). Features: latitude, longitude, elevation, distance, and clutter height.
ModelAprilMayJuneCombined
R2MAE
(dB)
RMSE
(dB)
R2MAE
(dB)
RMSE
(dB)
R2MAE
(dB)
RMSE
(dB)
R2MAE
(dB)
RMSE
(dB)
RF0.8812.052.740.8912.232.900.9242.082.640.8522.933.78
XGBoost0.8902.002.630.8862.282.970.9212.092.680.8512.923.80
MLP0.8592.232.980.8772.363.090.9232.072.660.8453.003.88
EBM0.8832.042.720.8712.493.160.9192.112.710.8413.093.93
GAM0.8232.623.340.7963.223.970.8343.013.900.7663.874.76
Table 13. Comprehensive model performance comparison (path loss prediction) with NDVI. Features: latitude, longitude, elevation, distance, clutter height, and NDVI.
Table 13. Comprehensive model performance comparison (path loss prediction) with NDVI. Features: latitude, longitude, elevation, distance, clutter height, and NDVI.
ModelAprilMayJuneCombined
R2MAE
(dB)
RMSE
(dB)
R2MAE
(dB)
RMSE
(dB)
R2MAE
(dB)
RMSE
(dB)
R2MAE
(dB)
RMSE
(dB)
RF0.8822.052.730.8882.262.940.9222.072.670.8912.473.24
XGBoost0.8921.972.610.8832.303.010.9222.072.670.8882.453.29
MLP0.8552.353.030.8702.453.170.9142.212.810.8792.693.43
EBM0.8772.052.780.8722.483.140.9102.222.860.8762.683.46
GAM0.8222.623.350.7983.193.950.8532.833.660.7983.534.42
Table 14. Hybrid model comparison (path loss prediction) without NDVI, where the hybrid model is always a combination of LNS with XGBoost, i.e., XGBoost operates on the residual path loss after LNS subtraction. The P L P A T rows use the same hyperparameters tuned on P L E I R P and re-tuned (ret.) rows use hyperparameters optimized on P L P A T . Features: latitude, longitude, elevation, distance, and clutter height. The LNS model is ignorant of features other than distance.
Table 14. Hybrid model comparison (path loss prediction) without NDVI, where the hybrid model is always a combination of LNS with XGBoost, i.e., XGBoost operates on the residual path loss after LNS subtraction. The P L P A T rows use the same hyperparameters tuned on P L E I R P and re-tuned (ret.) rows use hyperparameters optimized on P L P A T . Features: latitude, longitude, elevation, distance, and clutter height. The LNS model is ignorant of features other than distance.
ModelAprilMayJuneCombined
R2MAE
(dB)
RMSE
(dB)
R2MAE
(dB)
RMSE
(dB)
R2MAE
(dB)
RMSE
(dB)
R2MAE
(dB)
RMSE
(dB)
Target: P L E I R P
LNS only0.5274.195.460.5454.215.930.6514.455.650.6594.365.75
XGBoost only0.8902.002.630.8862.282.970.9212.092.680.8512.923.80
Hybrid0.8852.032.690.8822.323.020.9212.082.690.8642.763.64
Target: P L P A T
LNS only0.5924.816.480.6214.216.160.7604.305.560.6954.365.99
XGBoost only0.9272.072.740.9082.343.030.9442.082.690.8752.943.83
Hybrid0.9232.102.810.9072.363.060.9402.132.770.8932.733.55
XGBoost only, ret.0.9272.052.740.9082.343.030.9442.082.670.8762.943.82
Hybrid, ret.0.9272.082.740.9072.363.060.9402.132.780.8922.723.56
Table 15. Hybrid model comparison (path loss prediction) with NDVI, where the hybrid model is always a combination of LNS with XGBoost, i.e., XGBoost operates on the residual path loss after LNS subtraction. The P L P A T rows use the same hyperparameters tuned on P L E I R P and re-tuned (ret.) rows use hyperparameters optimized on P L P A T . Features: latitude, longitude, elevation, distance, clutter height, and NDVI. The LNS model is ignorant of features other than distance, in particular of NDVI.
Table 15. Hybrid model comparison (path loss prediction) with NDVI, where the hybrid model is always a combination of LNS with XGBoost, i.e., XGBoost operates on the residual path loss after LNS subtraction. The P L P A T rows use the same hyperparameters tuned on P L E I R P and re-tuned (ret.) rows use hyperparameters optimized on P L P A T . Features: latitude, longitude, elevation, distance, clutter height, and NDVI. The LNS model is ignorant of features other than distance, in particular of NDVI.
ModelAprilMayJuneCombined
R2MAE
(dB)
RMSE
(dB)
R2MAE
(dB)
RMSE
(dB)
R2MAE
(dB)
RMSE
(dB)
R2MAE
(dB)
RMSE
(dB)
Target: P L E I R P
LNS only0.5274.195.460.5454.215.930.6514.455.650.6594.365.75
XGBoost only0.8921.972.610.8832.303.010.9222.072.670.8882.453.29
Hybrid0.8911.982.630.8822.313.020.9192.082.720.8832.513.36
Target: P L P A T
LNS only0.5924.816.480.6214.216.160.7604.305.560.6954.365.99
XGBoost only0.9282.032.720.9052.353.080.9452.072.670.9072.463.31
Hybrid0.9252.062.770.9062.363.070.9422.102.730.9072.503.31
XGBoost only, ret.0.9301.982.680.9052.353.080.9452.072.670.9072.463.31
Hybrid, ret.0.9272.062.740.9062.363.070.9422.102.730.9072.503.31
Table 16. EBM and XGBoost Cross-Campaign R 2 Performance (Without NDVI).
Table 16. EBM and XGBoost Cross-Campaign R 2 Performance (Without NDVI).
Training/TestingAprilMayJune
April EBM0.9410.5700.068
May EBM0.5990.8970.465
June EBM−0.1600.2910.955
April XGBoost0.9610.5740.116
May XGBoost0.5480.9440.476
June XGBoost−0.1270.2930.969
Table 17. EBM and XGBoost Cross-Campaign R 2 Performance (With NDVI).
Table 17. EBM and XGBoost Cross-Campaign R 2 Performance (With NDVI).
Training/TestingAprilMayJune
April EBM0.9430.3970.281
May EBM0.5520.907−0.523
June EBM−1.038−2.1510.954
April XGBoost0.9650.5530.197
May XGBoost0.5710.9440.467
June XGBoost−0.4130.0160.970
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Schneider, D.; Jehangiri, A.I.; Müller, D.; Frey, H.; Wimmer, M.A. Explaining Seasonal 5G Path Loss in a Vineyard: From Empirical Models to Interpretable Machine Learning. Future Internet 2026, 18, 237. https://doi.org/10.3390/fi18050237

AMA Style

Schneider D, Jehangiri AI, Müller D, Frey H, Wimmer MA. Explaining Seasonal 5G Path Loss in a Vineyard: From Empirical Models to Interpretable Machine Learning. Future Internet. 2026; 18(5):237. https://doi.org/10.3390/fi18050237

Chicago/Turabian Style

Schneider, Daniel, Ali Imran Jehangiri, Daniel Müller, Hannes Frey, and Maria Anna Wimmer. 2026. "Explaining Seasonal 5G Path Loss in a Vineyard: From Empirical Models to Interpretable Machine Learning" Future Internet 18, no. 5: 237. https://doi.org/10.3390/fi18050237

APA Style

Schneider, D., Jehangiri, A. I., Müller, D., Frey, H., & Wimmer, M. A. (2026). Explaining Seasonal 5G Path Loss in a Vineyard: From Empirical Models to Interpretable Machine Learning. Future Internet, 18(5), 237. https://doi.org/10.3390/fi18050237

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop