1. Introduction
Water hyacinth (
Eichhornia crassipes) ranks among the world’s most problematic invasive aquatic plants, causing severe ecological and socioeconomic impacts across tropical and subtropical freshwater systems [
1,
2,
3]. Native to the Amazon basin of South America, this free-floating macrophyte has colonized freshwater bodies across six continents, facilitated by its ornamental trade value, rapid vegetative reproduction, and broad environmental tolerance [
3,
4]. Under favourable nutrient conditions, water hyacinth populations exhibit exponential growth rates, with biomass doubling times of 6–18 days, enabling rapid transformation of open water into dense vegetative mats [
5,
6]. Water hyacinth invasion causes severe ecological impacts, including reduced light penetration, hypoxic conditions, and disrupted aquatic food webs [
7,
8], with socioeconomic costs in Africa alone exceeding USD 20–50 million annually through navigation obstruction, infrastructure interference, and reduced fisheries productivity [
9,
10].
Lake Victoria, the world’s second-largest freshwater lake by surface area (68,800 km
2) and largest tropical lake, has experienced recurrent water hyacinth infestations since the species was first documented in 1989 [
11,
12]. The lake supports over 40 million people across Kenya, Uganda, and Tanzania through artisanal fisheries, transportation, and domestic water supply [
12,
13]. Water hyacinth proliferation in Lake Victoria has been linked to progressive eutrophication driven by nutrient loading from agricultural runoff, urban wastewater discharge, and atmospheric deposition [
14,
15], creating conditions that favour explosive macrophyte growth. Winam Gulf, a shallow semi-enclosed embayment in northeastern Lake Victoria, has been identified as one of the most severely affected areas, with historical coverage exceeding 35% of gulf surface area during the 1997–1998 outbreak peak [
16]. Despite decades of mechanical harvesting and biological control efforts, water hyacinth remains a persistent management challenge [
17,
18].
Remote sensing has emerged as an essential tool for monitoring aquatic invasive species across large spatial extents with high temporal frequency [
19,
20,
21]. Satellite-based multispectral sensors can detect floating vegetation through characteristic spectral signatures, particularly the pronounced increase in near-infrared (NIR) reflectance caused by internal leaf scattering, contrasted with strong NIR absorption by water [
22,
23]. This spectral contrast enables discrimination of floating macrophytes from open water using vegetation indices such as the Normalized Difference Vegetation Index [
24], Normalized Difference Water Index [
25], and Floating Algae Index [
26]. Numerous studies have applied remote sensing to map water hyacinth in African lakes [
27,
28,
29,
30,
31,
32,
33,
34,
35]. Fusilli et al. [
36] used MODIS imagery to track water hyacinth dynamics in Lake Victoria from 2000 to 2010, documenting seasonal expansion-contraction cycles. Dube et al. [
27] achieved classification accuracies exceeding 90% for water hyacinth in South African impoundments using Landsat 8 and machine learning algorithms, while Thamaga and Dube [
28] demonstrated the value of Sentinel-2 red edge bands for discriminating water hyacinth from other aquatic vegetation.
Among available satellite platforms, Sentinel-2 is particularly well-suited for monitoring dynamic aquatic macrophyte populations owing to its 10 m spatial resolution, red-edge spectral bands, and 5-day revisit capability [
37,
38,
39,
40]. Machine learning classifiers, especially Support Vector Machines, have shown superior performance for aquatic vegetation mapping in optically complex inland waters [
27,
28], and polygon-mean feature extraction approaches have been shown to reduce spectral noise relative to pixel-based methods [
21,
41].
While extensive research has focused on mapping water hyacinth extent and temporal dynamics, quantitative understanding of spatial distribution patterns within affected water bodies remains limited [
9,
18]. Most studies treat water hyacinth distribution as spatially homogeneous within mapped extents [
27,
36], yet field observations consistently indicate heterogeneous distributions with concentrations in sheltered bays, along lee shorelines, and near nutrient sources [
9,
16]. The relationship between floating vegetation distribution and shoreline proximity has been qualitatively documented but not rigorously quantified. Kateregga and Sterner [
9] noted that water hyacinth in Lake Victoria “tends to accumulate along shorelines and in protected bays,” attributing this pattern to wind-driven transport and reduced wave exposure. Albright et al. [
16] observed that persistent infestations in Winam Gulf occurred “within 2–3 km of shore.” However, these observations lack quantitative precision needed to inform spatial management strategies.
Wind-driven wave action is a primary mechanism governing the spatial distribution of floating vegetation in large water bodies [
42,
43]. Surface waves generate net mass transport in the direction of wave propagation through Stokes drift, which can substantially affect floating objects [
44,
45]. In fetch-limited shallow water environments, wave energy decays exponentially with distance from shore due to bottom friction and geometric spreading [
46,
47]. This process creates distinct hydrodynamic zones: high-energy offshore areas where vegetation disperses, transition zones where accumulation begins, and low-energy nearshore zones where stable mat formation occurs [
42,
48]. If floating vegetation distribution is governed primarily by wave-driven transport, occurrence probability should decline exponentially with shoreline distance, following wave energy attenuation patterns [
43,
46]. However, alternative functional forms such as power law relationships may emerge if bathymetric gradients are non-linear or if secondary accumulation zones exist at intermediate distances [
49,
50]. Testing these alternative models empirically requires spatially explicit quantitative analysis.
Despite extensive water hyacinth mapping research, quantitative models of shoreline preference remain absent for Lake Victoria. Understanding spatial distribution patterns has practical significance: mechanical harvesting operations are typically limited to within 2–5 km of shore-based facilities due to operational constraints [
18], yet the proportion of biomass within accessible zones remains unquantified. Furthermore, seasonal variability in wind patterns may modulate shoreline preference, but temporal dynamics of spatial distribution have not been systematically analysed [
51,
52]. This study aims to address these knowledge gaps by quantifying the spatial relationship between water hyacinth distribution and shoreline proximity in Winam Gulf, Lake Victoria. The specific objectives are: (1) to clarify the spatial distribution pattern of water hyacinth and quantify its seasonal variability using Sentinel-2 multispectral imagery; (2) to establish and compare quantitative models of the decay in shoreline preference with distance, tested against multiple functional forms; and (3) to translate spatial findings into evidence-based recommendations for prioritising mechanical harvesting zones.
This study provides the first quantitative estimate of shoreline preference magnitude for water hyacinth in Lake Victoria, develops theoretically motivated spatial decay models testable across sites and time periods, demonstrates a rigorous methodological framework integrating high-resolution satellite imagery with machine learning classification and multi-model spatial analysis, and translates findings into actionable management guidance based on quantitative evidence. The approach is transferable to other aquatic invasive species and water bodies globally.
2. Materials and Methods
2.1. Study Area
Winam Gulf (also known as Nyanza Gulf or Kavirondo Gulf) is a shallow, semi-enclosed embayment located in the northeastern corner of Lake Victoria, Kenya (approximately 0°06′ S to 0°32′ S, 34°13′ E to 34°52′ E). The gulf extends approximately 74 km in length and 50 km in width, with a total surface area of approximately 1493 km
2 and a mean depth of 4–6 m [
51]. The gulf connects to the main body of Lake Victoria through a narrow opening (~3 km wide) near Rusinga Island, which restricts water exchange and creates a semi-enclosed system conducive to water hyacinth accumulation.
Winam Gulf receives substantial nutrient inputs from the Nyando, Sondu-Miriu, and Kibos rivers, as well as urban and agricultural runoff from Kisumu City and surrounding catchments. These nutrient loadings, combined with the gulf’s sheltered morphology, shallow bathymetry, and warm tropical climate (mean annual temperature: 23–25 °C), create optimal conditions for water hyacinth proliferation. The gulf has experienced recurrent water hyacinth outbreaks since the 1990s, making it one of the most severely affected areas in Lake Victoria [
16,
36].
The gulf experiences distinct seasonal hydrodynamic regimes. During the dry season (June–October), persistent southeasterly trade winds generate surface waves that propagate toward the northern and western shorelines, potentially driving shoreward transport of floating vegetation. During the wet season (March–May), variable wind directions, elevated rainfall, higher lake levels, and increased riverine discharge create more complex circulation patterns. This seasonal contrast motivated the multi-temporal design of this study (
Figure 1).
2.2. Satellite Data Acquisition and Image Selection
Three cloud-free Sentinel-2 images were selected in advance to capture contrasting seasonal hydrodynamic conditions in Winam Gulf. The first image, from 31 May 2020, represents the wet-dry seasonal transition, a period marked by variable wind directions, elevated rainfall, and increased riverine discharge. The second, from 8 October 2020, corresponds to the early dry season, reflecting the onset of persistent southeasterly trade winds. The third, from 23 October 2020, represents the peak dry season, when southeasterly wind forcing is sustained and dominant.
Date selection was based on: (1) cloud cover <5% over the study area, (2) representation of distinct seasonal hydrodynamic regimes documented in prior Lake Victoria studies [
51], and (3) availability of high-quality Level-2A atmospherically corrected imagery [
53]. The October dates were specifically chosen to capture potential temporal variation within the dry season, as previous studies have documented intra-seasonal variability in water hyacinth distribution [
16].
2.3. Training Data Collection
Training polygons for supervised classification were manually digitised through visual interpretation of the Sentinel-2 imagery using a false-colour composite (NIR-Red-Green) to enhance vegetation detection. Two classes were defined: water hyacinth (H) and open water (W). Polygons were delineated based on spectral characteristics, with water hyacinth appearing as bright magenta-red patches due to high NIR reflectance from chlorophyll-rich floating vegetation, and open water appearing dark due to strong absorption in NIR wavelengths. For the peak dry season image, a total of 246 training polygons were digitised, comprising 174 water hyacinth polygons (70.7%) and 72 open water polygons (29.3%). Training samples were distributed across the entire gulf to capture spectral variability associated with different water hyacinth mat densities, water depths, and turbidity conditions. Polygon sizes ranged from approximately 500 m2 to 50,000 m2, ensuring representation of both small, fragmented patches and large contiguous mats.
2.4. Spectral Characterisation and Separability Analysis
Prior to classification, spectral separability between water hyacinth and open water was quantified to assess the feasibility of accurate discrimination and to inform classification approach selection. Mean reflectance values were extracted for each training polygon across all ten spectral bands, yielding a representative spectral signature for each training sample.
Spectral separability was quantified using the Jeffries–Matusita (JM) distance [
54], a widely used metric in remote sensing that ranges from 0 (complete spectral overlap) to 2 (complete separation). For each band, the JM distance was calculated as follows:
where B is the Bhattacharyya distance:
where μ
1 and μ
2 are the class means, σ
1 and σ
2 are the class standard deviations, and σ
2 = (σ
12 + σ
22)/2 is the pooled variance. Following Richards and Jia [
55], JM values exceeding 1.9 indicate excellent separability, values between 1.7 and 1.9 suggest good separability, and values below 1.0 indicate poor separability.
To provide independent validation of the SVM classification, three established vegetation indices were calculated for the classified image: (1) Normalized Difference Vegetation Index (NDVI = (NIR − Red)/(NIR + Red) [
24]), (2) Normalized Difference Water Index (NDWI = (Green − NIR)/(Green + NIR) [
25]), and (3) Floating Algae Index (FAI = NIR − (Red + (SWIR1 − Red) × (λ
NIR − λ
Red)/(λ
SWIR1 − λ
Red)) [
26]). In the FAI formula, λ
NIR, λ
Red, and λ
SWIR1 denote the central wavelengths of the NIR (B8, 842 nm), Red (B4, 665 nm), and SWIR1 (B11, 1610 nm) bands, respectively. Mean values of each index were compared between SVM-classified water hyacinth and open water pixels to confirm alignment with expected patterns documented in the literature.
2.5. Image Classification
A Support Vector Machine (SVM) classifier with a radial basis function (RBF) kernel was employed for supervised classification. SVM was selected for its demonstrated effectiveness in classifying aquatic vegetation in optically complex inland waters [
27,
56,
57,
58] and its robustness to high-dimensional feature spaces; although forest classifiers have also shown strong performance in remote sensing applications [
59]. To mitigate overfitting and reduce noise inherent in pixel-level classification, a polygon-mean approach was adopted [
41,
60]. For each training polygon, the mean spectral signature was calculated by averaging reflectance values across all pixels within the polygon for each of the ten spectral bands.
Prior to model training, features were standardised using z-score normalisation (mean = 0, standard deviation = 1) to ensure equal contribution of all spectral bands to the classification decision. The classification decision was therefore made at the polygon level (using mean spectral signatures), not on individual pixels; the resulting polygon-level labels were then projected spatially to produce the final binary raster. The SVM classifier was configured with the following parameters: C = 1.0 (regularisation parameter), gamma = ‘scale’ (kernel coefficient set to 1/(n_features × variance)), and balanced class weights (class_weight = ‘balanced’) to explicitly counteract the training polygon imbalance between hyacinth (n = 174) and water (n = 72) classes; this weighting increases the penalty for misclassifying minority-class samples, preventing the classifier from simply predicting the majority class.
The training dataset was partitioned using stratified random sampling, with 70% (
n = 172 polygons) allocated for model training and 30% (
n = 74 polygons) reserved for independent accuracy assessment. Model performance was further evaluated using 5-fold cross-validation on the training subset to assess generalisation ability and detect potential overfitting. Classification accuracy was assessed [
61,
62] using the following metrics: (1) overall accuracy (OA), defined as the proportion of correctly classified test samples; (2) Cohen’s kappa coefficient (κ), which accounts for chance agreement; (3) precision (user’s accuracy for water hyacinth); (4) recall (producer’s accuracy for water hyacinth); and (5) F1-score, the harmonic mean of precision and recall. A confusion matrix was generated to quantify commission and omission errors for each class.
Following validation, the trained SVM model was applied to the entire Sentinel-2 image to generate a binary classification map. Classification was performed using a moving window approach with 1000 × 1000-pixel tiles to manage computational memory requirements.
2.6. Shoreline Distance Calculation
Shoreline distance was calculated as the Euclidean distance from each pixel centroid to the nearest point on the Winam Gulf shoreline boundary. The shoreline boundary was extracted from a digitised vector polygon of the study area in the projected coordinate system UTM Zone 36S [
63] to ensure distance measurements in metres.
For each pixel location within the classified water area, the minimum distance to the shoreline was computed using the following procedure: (1) pixel centroids were derived from the raster coordinate system using affine transformation parameters; (2) shoreline vertices were extracted from the boundary polygon (40,933 coordinate points); and (3) the minimum Euclidean distance between each pixel centroid and all shoreline vertices was calculated using the cdist function from the SciPy spatial distance module [
64]. This process generated a continuous distance-to-shore raster with values ranging from 0 m (at shoreline) to a maximum of approximately 24,200 m at the gulf centre.
2.7. Shoreline Preference Analysis
2.7.1. Descriptive Statistics
Distance-to-shore values were extracted separately for pixels classified as water hyacinth and open water. Descriptive statistics [
65] including mean, median, standard deviation, 25th percentile (Q1), and 75th percentile (Q3) were calculated for each class to characterise their respective distance distributions.
Shoreline preference strength was quantified as the difference between the mean distance of open water pixels and the mean distance of water hyacinth pixels:
where
μwater and
μhyacinth are the mean distances for open water and water hyacinth, respectively. Positive values indicate that water hyacinth occurs closer to the shoreline than would be expected under a spatially random distribution.
2.7.2. Bootstrap Confidence Intervals
To quantify uncertainty in the preference strength estimate, 95% confidence intervals were calculated using a bootstrap resampling procedure [
66]. A total of 1000 bootstrap iterations were performed, with each iteration randomly sampling (with replacement) 10,000 pixels from each class to ensure computational feasibility while maintaining statistical representativeness. The preference strength was recalculated for each bootstrap sample, and the 2.5th and 97.5th percentiles of the resulting distribution defined the 95% confidence interval bounds.
2.7.3. Shoreline Buffer Zone Analysis
Five concentric buffer zones were delineated from the shoreline at distances of 100 m, 250 m, 500 m, 1000 m, and 2000 m to examine the spatial distribution of water hyacinth relative to shoreline proximity. These distances were selected to capture near-shore accumulation patterns (100–500 m), intermediate zones (1000 m), and the transition to open water (2000 m).
For each buffer zone, the following metrics were calculated:
Proportion within buffer: The percentage of total water hyacinth pixels occurring within each buffer distance from shore.
Odds ratio (OR): The odds of water hyacinth presence within the buffer relative to outside the buffer, compared to the equivalent odds for open water:
where
Hin and
Hout represent water hyacinth pixel counts inside and outside the buffer, respectively, and
Win and
Wout represent open water pixel counts. OR > 1 indicates preferential association with the near-shore zone.
95% confidence intervals for odds ratios: Calculated using the standard error of the log odds ratio under the assumption of independent observations:
2.7.4. Statistical Testing with Multiple Comparisons Correction
The Mann–Whitney U test [
67] was employed to assess whether the distance distributions of water hyacinth and open water pixels differed significantly. This non-parametric test [
68] was selected because distance data exhibited right-skewed distributions inconsistent with normality assumptions. Chi-square tests of independence (χ
2) [
69] were conducted for each buffer zone to test the null hypothesis that water hyacinth and open water were distributed independently of shoreline proximity. To account for multiple hypothesis testing across five buffer zones, the Bonferroni correction [
70] was applied, adjusting the significance threshold from α = 0.05 to α = 0.01 (0.05/5 tests).
Effect size was quantified using Cohen’s d, calculated as the standardised mean difference:
where
σpooled = √((
σ2water +
σ2hyacinth)/2) is the pooled standard deviation.
2.7.5. Spatial Decay Modelling
To characterise the relationship between odds ratio and buffer distance, four candidate models were fitted and compared:
where OR(d) is the odds ratio at distance d (metres), and a and b are model parameters. The offset of +1 in the power law and logarithmic models prevents undefined values at d = 0. Model parameters were estimated using non-linear least squares regression via the Levenberg–Marquardt algorithm [
71] implemented in SciPy’s curve_fit function [
64]. Goodness-of-fit was assessed using the coefficient of determination (R
2):
Model selection was based on Akaike Information Criterion [
72,
73]:
where
n is the number of observations (buffer zones), and k is the number of model parameters. Lower AIC values indicate better model fit relative to complexity. For the exponential model, the half-distance (d
1/2), defined as the distance at which the odds ratio declines to half of its initial value, was calculated as follows:
2.8. Temporal Validation
To assess the temporal robustness of observed spatial patterns, the complete classification and shoreline preference analysis pipeline was applied independently to two additional Sentinel-2 images: 31 May 2020 (wet-dry transition) and 8 October 2020 (early dry season). Each date was processed with its own independently digitised training polygons to avoid cross-date contamination.
Consistency was evaluated by comparing preference strength, odds ratios, model parameters, and statistical significance across dates. The coefficient of variation (CV) was calculated to quantify temporal variability in preference strength:
where σ is the standard deviation and μ is the mean of preference strength across dates.
2.9. Software and Reproducibility
All analyses were conducted using Python 3.9 [
74] with the following libraries: Rasterio 1.3 [
75] for raster input/output and spatial operations, GeoPandas 0.12 [
76] for vector processing and coordinate transformations, Scikit-learn 1.2 [
77] for machine learning classification and cross-validation, SciPy 1.10 [
64] for statistical tests and non-linear optimisation, NumPy 1.24 [
78] for numerical array operations, Matplotlib 3.6 [
79] and Seaborn 0.12 [
80] for data visualisation.