1. Introduction
One of the greatest tasks facing our world today is to harmonize financial advancement, equitable societal participation, and the preservation of natural ecosystems in order to ensure long-term global progress [
1,
2,
3,
4]. In 2015, the Member Countries of the United Nations approved the 2030 Agenda, which identifies universal access to safe water as one of its Sustainable Development Goals [
5]. Across the world, over 2.5 billion individuals rely each day on subsurface freshwater reserves, widely recognized as a reliable and uncontaminated supply for human consumption [
6,
7]. On a global scale, the expansion of cities and densely populated areas has severely undermined the condition and purity of underground water resources [
8,
9]. The expansion of urban areas often stimulates financial development and enhances overall quality of life; nevertheless, it simultaneously generates multiple contamination sources, including industrial discharges, household wastewater, and leaking sewer infrastructure [
10,
11], all of which pose serious risks to subsurface water systems. Although numerous investigations have documented the adverse consequences associated with urban growth [
8,
12,
13], the measurable and statistically defined linkage between urban expansion processes and the condition of groundwater resources remains insufficiently clarified. Accordingly, clarifying this interconnection is crucial, particularly with regard to shifts in land use and land cover driven by urban growth and broader anthropogenic pressures, in order to equip policymakers with solid scientific foundations for the sustainable oversight and safeguarding of groundwater resources.
Investigations examining how variations in land utilization and surface cover patterns influence the condition of subsurface water systems, particularly under intensifying human-induced stresses, trace their origins to 1984, when pioneering work was launched by the United States Geological Survey [
14]. From that point forward, researchers have applied a wide range of methodological approaches to investigate this interaction. Conventional analytical tools, including statistical correlation measures and hypothesis-based significance assessments, primarily identify associative patterns or levels of statistical relevance, yet they exhibit substantial methodological constraints [
15,
16,
17]. In pursuit of greater analytical rigor and enhanced predictive reliability, researchers have introduced alternative quantitative frameworks, including geographically weighted regression (GWR), structural equation modelling (SEM), and random forest (RF) algorithms [
18,
19,
20]. Despite these advancements, geographically weighted regression (GWR) and structural equation modeling (SEM) face constraints related to spatial sampling density and model specification, whereas random forest (RF) approaches are primarily confined to patterns detectable within the training dataset and may struggle with extrapolating beyond observed trends [
16,
18,
20]. Moreover, the uneven spatial distribution of groundwater monitoring points limits the accuracy of traditional models. To overcome this, the study introduces a novel quantitative framework that applies spatial grids to maps of human-driven land changes, enhancing both the precision and efficiency of detecting variations in groundwater quality. In recent years, machine learning (ML) techniques have transformed environmental predictive modeling, offering more robust and data-driven insights under the influence of anthropogenic pressures [
21,
22,
23]. Machine learning offers advanced tools to interpret intricate, nonlinear datasets, revealing hidden patterns and relationships that traditional statistical techniques often fail to detect [
24]. Nonetheless, the use of machine learning to examine human-driven land changes at a fine spatial grid scale has received limited attention. Consequently, this study emphasizes combining a grid-based framework with ML algorithms to generate more robust and dependable insights.
A transformative phase in the safeguarding of subsurface water systems has emerged through the adoption of machine learning frameworks, with artificial intelligence technologies playing a central and pioneering role in advancing groundwater resource protection [
21,
22,
24]. For machine learning applications, the attainment of accurate and dependable predictive performance fundamentally depends on the availability of comprehensive, well-structured, and reliable data for model training [
21,
22]. Accordingly, developing reliable training datasets requires a rigorous and accurate assessment of groundwater conditions. Conventional evaluation approaches generally rely either on expert-based qualitative judgments or exclusively on standardized quantitative indicators, rather than integrating both perspectives within a unified framework [
16,
17,
18,
19,
20]. Approaches grounded in professional judgment draw heavily on individual expertise; however, they are intrinsically susceptible to bias because specialists often assign unequal importance to influencing variables [
23,
24]. Moreover, these techniques can fail to detect latent patterns and subtle relationships embedded within groundwater data records [
25,
26]. Methods based exclusively on computational or mathematical formulations can disregard the interpretative knowledge of specialists, which may ultimately compromise the reliability and validity of the resulting assessments [
26]. To address these shortcomings, contemporary approaches increasingly combine expert-driven evaluations with quantitative analytical frameworks. Recent advances in Geo-AI ensemble modeling have enabled more robust and precise assessments of groundwater contamination under complex anthropogenic pressures. In particular, Graph Neural Networks (GNNs) capture spatial dependencies among monitoring wells [
27,
28,
29], Light Gradient Boosting Machine (LightGBM) models identify nonlinear interactions between environmental predictors and contaminant concentrations [
23,
30,
31], and Deep Long Short-Term Memory Networks (LSTM) account for temporal trends in groundwater quality over extended periods [
32,
33,
34,
35]. Building on these developments, this study develops a Composite Contamination Index (CCI), which integrates outputs from GNN, LightGBM, and LSTM within an optimization framework to reconcile expert-driven and data-driven weightings, thereby maximizing the combined informational value of spatial, temporal, and predictor-contaminant relationships for a comprehensive assessment of aquifer degradation.
Understanding the influence of human-driven land and resource pressures on groundwater quality is critical for safeguarding aquifer systems globally. Despite its importance, efforts to develop frameworks that capture these relationships at a fine, point-by-point spatial resolution remain scarce. To address this limitation, we propose a novel approach that integrates spatial grid sampling with advanced machine learning models, enabling a detailed and data-driven analysis of how anthropogenic pressures affect groundwater contamination patterns. The Manouba peri-urban basin, located in the western suburbs of Tunis, serves as a key agricultural and residential zone where groundwater constitutes the primary source for both domestic consumption and irrigation [
36]. Between 2005 and 2025, the basin experienced rapid urban expansion and population growth, increasing from approximately 0.35 million to 0.52 million inhabitants, driven by suburban development and intensified peri-urban agriculture. This accelerated anthropogenic activity has introduced multiple contaminants into groundwater, including nitrates from fertilizers, wastewater infiltration, and emerging organic pollutants. Human-driven land changes and urban expansion in Manouba are therefore major pressures degrading aquifer quality and heightening the vulnerability of the local population.
Given these conditions, the Manouba peri-urban basin represents an ideal case study for examining the interactions between anthropogenic pressures and groundwater contamination through our Geo-AI ensemble framework. The specific objectives of this research are: (1) to characterise the spatial and temporal patterns of human-driven land changes and urbanisation in the Manouba basin; (2) to evaluate groundwater quality by integrating subjective expert judgment with objective data-driven methods; and (3) to implement a novel grid-based framework that couples machine learning models with spatial sampling to quantify the impacts of anthropogenic pressures on groundwater contamination. The findings offer a data-driven, transferable framework to support sustainable groundwater management under increasing anthropogenic pressures in peri-urban agricultural aquifers.
2. Study Area and Data Collection
2.1. Study Area
The study area is the Manouba peri-urban basin, situated in the western suburban fringe of Tunis, northern Tunisia (
Figure 1). The basin extends over approximately 160 km
2, with elevations generally ranging between 10 and 120 m above sea level. It represents a strategic peri-urban agricultural zone where rapid residential expansion and intensive farming coexist [
37,
38]. The basin is bordered by the Tebourba plain to the west and the urban agglomeration of Tunis to the east, forming a transitional interface between metropolitan development and irrigated agricultural lands. The terrain is relatively flat to gently undulating, facilitating both urban sprawl and agricultural intensification [
39].
The Medjerda River, the principal fluvial system in northern Tunisia, flows in proximity to the basin and plays a significant role in regional hydrological dynamics [
36,
39]. The area is characterized by a Mediterranean semi-arid climate, marked by mild, wet winters and hot, dry summers [
36,
40]. The mean annual temperature is approximately 18 °C, with summer maxima frequently exceeding 40 °C [
36]. Average annual precipitation ranges between 400 and 450 mm, occurring predominantly from October to March, while potential evapotranspiration surpasses rainfall during most of the year, contributing to water stress conditions [
36,
39,
40].
Geologically, the Manouba basin is composed mainly of Quaternary alluvial deposits overlaying Mio-Pliocene formations [
36,
40,
41,
42]. The shallow aquifer system consists primarily of unconsolidated sands, silts, gravels, and clayey intercalations with variable thickness, typically between 20 and 80 m. Hydrogeologically, the aquifer is predominantly phreatic and highly vulnerable to surface-derived contamination due to its shallow water table and permeable sedimentary structure [
36,
41,
42,
43]. Recharge occurs mainly through rainfall infiltration, irrigation return flows, and lateral contributions from surrounding highlands, while groundwater extraction for domestic supply and agricultural irrigation constitutes the dominant discharge mechanism. Regional hydraulic gradients generally direct groundwater flow toward the northeast, following the topographic slope and converging toward lower-lying zones influenced by urban development and agricultural activity [
41,
43].
2.2. Data Collection and Processing
Groundwater monitoring in the Manouba peri-urban basin was conducted through two sampling campaigns representing distinct stages of anthropogenic pressure. In total, 295 samples were collected from 85 shallow phreatic wells: 223 samples obtained from 54 monitoring wells in 2005 and 72 samples from 31 wells in 2025. All samples were drawn from the unconfined aquifer system, which is particularly vulnerable to surface-derived contamination due to its limited depth and permeable alluvial deposits. The geographic coordinates of each sampling location were recorded to enable precise spatial analysis. Prior to sampling, wells were purged to ensure representative groundwater conditions, after which samples were collected, preserved under controlled temperature conditions, and transported for laboratory analysis within an appropriate timeframe to maintain data integrity.
It is acknowledged that the groundwater samples collected in 2005 and 2025 were obtained from partially different sets of monitoring wells. As a result, the temporal comparison does not rely on paired observations from identical locations. To address this limitation, both datasets were analyzed within a consistent geospatial framework. Hydrochemical parameters from each period were interpolated and aggregated onto a common spatial grid, allowing for a harmonized comparison of groundwater quality patterns across the study area.
This approach reduces the influence of individual well variability and enables the assessment of regional-scale temporal changes in groundwater quality. While the absence of identical sampling points may introduce some uncertainty in local-scale comparisons, the relatively large number of samples in each period supports a robust characterization of spatial trends and their evolution over time.
The hydrochemical assessment included major physicochemical indicators and ionic constituents, namely pH, total dissolved solids (TDS), potassium (K+), sodium (Na+), calcium (Ca2+), magnesium (Mg2+), chloride (Cl−), bicarbonate (HCO3−), sulfate (SO42−), and nitrate (NO3−). Field measurements were conducted for in situ parameters, while laboratory analyses were performed using standardized analytical procedures to quantify major cations and anions. Data reliability was ensured through systematic quality control protocols, including duplicate sampling, blank controls, and verification of analytical precision. The ionic charge balance error was calculated for each sample, and only results within an acceptable uncertainty threshold (±5%) were retained for further analysis.
In addition to hydrochemical data obtained from groundwater sampling campaigns, ancillary datasets were incorporated to account for potential anthropogenic drivers influencing groundwater quality. These include land use/land cover maps derived from remote sensing data (e.g., urban areas, cultivated lands, and natural covers), which were used as proxy indicators of human pressure on the aquifer system. These spatial datasets enable the characterization of surface activities that may contribute to groundwater contamination through recharge modification, pollutant infiltration, and surface runoff processes.
Although direct measurements of groundwater abstraction volumes and discharge rates were not available at the scale of the study, land-use indicators, particularly the extent of urbanized and agricultural areas, were used as indirect proxies for anthropogenic pressure. These proxies are commonly adopted in hydrogeological studies when direct pumping or discharge records are limited or unavailable.
The groundwater datasets collected in 2005 and 2025 are treated as two independent temporal snapshots representing distinct hydrochemical conditions of the aquifer system. The optimization procedures applied in this study are not based on the assumption of temporal stationarity of groundwater contamination. Instead, they are used to derive weights and indices that are applicable within each dataset and to enable spatial comparison between the two time periods. Consequently, the approach does not assume that contamination levels remain constant over time, but rather quantifies their evolution by comparing two independent observations of the system at different dates.
3. Composite Contamination Index (CCI)
Reliable quantification of groundwater contamination under intensifying anthropogenic pressures requires a weighting strategy that reconciles hydrogeochemical variability with expert-driven environmental interpretation. Weighting schemes derived solely from professional judgment may introduce epistemic uncertainty due to differences in expert perception of contaminant relevance, whereas purely statistical approaches can overlook hydrogeological processes and contextual aquifer vulnerability [
44,
45,
46]. In peri-urban agricultural systems such as the Manouba basin, where contamination arises from interacting urban expansion, agricultural intensification, and groundwater abstraction, an integrated weighting framework is therefore essential to produce a robust and interpretable contamination metric. To address this challenge, we developed a Composite Contamination Index (CCI) based on a hybrid methodology that combines structured subjective prioritization through the Best-Worst Method (BWM) with objective data-driven weighting using the CRITIC approach [
47,
48,
49,
50,
51]. Final weights are determined via constrained quadratic optimization, ensuring mathematical consistency while preserving both conceptual and statistical information.
3.1. Indicator Standardization
Ten hydrochemical indicators were considered: pH, total dissolved solids (TDSs), potassium (K+), sodium (Na+), calcium (Ca2+), magnesium (Mg2+), chloride (Cl−), bicarbonate (HCO3−), sulfate (SO42−), and nitrate (NO3−). Because these variables differ in scale, unit, and environmental significance, direct aggregation would distort their relative contribution. Therefore, normalization was required to transform all indicators into dimensionless contamination intensities.
For contamination-type parameters whose impact increases proportionally with concentration, normalization was performed as (Equation (1)):
where
xij represents the value of the j-th parameter for the i-th sample,
sj denotes the standard deviation of the j-th parameter across all samples, and
zij is the standardized value (z-score) of the j-th parameter for the i-th sample. For pH, whose environmental risk depends on deviation from an acceptable standard value, the following formulation was used (Equation (2)):
In this formulation, xij represents the value of indicator j for sample i, while Sj denotes the reference value (e.g., the mean or standard deviation depending on the adopted normalization scheme) of indicator j across all samples. The term zij is the normalized value of indicator j for sample i, computed as a relative deviation from the reference value Sj. This transformation allows the indicators to be expressed on a comparable scale by accounting for their relative variability and magnitude. The resulting normalized matrix is subsequently used in the aggregation and weighting procedures of the contamination assessment framework.
3.2. Subjective Weighting Using the Best–Worst Method
The Best–Worst Method (BWM) was applied to derive structured subjective weights based on hydrogeological expertise regarding contamination drivers in the peri-urban aquifer. The preference judgments were established by the author based on a synthesis of published studies and prior experience in similar hydrogeological settings. Experts first identified the most influential indicator (Best) and the least influential indicator (Worst) in terms of groundwater degradation. Pairwise preference ratios were then assigned to express the relative importance between the Best indicator and each of the remaining indicators, as well as between all indicators and the Worst one.
The optimization problem is formulated as (Equation (3)):
subject to (Equation (4))
In this formulation, wj represents the subjective weight assigned to indicator j. The terms wB and wW correspond to the weights of the Best and Worst indicators, respectively. The coefficient aBj denotes the preference intensity of the Best indicator over indicator j, while ajW expresses the preference of indicator j over the Worst indicator. The scalar ξ represents the maximum deviation to be minimized, ensuring consistency in pairwise comparisons. The solution of this constrained optimization yields the subjective weight vector that best satisfies the expert preference structure while maintaining normalization.
3.3. Objective Weighting Using the CRITIC Method
To complement the expert-based approach, the CRITIC method was implemented to quantify the intrinsic informational contribution of each indicator. This method evaluates both the variability of each variable and its redundancy relative to others.
The informational content of indicator j is computed as (Equation (5)):
and the corresponding objective weight is obtained through normalization (Equation (6)):
Here, σj denotes the standard deviation of the normalized values of indicator j, reflecting its contrast intensity across the study area. The coefficient rjk represents the Pearson correlation coefficient between indicators j and k, quantifying redundancy. The term Cj therefore captures the combined effect of dispersion and independence. Indicators exhibiting high variability and low correlation with others receive greater objective weight because they provide more unique information about contamination processes.
3.4. Weight Optimization and Classification of CCI
To reconcile subjective and objective weighting schemes, an integrated weight vector was obtained by minimizing the total squared deviation between the final weights and both preliminary weight sets (Equation (7)):
subject to (Equation (8)):
In this equation, wjw_jwj represents the optimized integrated weight for indicator j, while wjBWM and wjCRITIC denote the subjective and objective weights, respectively. The quadratic minimization ensures that the final weights remain statistically stable while preserving expert knowledge. This dual-consistency mechanism enhances methodological robustness and reduces weighting bias.
The Composite Contamination Index for groundwater sample
i is calculated as (Equation (9)):
where
CCIi represents the aggregated contamination intensity at location
i, obtained by summing the product of the optimized weight
wj and the standardized contamination score
zij across all indicators. This weighted aggregation captures both exceedance magnitude and relative importance.
To facilitate interpretation for groundwater management, the CCI values were classified into five contamination categories as shown in
Table 1.
These thresholds reflect increasing exceedance intensity relative to regulatory standards and were calibrated based on the statistical distribution of contamination scores within the peri-urban aquifer system.
4. Geo-AI Grid-Based Modeling
4.1. Methodological Framework
To quantify land transformations and related anthropogenic pressures in the Manouba peri-urban aquifer from 2005 to 2025, a spatially explicit grid-based framework was applied (
Figure 2). Rather than relying on administrative boundaries, the study employed a regular lattice in which each cell represented an independent spatial-temporal unit. Land-change datasets for 2005 and 2025 were overlaid onto this grid to enable consistent assessment of surface transitions and pressure dynamics over the twenty-year period.
It is important to note that the groundwater sampling campaigns conducted in 2005 and 2025 do not consist of identical monitoring wells. Therefore, temporal changes in groundwater quality are not evaluated at the level of individual wells, but rather at the spatial distribution level across the study area.
To address this, hydrochemical data from both periods were integrated within a geospatial framework, and spatial interpolation combined with grid-based aggregation was used to ensure comparability between the two time slices. This approach allows the characterization of overall temporal trends in groundwater quality by comparing spatial patterns of contamination indicators (e.g., CCI) between 2005 and 2025, rather than relying on paired time-series observations at fixed monitoring points.
Each grid cell measures 500 m × 500 m (0.25 km2), ensuring a suitable compromise between spatial resolution and computational efficiency. At the raster scale, each cell contains 18 × 18 pixels, allowing aggregation of pixel-level land information into standardized analytical units while reducing local noise. The grid was generated using ArcGIS 10.8 to ensure geometric consistency and accurate alignment of multi-temporal data. After removing cells lacking valid land-change information, 900 effective grid units were retained to support the analysis of spatiotemporal anthropogenic pressure intensification.
A centroid sampling point was assigned to each grid cell to ensure consistent data extraction. For every unit, the proportional extent of key land categories linked to anthropogenic pressures (urban areas, bare land, agriculture, forest, and water bodies) was calculated for 2005 and 2025 (
Figure 3). The LULC maps were generated through remote sensing data processing using Landsat 7 ETM+ imagery for 2005 and Landsat 8 OLI imagery for 2025. Temporal change was quantified by computing percentage differences between the two years, reflecting both the magnitude and direction of land-use transformation.
To link land dynamics with groundwater conditions, hydrochemical parameters and Composite Contamination Index (CCI) values for 2005 and 2025 were averaged within each grid cell. The evolution of groundwater contamination was then expressed as a relative change rate derived from CCI values (Equation (10)).
In this formulation, E represents the proportional variation in contamination intensity over the 2005–2025 interval, while CCI2005 and CCI2025 represent the Composite Contamination Index values at the initial and final years, respectively. This normalized indicator allows comparison of contamination progression irrespective of baseline conditions.
The calculated land-transformation proportions and aggregated hydrochemical variables were subsequently used as predictor features in machine learning models to simulate and predict the spatial evolution of CCI over the study period. For clearer interpretation and mapping, the predicted CCI change rates for 2025 were converted back into absolute CCI values, allowing direct comparison with observed contamination patterns. The overall methodological framework is summarized in
Figure 2.
4.2. Geo-AI Model Selection and Architecture
The expansion of groundwater monitoring in the Manouba peri-urban aquifer has produced high-dimensional datasets with spatial heterogeneity, temporal variability, and nonlinear interactions driven by land changes and anthropogenic pressures. Traditional statistical methods struggle to capture these complexities, whereas advanced Geo-AI approaches can model nonlinear relationships, learn hierarchical features, and integrate spatial structure for more accurate predictions.
To develop the Geo-AI Ensemble Modeling Framework for Assessing Groundwater Contamination under Anthropogenic Pressures, three complementary algorithms were employed: Graph Neural Networks (GNNs) for spatial connectivity, Light Gradient Boosting Machine (LightGBM) for efficient nonlinear feature modeling, and Deep Long Short-Term Memory Networks (LSTM) for temporal dynamics. Collectively, these models capture the spatial, structural, and temporal dimensions of contamination under evolving land transformations.
4.2.1. Graph Neural Networks (GNNs)
Groundwater contamination patterns exhibit strong spatial dependency due to hydraulic connectivity and shared anthropogenic drivers [
27,
28,
29]. To explicitly incorporate spatial structure, the study area was represented as a graph
G = (
V,E), where grid cells correspond to nodes
V, and edges
E represent spatial adjacency or hydrogeological connectivity.
The graph convolution operation is expressed as (Equation (11)):
where
à =
A +
I denotes the adjacency matrix with self-loops,
Ñ is the corresponding degree matrix,
H(l) represents node features at layer
l,
W(l) is the trainable weight matrix, and
σ is a nonlinear activation function. The initial feature matrix
H(0) includes land-change proportions, anthropogenic pressure indicators, and hydrochemical variables.
The final node representation is mapped to predicted CCI change rates through (Equation (12)):
where
Hi(L) denotes the learned embedding of node
i after
L graph convolution layers, and
ŷi represents the predicted contamination change rate. By propagating information across neighboring cells, GNNs effectively capture spatial diffusion effects and localized anthropogenic influences.
4.2.2. Light Gradient Boosting Machine (LightGBM)
To model complex nonlinear relationships between land transformations and groundwater contamination, LightGBM was employed as a high-performance gradient boosting framework [
23,
30,
31]. LightGBM constructs an ensemble of decision trees in a sequential manner, minimizing a differentiable loss function.
At iteration t, the objective function is defined as (Equation (13)):
where
yi is the observed CCI change rate,
ŷi(t−1) is the previous prediction,
ft represents the newly added decision tree, and
Ω(
ft) is a regularization term controlling model complexity.
The prediction at iteration ttt is updated as (Equation (14)):
where
η denotes the learning rate. By iteratively correcting residual errors, LightGBM efficiently captures nonlinear interactions among land-change proportions, anthropogenic indicators, and hydrochemical variables while reducing overfitting through regularization.
4.2.3. Deep Long Short-Term Memory Networks (LSTM)
Groundwater contamination evolution reflects cumulative and delayed effects of anthropogenic pressures [
32,
33,
34,
35]. To model temporal dependencies across the 2005–2025 interval, a deep LSTM architecture was implemented. LSTM networks incorporate gated memory mechanisms that regulate information flow across time steps.
The forget, input, and output gates are formulated as (Equation (15)):
where
xt represents the input feature vector at time
t,
ht−1 is the previous hidden state, and
W,
U, and
b denote trainable parameters. The gates control memory updating and state propagation, enabling the network to learn long-term contamination trends driven by progressive land transformations.
The final output layer produces predicted CCI change rates based on the learned temporal representation. By capturing delayed hydrochemical responses to anthropogenic pressures, LSTM enhances temporal forecasting capability.
4.2.4. Ensemble Integration
To leverage complementary strengths, outputs from GNN, LightGBM, and LSTM models were integrated within an ensemble framework. Spatial dependencies captured by GNN, nonlinear feature interactions modeled by LightGBM, and temporal evolution learned by LSTM collectively improve predictive robustness. This integrated Geo-AI architecture strengthens the assessment of groundwater contamination under land changes and anthropogenic pressures, directly supporting sustainable groundwater management in the extensive peri-urban agricultural aquifer.
4.3. Data Preparation and Normalization
For the Geo-AI ensemble framework, the input features consisted of land-change proportions, anthropogenic pressure indicators, and hydrochemical parameters from 2005 and 2025, while the output was defined as the change rate of the Composite Contamination Index (CCI) over the 2005–2025 period. Prior to model training, all input variables were normalized to eliminate differences in units and scales among hydrochemical and land-change parameters, producing dimensionless values suitable for machine learning.
To mitigate overfitting and enhance model generalization, L2 regularization was applied during preprocessing. The regularized loss function is expressed as (Equation (16)):
where
L(
w) is the original loss,
wj are the model weights,
n is the number of input features, and
α ∈ [0, 1] is the regularization coefficient controlling the strength of penalization. All preprocessing steps and normalization procedures were implemented using Python, ensuring consistency across the GNN, LightGBM, and LSTM models.
4.4. Model Performance Evaluation
Model performance was assessed using three standard metrics: the root mean squared error (RMSE), mean absolute error (MAE), and the coefficient of determination (R
2), calculated as follows (Equations (17)–(19)):
where
yi and
ŷi are the observed and predicted CCI change rates for grid cell
i,
ȳ is the mean of observed CCI change rates, and
m is the total number of samples. High predictive accuracy is indicated by low RMSE and MAE values and an R
2 value approaching 1.
To ensure reliable model evaluation and mitigate overfitting, the dataset was divided into training and testing subsets using a hold-out validation strategy. Model performance was assessed on unseen data, and the process was repeated using multiple random partitions to improve robustness and reduce the influence of sampling variability. The reported performance metrics correspond to the average results obtained across these repetitions. This approach is appropriate for the available dataset, which consists of a moderate number of samples with spatial heterogeneity. In addition, inherent regularization mechanisms in LightGBM and early stopping criteria (where applicable) were used to further control model complexity and enhance generalization.
To evaluate the influence of land transformations and anthropogenic pressures on groundwater contamination predictions, a variable importance analysis was conducted using permutation-based techniques [
23,
34,
35]. This analysis highlights the relative contribution of each predictor to the modeled CCI dynamics, providing insight into which pressures most strongly drive contamination patterns in the Manouba peri-urban aquifer.
All data preprocessing, model training, and evaluation were implemented in Python (V 3.12.0) using the scikit-learn library for LightGBM and permutation analyses, PyTorch (V 2.6.0) for LSTM, and PyG (PyTorch Geometric) (V 2.6.0) for GNNs, ensuring reproducibility and computational efficiency.
6. Discussion
This study provides a comprehensive assessment of groundwater quality evolution in the Manouba peri-urban aquifer between 2005 and 2025, highlighting the dominant role of land transformation and anthropogenic pressures. The CCI results reveal a marked deterioration in groundwater quality over the study period, reflecting the combined effects of urban expansion and agricultural intensification.
Spatial patterns indicate that groundwater contamination is strongly structured by anthropogenic drivers rather than randomly distributed. High CCI values are concentrated in areas undergoing intensive urbanisation and cultivation, where impervious surfaces alter recharge processes and increase pollutant residence time in the subsurface [
7,
10,
14]. Nitrate emerged as the most influential parameter controlling CCI variability, confirming the dominant role of agricultural fertilisation and urban/industrial inputs [
39,
43]. This apparent dominance should be interpreted in a multivariate statistical sense, as it reflects the relative contribution of nitrate within the CCI framework and its broader concentration variability, rather than implying it is the sole controlling factor of hydrogeochemical processes. In this context, other indicators such as chloride and total dissolved solids also point to concurrent salinisation processes, indicating that groundwater quality degradation results from multiple interacting anthropogenic and geogenic drivers. The spatial shift of nitrate hotspots toward more urbanised zones by 2025 reflects the evolving nature of contamination sources under progressive land transformation.
From a hydrogeochemical perspective, two main mechanisms explain the observed trends. First, urbanisation modifies the natural hydrological cycle by reducing infiltration and enhancing runoff and evapoconcentration, leading to increased solute accumulation. Second, agricultural and industrial activities contribute sustained contaminant loading, particularly nitrate and dissolved salts, resulting in progressive salinisation and localised contamination hotspots. These processes are widely documented in Mediterranean peri-urban aquifers experiencing rapid land-use change [
14,
15,
23].
The Geo-AI ensemble framework provides further insight into these dynamics. Among the tested models, LightGBM demonstrated the highest predictive accuracy, highlighting its capacity to capture nonlinear interactions and heterogeneous hydro-chemical responses. GNNs effectively represented spatial dependencies related to hydrogeological connectivity; however, the GNN architecture is based on spatial adjacency and does not explicitly incorporate directional hydraulic gradients or groundwater flow vectors; therefore, it captures spatial connectivity patterns rather than physically constrained advective transport pathways. LSTM showed comparatively lower performance, which can be attributed both to the limited temporal resolution of the dataset (two discrete time points) and to the predominance of spatially driven contamination processes. It should be noted that deep LSTM networks are typically designed to exploit multiple sequential time steps, and their use in this study is therefore exploratory given the limited temporal sampling. This result confirms that ensemble models such as LightGBM are more suitable for the present study configuration.
A comparison with conventional groundwater quality indices, such as the Water Quality Index (WQI), highlights the added value of the proposed Composite Contamination Index (CCI). While WQI remains widely used due to its simplicity, it typically relies on fixed or expert-based weighting schemes that may not fully capture complex interactions among hydrochemical variables. In contrast, the CCI integrates subjective (BWM) and objective (CRITIC) weighting within an optimization framework, allowing a more adaptive and data-driven representation of contamination processes. Furthermore, coupling CCI with Geo-AI models enables spatial prediction and dynamic analysis, extending beyond the static assessment provided by traditional indices. This integration enhances both interpretability and predictive capability, positioning CCI as a more suitable tool for dynamic groundwater quality assessment in rapidly evolving peri-urban systems.
Although machine learning models are typically data-hungry, the present study operates with a relatively moderate dataset consisting of 295 samples distributed across 85 wells over two time periods. To mitigate the limitation of limited data availability, several strategies were adopted. First, the dataset was structured using spatially distributed samples rather than relying on a single temporal sequence, which increases the effective diversity of the training data. Second, ensemble-based and tree-based models such as LightGBM were included, as they are known to perform well with limited to moderate-sized datasets and are less prone to overfitting compared to deep neural networks.
In addition, cross-validation and appropriate regularization mechanisms were employed to ensure model generalization and robustness. Data balancing techniques such as synthetic oversampling (e.g., SMOTE) were not applied in this study, as they may introduce artificial variability in continuous hydrochemical datasets and potentially affect the physical interpretability of model outputs.
From a management perspective, the findings provide important implications for groundwater protection. Urban planning should incorporate strategies to preserve recharge zones and limit excessive soil sealing. At the same time, improved regulation of agricultural practices and wastewater management is essential to control nitrate inputs and mitigate salinisation. The integration of grid-based monitoring with AI-driven predictive tools offers a promising approach for early detection of contamination trends and adaptive groundwater management.
It should be emphasized that the methodology does not assume temporal stationarity of groundwater contamination. The 2005 and 2025 datasets represent two independent snapshots of the aquifer system, and the observed differences reflect real temporal changes driven by evolving anthropogenic and environmental conditions, rather than a constant underlying contamination state. It should also be acknowledged that land-use practices, irrigation strategies, and regulatory frameworks may have evolved between the two periods. However, these changes are implicitly reflected in the observed spatial differences, as the study treats each year as an independent snapshot rather than modeling continuous temporal transitions.
A key limitation of this study lies in both the sampling design and the spatial aggregation approach. Groundwater samples collected in 2005 and 2025 originate from partially overlapping but non-identical sets of monitoring wells, meaning that the 2025 dataset does not represent a strict subset of the 2005 network. This prevents direct paired comparisons at identical locations and introduces uncertainty in point-based temporal interpretation. Consequently, temporal variations are evaluated at the level of spatial distributions rather than point-based trajectories. In addition, the dataset exhibits uneven temporal density, with more extensive sampling in 2005 than in 2025, as well as a non-uniform spatial distribution of wells. To address these constraints, hydrochemical parameters and CCI values were aggregated within a 500 m × 500 m grid framework, enabling consistent regional-scale comparison between the two periods. Although high-precision GPS coordinates were available for individual wells, the chosen grid resolution represents a compromise between spatial accuracy and data density, ensuring robust interpolation while maintaining consistency with the sampling distribution. However, this aggregation may smooth local variability and partially mask small-scale contamination hotspots, which are particularly relevant in peri-urban environments characterized by heterogeneous pollution sources. Similarly, the aggregation of land-use information at the pixel level (18 × 18) may also reduce high-frequency spatial variability associated with localized urban-industrial contamination sources, although this trade-off was necessary to reduce local noise and ensure consistency with the scale of hydrochemical observations. Despite this limitation, the combined use of geostatistical interpolation and spatial modelling preserves the dominant contamination gradients and allows robust identification of regional patterns. Future studies should incorporate higher-resolution sampling strategies and denser monitoring networks to better capture fine-scale heterogeneity and localized contamination processes.
Future work should focus on increasing sampling density, incorporating multi-temporal monitoring datasets, and integrating additional tracers such as emerging contaminants and isotopic indicators. It should be noted that the present study is limited to inorganic hydrochemical parameters, while emerging contaminants were not included in the analysis and are considered as a direction for future research rather than part of the current dataset. These improvements would enhance both process understanding and the predictive capability of Geo-AI frameworks, supporting more robust groundwater management in peri-urban environments.
Another avenue for improvement concerns the integration of physical hydrological processes into the modelling framework. Incorporating groundwater flow information, such as Darcy-based flow vectors, into the graph structure of GNN models could enhance the representation of contaminant transport pathways. However, such an approach requires detailed hydraulic and piezometric data, which were not available for the full study period. Future work should explore hybrid modelling strategies combining data-driven approaches with physically based constraints to improve process realism.
More advanced hybrid modeling approaches, such as Physics-Informed Neural Networks (PINNs), offer a promising avenue to further enhance groundwater modeling by embedding physical laws directly into the learning process through physics-based constraints in the loss function. Such approaches have recently been applied in groundwater and water resources studies and hydrogeological studies to improve model generalization, particularly in data-scarce contexts. However, the implementation of PINNs typically requires well-defined governing equations and additional calibration of physical parameters, which are beyond the scope of the present study. Future work could explore the integration of physics-informed learning frameworks to further improve predictive performance and physical consistency of groundwater quality models [
52,
53,
54].
7. Conclusions
This study presents a comprehensive assessment of groundwater quality in the Manouba peri-urban aquifer between 2005 and 2025 by integrating hydrochemical analyses, land-change dynamics, and advanced machine-learning approaches. The Composite Contamination Index (CCI) revealed a clear deterioration of groundwater quality over the study period. In 2005, groundwater conditions were relatively balanced, with low contamination covering 25% of the area and slight contamination 20%, while moderate, high, and severe contamination accounted for 30%, 18%, and 7%, respectively. By 2025, the spatial distribution of CCI classes shifted markedly toward more degraded conditions. Severe contamination became dominant, representing 55% of the area, followed by high contamination (15%) and moderate contamination (10%), whereas slight and low contamination declined to 12% and 8%, respectively. This evolution highlights the progressive expansion of highly contaminated zones within the aquifer system.
Machine-learning analyses demonstrated the effectiveness of advanced algorithms in modelling complex interactions between hydrochemical variables and land changes. Among the tested models, LightGBM achieved the highest predictive performance, with R2 = 0.986, RMSE = 13.14, and MAE = 1.72 for CCI prediction. GNNs also showed strong predictive capability (R2 = 0.972, RMSE = 18.73, MAE = 2.11), particularly in capturing spatial dependencies related to hydrogeological connectivity. In contrast, LSTM exhibited comparatively lower accuracy (R2 = 0.943, RMSE = 38.92, MAE = 4.83), reflecting the predominantly spatial rather than temporal nature of groundwater contamination processes. Residual analyses further confirmed that LightGBM and GNNs produced stable and largely unbiased predictions across the contamination spectrum, whereas LSTM displayed greater dispersion near CCI threshold boundaries.
Variable importance and sensitivity analyses highlighted the dominant influence of artificial surfaces and cultivated lands, which explained most of the spatial and temporal variability in CCI. Other land-cover types, including forests, wetlands, shrublands, and water bodies, contributed marginally due to their limited spatial extent. These results demonstrate that urban expansion and agricultural intensification are the principal drivers of groundwater degradation in the Manouba aquifer.
Two major mechanisms explain the influence of urbanisation on groundwater quality. First, the conversion of natural or agricultural land into artificial surfaces modifies recharge, flow, and discharge processes, leading to increased concentration of hydrochemical parameters in certain zones. Second, the concentration of industrial and urban activities in pollution-intensive areas further increases groundwater vulnerability, particularly to nitrate contamination.
The integration of hydrochemical monitoring, CCI-based contamination assessment, land-change analysis, and machine-learning modelling provides a robust framework for understanding groundwater quality evolution under increasing anthropogenic pressure. The results clearly indicate a progressive shift toward higher contamination levels between 2005 and 2025, with severe contamination expanding across large parts of the aquifer. Among the evaluated models, LightGBM emerged as the most reliable predictive framework, offering high accuracy and stability under heterogeneous hydrochemical conditions. These findings provide a strong scientific basis for improving groundwater management, guiding sustainable land-use planning, and implementing targeted pollution mitigation strategies in the Manouba region and in similar peri-urban aquifer systems.