1. Introduction
Pavement infrastructure has historically been conceived as a passive structural system whose principal purpose is to sustain traffic loading, distribute stresses, and maintain serviceability under long-term environmental exposure. Although this conventional view remains central to pavement engineering practice, it is increasingly insufficient for the emerging demands of sustainable transportation infrastructure, where roadway systems are expected not only to provide structural performance but also to contribute to resilience, intelligence, and resource efficiency. In this broader context, renewable energy-integrated pavement systems have attracted growing interest as a new class of multifunctional infrastructure in which structural layers are no longer treated solely as load-bearing media, but also as physically active domains capable of thermal and electromechanical energy harvesting [
1]. Such systems are particularly attractive because pavements are continuously exposed to wheel-induced mechanical excitation, solar-driven thermal gradients, and hydro-environmental actions, all of which simultaneously act as deterioration drivers and potential energy sources [
2]. The transition from passive pavements to energy-responsive smart pavement systems therefore represents not merely an incremental modification of pavement design, but a conceptual shift toward infrastructure that can structurally perform, environmentally interact, and functionally generate useful engineering outputs within the same physical domain [
3].
The analytical treatment of renewable energy-integrated pavements remains far more complex than that of conventional layered sections, since their response arises from the simultaneous interaction of structural mechanics, thermal transfer, moisture migration, and embedded energy-conversion mechanisms. In the present problem, the pavement cannot be interpreted as a purely mechanical body governed only by deflection or stress transmission. Instead, it must be understood as a layered thermo-hydro-mechanical continuum in which the asphalt surface and intermediate binder govern traffic interaction and near-surface conditioning, the base layer controls the main structural load-spreading function, the thermoelectric layer exploits vertical temperature differentials, the piezoelectric zone responds to traffic-induced compressive excitation, and the lower subbase and subgrade regulate hydraulic moderation, confinement, and global support stability. The final system behavior therefore emerges from a coupled field environment rather than from isolated layer contributions [
4].
A second research gap arises at the artificial intelligence stage. While machine learning and deep learning have been increasingly used for pavement condition prediction and infrastructure assessment, many available approaches still rely on flat feature representations that do not preserve the physical hierarchy of multilayer pavement systems. This limitation becomes even more critical in renewable energy-integrated pavements, where the predictive targets are not confined to conventional mechanical indicators but extend to temperature gradients, moisture-sensitive support response, electromechanical harvesting outputs, and probabilistic reliability descriptors. Under such conditions, the predictive problem is not merely one of nonlinear regression or classification, but one of learning how disturbances propagate across a structured infrastructure medium whose internal dependency pattern is intrinsically layered [
5]. A model that ignores this ordered architecture risks learning statistical associations without capturing the physical pathways through which load, heat, moisture, and energy interaction jointly evolve. For this reason, the present study departs from conventional tabular learning and reformulates each simulation case as a graph-structured pavement representation, in which the seven pavement constituents are treated as interacting nodes and the dominant structural, thermal, hydraulic, and functional transfer mechanisms are encoded through interlayer relations [
6]. This idea forms the conceptual basis of the proposed Layer-Coupled Reliability Graph Operator Network (LaRGO-Net), whose design is explicitly matched to the stratified physical logic of the pavement rather than imposed as a generic black-box predictor [
7].
To address these methodological and engineering gaps, this study proposes a novel simulation-driven intelligent framework for renewable energy-integrated pavement assessment that unifies multiphysics finite element modeling, high-fidelity data generation, and graph-based reliability-aware learning within a single computational pipeline. First, a three-dimensional coupled thermo-hydro-mechanical finite element model is established in Abaqus to simulate the concurrent effects of moving wheel load, solar heat flux, rainfall infiltration, and internal moisture diffusion in a seven-layer pavement with embedded thermoelectric and piezoelectric functional zones. Second, the simulation environment is used not merely for response visualization, but as a physics-based data-generation engine from which a structured AI-ready dataset is constructed, containing geometric, material, environmental, traffic, uncertainty, structural, thermal, hydraulic, energy-harvesting, and reliability-related descriptors. Third, LaRGO-Net is developed to preserve the layered physical hierarchy of the pavement in the learning process through node-wise feature encoding, global scenario conditioning, adaptive interlayer coupling, graph-level pooling, and multi-head prediction of structural, thermal, hydraulic, renewable-energy, and reliability targets. The framework is further strengthened by physics-consistency and coupling-consistency objectives, so that the model is optimized not only for predictive accuracy but also for engineering plausibility and inter-output coherence. Consequently, the research builds upon both traditional pavement modeling and traditional AI modeling by developing a reliability-focused intelligent surrogate model that is physically structured, multifunctional, and specifically designed for the new generation of renewable energy-powered pavements.
The importance of this work is its ability to bridge three fields that are traditionally investigated separately: multiphysics of advanced pavement systems, integrated renewable-energy functionality, and explainable artificial intelligence for infrastructure reliability. In the proposed approach, the finite element model is not merely regarded as a stand-alone engineering analysis, but rather as the source of physically relevant evidence; the dataset is not regarded as a numerical dataset, but instead as a representation of causal pavement physics; and the AI model is not seen as a black-box predictor, but as a layer-sensitive graph operator that learns how the pavement transitions into safe, marginal, or unsafe states under uncertain operating conditions. This holistic view is critical for next-generation infrastructure where pavement systems are expected to simultaneously provide traffic services, adapt to climate-sensitive operating conditions, harness distributed energy opportunities, and provide actionable information for predictive maintenance. As such, the current study is part of a larger infrastructure movement that transforms transportation surfaces from passive to active, intelligent, and functional energy systems.
Compared with existing thermo-hydro-mechanical pavement modeling, energy-harvesting pavement research, and graph-neural-network-based pavement prediction methods, the main innovation of this study lies in the unified integration of these directions within a single reliability-oriented framework. Existing coupled pavement models mainly focus on structural-environmental response, energy-harvesting studies often evaluate thermoelectric or piezoelectric functionality as separate modules, and graph-based pavement prediction methods generally emphasize structural response prediction without explicitly combining renewable-energy functionality and reliability interpretation. In contrast, the proposed framework integrates a seven-layer renewable energy-embedded pavement configuration, coupled Abaqus-based thermo-hydro-mechanical simulation, a 6000-case and 68-variable simulation-generated dataset, and the Layer-Coupled Reliability Graph Operator Network (LaRGO-Net). This integration enables simultaneous assessment of structural, thermal, hydraulic, thermoelectric, piezoelectric, and reliability indicators while preserving the physical hierarchy of the layered pavement system.
The principal contributions of this study are summarized as follows:
A seven-layer renewable energy-integrated pavement system is introduced, in which conventional structural layers are coupled with thermoelectric and piezoelectric harvesting zones to enable simultaneous structural support, thermal gradient utilization, traffic-induced energy harvesting, and reliability-oriented assessment.
A coupled thermo-hydro-mechanical Abaqus-based simulation framework is developed as a physics-based data-generation engine, producing an AI-ready dataset of 6000 cases and 68 variables covering geometric, material, environmental, traffic, uncertainty, structural, thermal, hydraulic, energy-harvesting, and reliability descriptors.
A Layer-Coupled Reliability Graph Operator Network (LaRGO-Net) is proposed to preserve the physical hierarchy of the multilayer pavement system through node-wise layer encoding, adaptive interlayer coupling, global scenario embedding, graph-level aggregation, and physics-aware multi-task prediction.
Figure 1 provides a high-level conceptual synthesis of the research problem addressed in this study by contrasting the functional limitations of conventional pavement systems with the engineering potential of renewable energy-integrated intelligent pavements. As illustrated on the left side, the conventional pavement is represented as a passive structural system exposed to the combined effects of traffic loading, solar heating, and rainfall-induced moisture ingress, where the dominant outcome of these coupled actions is progressive deterioration in the form of fatigue cracking, rutting, moisture damage, and pothole formation. This representation emphasizes that, within the traditional pavement paradigm, the layered structure primarily serves a load-bearing role and remains vulnerable to multi-physics degradation without providing any additional functional benefit beyond structural serviceability. The figure therefore captures an important gap in current pavement engineering practice, namely, that conventional assessment approaches are largely centered on distress identification and do not explicitly exploit the embedded physical processes of the pavement body for energy functionality or intelligent reliability-oriented interpretation. In contrast, the right side of the figure presents the proposed renewable energy-integrated intelligent pavement system as a multifunctional infrastructure platform in which the seven-layer pavement configuration is not only designed to resist coupled thermo-hydro-mechanical actions, but also to convert part of these actions into useful engineering functionality. In particular, the thermoelectric energy layer is positioned to utilize vertical heat gradients generated through solar exposure, while the piezoelectric insert zone is located within the mechanically active region to harvest traffic-induced compressive excitation, thereby transforming the pavement from a passive transport surface into an active energy-responsive medium. At the same time, the figure highlights that the proposed system supports a broader analytical interpretation through integrated structural, thermal, hydraulic, renewable-energy, and reliability response domains, which together define the need for intelligent performance prediction rather than isolated distress observation.
Mou et al. (2025) [
1] offer a targeted review of pavement piezoelectric energy harvesting, and they synthesize the current state of the art in terms of piezoelectric mechanisms, material types, harvester designs, finite element-based analysis, packaging, circuitry, and pavement applications. The research question they seek to answer is how the mechanical vibration due to pavement loads can be transformed into electrical energy for low-power, transportation-related functions, and the research problem is seen as the lack of a consolidated and still-emerging engineering view on road-integrated piezoelectric harvesting systems. The paper is a review, rather than a novel predictive or simulation-based approach, and synthesizes past research in terms of design, modeling and implementation. Their key finding is an engineering landscape of piezoelectric pavement harvesting, with emphasis on its potential and limitations in terms of device durability, integration, and efficiency. However, the main contribution missing from their paper is that the review does not integrate a pavement-scale model to model piezoelectric harvesting together with thermal gradients, moisture, pavement response, and probabilistic prediction of reliability within the same platform; this is the novelty of the current study, which uses a seven-layer thermo-hydro-mechanical Abaqus model combined with graph-based multi-task reliability learning for predicting the reliability.
Priyanka and Gao (2026) [
3] offer a review of pavement energy harvesting from the perspective of pavement engineering, focusing on traffic load-driven, solar-driven, and thermal gradient-driven energy harvesting concepts. The scope of their work is to assess the state of the art in pavement energy harvesting technologies and outline the main engineering opportunities for transforming road infrastructure into multifunctional responsive systems capable of harvesting energy, with the research question focusing on the need for a comprehensive engineering review of the field, given the different harvesting principles. Their approach is a review of a large collection of published papers, which classifies the literature based on harvesting mechanisms, performance tendencies and challenges in their implementation. The primary finding of the paper is that pavements can be used as multifunctional energy platforms for distributed harvesting, but deployment is highly dependent on structural compatibility, harvesting efficiency and operating climatic conditions. The key gap in this paper (which is addressed in the current work) is that their work is review-based and does not provide a simulation-based, layer-resolved predictive model that can jointly capture thermal, hydrologic, structural, thermoelectric, piezoelectric, and reliability processes in a single data-driven computational framework.
Liu and Al-Qadi (2025) [
6] propose a graph-neural-network-based framework for modeling three-dimensional asphalt concrete pavement responses from finite element simulation data, making their work one of the closest methodological references to the AI component of the present study. Their objective is to build a graph-based surrogate capable of reproducing 3D pavement responses with lower computational cost than repeated full finite element analyses, while the research problem they define is that conventional pavement-response modeling is computationally intensive and that unstructured learning representations do not fully exploit the underlying organization of pavement systems. Methodologically, they develop a validated 3D FE pavement model, generate hundreds of simulation cases, convert the FE output into graph structures, and train a GNN-based pavement simulator to predict structural responses. Their results show that graph-based learning can reproduce 3D FE pavement behavior effectively and can function as an efficient surrogate modeling tool. The research gap that remains, and that the present study addresses, is that their framework is centered on structural asphalt-concrete response modeling and does not extend toward a renewable energy-integrated pavement with coupled thermo-hydro-mechanical behavior, embedded thermoelectric and piezoelectric layers, multi-domain targets, and reliability-oriented physics-aware optimization.
Chen et al. (2025) [
5] develop an improved recurrent-neural-network-based pavement performance prediction model using the LTPP database, focusing on the prediction of pavement deterioration through deep learning. The objective of their study is to enhance pavement performance forecasting by incorporating multiple influential factors, while the research problem is defined as the limited capability of earlier prediction models to capture nonlinear, multivariate deterioration patterns in long-term pavement datasets. Their methodology consists of identifying major contributing factors such as initial condition indicators, traffic load, weather, pavement structure, and maintenance measures, and then training an RNN-based prediction model using LTPP data. Their results demonstrate that deep learning improves pavement performance prediction and can model nonlinear degradation behavior more effectively than simpler approaches. The key gap that the present study addresses is that their model remains fundamentally tabular and performance-oriented, without explicitly preserving pavement–layer hierarchy, without embedding coupled thermo-hydro-mechanical field behavior, and without incorporating renewable-energy outputs and probabilistic reliability descriptors into a single graph-structured predictive framework.
Hei et al. (2026) [
8] propose a Bayesian-neural-network-based model for pavement management engineering, referred to as an all-in-one performance prediction framework. The objective of their study is to develop a generalized pavement prediction model capable of handling multiple management-related tasks within one probabilistic deep learning architecture, while the research problem is that many pavement prediction methods either address only specific tasks or do not explicitly quantify predictive uncertainty for engineering decision support. Methodologically, the study adopts a Bayesian neural network formulation to produce unified prediction outputs together with uncertainty-aware estimates that are relevant for pavement management applications. Their reported results indicate that the Bayesian framework improves the generality of pavement-performance prediction and contributes a more uncertainty-conscious modeling strategy. The gap that remains, and that the present study directly addresses, is that their model is not formulated around a coupled renewable energy-integrated pavement system and does not encode pavement layers as interacting graph nodes linked by mechanical, thermal, hydraulic, and functional transfer pathways; in contrast, the present work combines graph-structured multiphysics learning with embedded energy-harvesting layers and reliability-specific targets such as
,
, safety class, and damage-risk level.
Rahmani and Kim (2025) [
2] develop a coupled thermo-mechanical finite element framework for reflective cracking in flexible pavements and thereby contribute an important recent reference on multi-field pavement simulation. The objective of their study is to simulate reflective cracking under combined thermal and mechanical loading using a more rigorous mechanistic framework, while the research problem is that classical pavement analyses often simplify thermal-mechanical interaction and therefore fail to represent cracking evolution under realistic service conditions. Their methodology integrates computational fracture mechanics, mixture-level fracture characterization of bituminous materials, and structural-level finite element modeling of flexible pavements under coupled thermal and mechanical actions. Their results show that combined thermal and mechanical loading intensifies reflective cracking behavior and that a coupled-field treatment is necessary for realistic damage assessment. The gap addressed by the present study is that, although their work is strong in thermo-mechanical pavement simulation, it does not extend to a broader thermo-hydro-mechanical environment with embedded renewable-energy functionality, simulation-generated AI-ready datasets, and graph-based multi-task reliability prediction; the current study fills that exact gap by adding hydraulic transport, thermoelectric and piezoelectric integration, and reliability-aware intelligent learning.
Ma et al. (2025) [
4] investigate asphalt-pavement damage mechanisms under coupled salt-thermal-mechanical effects using a multi-scale modeling framework, highlighting the growing importance of multi-field pavement analysis in recent literature. The objective of their study is to identify damage mechanisms under aggressive environmental and mechanical coupling, while the research problem is that single-field or simplified pavement models cannot adequately explain deterioration under chemically and thermally complex operating conditions. Methodologically, they combine multi-scale modeling with laboratory-based parameter identification and validation procedures, including experimental testing to calibrate and verify their numerical interpretation of damage evolution. Their results show that salt, thermal, and mechanical interactions can significantly accelerate pavement damage and that multi-field modeling is essential for realistic durability assessment. The research gap that the present study addresses is that their work, despite its strong multi-field mechanics, does not translate coupled pavement behavior into a simulation-driven AI framework with structured dataset generation, embedded thermoelectric and piezoelectric functionality, and graph-based multi-task prediction of structural, thermal, hydraulic, energy, and reliability responses; the present study explicitly resolves that gap by integrating those components into one unified framework.
Ma et al. (2025) [
4] investigate asphalt-pavement damage under coupled salt thermal mechanical effects using multi-scale modeling, making their work relevant to recent coupled pavement simulation research. Their objective is to explain how salt erosion, temperature variation, and mechanical loading jointly accelerate asphalt deterioration, while the research problem is that simplified single-field models cannot capture combined environmental mechanical degradation. They use pull-off tests, semicircular bending tests, cohesive zone modeling, and multi-scale finite element simulation. Their results show that after 48 h of salt treatment, cohesive strength decreased by 26.0 32.5% and adhesive strength by 33.5 63.8% at
, while at
, cohesive strength decreased by 24.9 44.0% and adhesive strength by 37.9 71.6%; fracture energy at
was also reduced by 4.2 13.0% compared with
, and salt penetration reached only about 10 mm after 48 h. However, their work remains focused on salt thermal mechanical damage and does not include renewable-energy layers, thermoelectric/piezoelectric outputs, AI-ready dataset generation, or graph-based reliability prediction, which are addressed in the present study.
Yin et al. (2026) [
9] propose a hybrid pavement module integrating photovoltaic and piezoelectric effects, providing a recent example of multi-source energy-harvesting pavement research. Their objective is to improve pavement energy-harvesting efficiency by combining solar conversion with traffic-induced piezoelectric generation, while the research problem is that single photovoltaic pavement systems suffer from intermittency caused by day night variation, weather changes, and vehicle shading. They combine finite element simulation, mathematical power-generation modeling, and physical prototype testing under a 20 kN standard vertical load. Their results report a simulated fatigue baseline of
cycles, a simulated piezoelectric power yield of about 22 mW at an optimal impedance of 680
, outdoor photovoltaic output power of 5.286 W, stable rectified piezoelectric mean power of 17.99 mW, and mean hybrid charging power of 1.55 W. Nevertheless, their work remains a prototype-level feasibility and safety-margin study under simplified boundary conditions, whereas the present study extends energy harvesting into a seven-layer thermo-hydro-mechanical pavement framework linked with structural, thermal, hydraulic, energy, and reliability prediction.
Liu and Al-Qadi (2025) [
7] develop a physics-informed graph neural network for three-dimensional spatiotemporal structural response prediction of flexible pavements, making their work closely related to the graph-learning component of the present study. Their objective is to reduce the computational cost of 3D finite element pavement response modeling while preserving physical consistency, and the research problem is that conventional 3D finite element models are computationally expensive whereas purely data-driven models may violate mechanics-based constraints. They transform 3D finite element pavement data into graph format and develop a physics-informed graph neural network-based pavement simulator using strain displacement and stress-equilibrium loss components. Their results show that both the data-driven graph pavement simulator and physics-informed graph pavement simulator achieved rollout prediction time of less than 8 s per finite element simulation case, compared with about 12 h for traditional 3D finite element pavement simulation, while maintaining strong short-term accuracy and better long-term robustness for the physics-informed model. However, their model focuses on structural pavement responses only and does not include renewable-energy layers, thermo-hydro-mechanical field coupling, thermoelectric/piezoelectric outputs, or reliability-oriented multi-task prediction, which are integrated in the present study.
Wang et al. (2025) [
10] develop an uncertainty-aware Bayesian dynamic linear model enhanced with a kernel regression basis function for high-precision multi-step prediction of structural health monitoring sensor streams under extreme typhoon events. Their objective is to improve SHM sensor forecasting, missing-data imputation, and uncertainty interpretation under severe environmental disturbance, while the research problem is that conventional deterministic models cannot robustly handle non-stationary and heteroscedastic sensor responses during typhoon conditions. They use a KR-BDLM framework and compare it with BP neural network and wavelet neural network predictors. Their results show that for single-step prediction, KR-BDLM achieved RMSE = 0.8049, which is 38.69% lower than the BP neural network RMSE of 1.3096; for multi-step prediction, KR-BDLM achieved RMSE = 3.5271, which is 9.44% lower than the BP neural network RMSE of 3.8946. The model also supports missing strain-data imputation and uncertainty estimation through prediction error, variance, and confidence-interval width. However, their work remains focused on SHM sensor-stream forecasting under typhoon events and does not include renewable-energy pavement layers, thermoelectric/piezoelectric outputs, thermo-hydro-mechanical pavement simulation, AI-ready dataset generation, or layer-coupled graph-based reliability prediction, which are addressed in the present study.
Cho et al. (2025) [
11] propose a real-time structural health monitoring framework based on Bayesian neural networks for distinguishing aleatoric and epistemic uncertainty in digital-twin-oriented diagnostics. Their objective is to reconstruct full-field strain distributions from sparse strain-gauge measurements while quantifying uncertainty, and the research problem is that SHM models must provide reliable confidence information rather than only deterministic predictions when sensor data are noisy, sparse, or affected by damage-induced strain discontinuities. They combine principal component analysis, Bayesian neural networks, and Hamiltonian Monte Carlo inference, and validate the framework through cyclic four-point bending tests on carbon fiber-reinforced polymer specimens with different crack lengths. Their results show that the proposed framework achieved accurate full-field strain reconstruction with
, while also generating real-time aleatoric and epistemic uncertainty fields. Nevertheless, their work focuses on experimental strain reconstruction for composite structural specimens, whereas the present study extends uncertainty-aware intelligent prediction to renewable energy-integrated multilayer pavements by coupling thermo-hydro-mechanical simulation, thermoelectric and piezoelectric energy-output prediction, reliability-index estimation, and layer-coupled graph-based multi-task learning.
3. Results
Table 9 shows a systematic progression from simple dense learning, through physics-informed and competitive graph/transformer baselines, to the full LaRGO-Net architecture. The baseline MLP in E1 represents the weakest configuration, using a 3-layer dense structure with
, dropout = 0.10, learning rate
, and no graph structure or physics-aware constraint. It produced the highest mean errors, with mean RMSE = 0.194, mean MAE = 0.152, mean
, mean F1 = 90.87%, and mean AUC = 94.18%. The energy block was the most difficult for this baseline, with RMSE = 0.218, MAE = 0.171,
, F1 = 89.92%, and AUC = 93.45%, confirming that a flat tabular model is not sufficient for representing coupled energy-harvesting behavior. Adding physical-consistency regularization in E2 substantially improved the dense baseline, reducing mean RMSE from 0.194 to 0.083 and MAE from 0.152 to 0.058, while increasing mean
from 0.861 to 0.974, mean F1 from 90.87% to 97.40%, and mean AUC from 94.18% to 98.23%. This improvement shows that physics-aware regularization helps constrain the response surface, but E2 still lacks explicit pavement–layer topology. The graph and attention-based baselines further improved performance: E3, the static GCN baseline, achieved mean RMSE = 0.071 and mean
, while E4, the GAT baseline, improved these values to mean RMSE = 0.059, mean MAE = 0.041, mean
, mean F1 = 98.49%, and mean AUC = 98.93%. However, because E3 uses fixed adjacency and E4 uses generic graph attention without pavement-specific adaptive interlayer coupling, both remain weaker than the final LaRGO-Net. Similarly, the transformer baselines in E5 and E6 confirm that strong tabular attention improves performance but does not fully replace the proposed physically structured graph operator. TabTransformer in E5 achieved mean RMSE = 0.066 and mean
, whereas FT-Transformer in E6 was the strongest non-LaRGO-Net competitor, with mean RMSE = 0.054, mean MAE = 0.038, mean
, mean F1 = 98.69%, and mean AUC = 99.09%.
The progression from E7 to E11 demonstrates the architectural value of LaRGO-Net before the addition of the full physics-aware objective. E7, the compact LaRGO-Net with , , mean pooling, dropout = 0.20, learning rate , and adaptive -based node/global embedding, achieved mean RMSE = 0.137, mean MAE = 0.105, mean , mean F1 = 94.48%, and mean AUC = 96.70%. Although this is weaker than the stronger transformer baselines, it shows the initial benefit of adaptive interlayer learning. Increasing the graph depth and latent dimension in E8, using , , residual blocks, LayerNorm, and shared fusion, reduced mean RMSE to 0.111 and improved mean to 0.952. The reliability block improved from RMSE = 0.118, MAE = 0.091, , F1 = 94.87%, and AUC = 97.02% in E7 to RMSE = 0.096, MAE = 0.072, , F1 = 96.18%, and AUC = 98.01% in E8. E9 further increased propagation depth to , reducing mean RMSE to 0.088 and increasing mean to 0.970, while the reliability block reached RMSE = 0.074, MAE = 0.055, , F1 = 97.24%, and AUC = 98.64%. Moving to E10, the wider attention-based version with , , attention pooling, and dense fusion reduced mean RMSE to 0.074 and improved mean F1 to 97.83%. This indicates that attention pooling helps the model assign nonuniform importance to pavement layers rather than treating all node embeddings equally. E11, the deep-wide variant with , , attention pooling, GELU activation, dropout = 0.35, and learning rate , achieved mean RMSE = 0.060, mean MAE = 0.043, mean , mean F1 = 98.46%, and mean AUC = 99.04%. This result is close to the GAT and FT-Transformer baselines, but it still benefits from the physically ordered pavement graph and adaptive layer-wise representation rather than generic attention alone.
The final improvement is obtained when the LaRGO-Net architecture is combined with physics-aware and coupling-consistency objectives. E12 adds to the regression and classification losses, using , , attention pooling, dropout = 0.35, learning rate , batch size 48, and 320 epochs. This reduced the mean RMSE from 0.060 in E11 to 0.048, reduced mean MAE from 0.043 to 0.034, and increased mean from 0.986 to 0.991. At the reliability-block level, E12 achieved RMSE = 0.036, MAE = 0.025, , F1 = 99.08%, and AUC = 99.37%, showing that physical-consistency regularization improves both continuous response prediction and reliability classification. The full LaRGO-Net in E13 further adds , uses dropout = 0.40, learning rate , batch size 32, 350 epochs, AdamW, weight decay , and early stopping. It achieved the best overall results, with mean RMSE = 0.040, mean MAE = 0.028, mean , mean F1 = 99.10%, and mean AUC = 99.45%. Compared with the strongest non-LaRGO-Net competitor, FT-Transformer in E6, the full model reduced mean RMSE from 0.054 to 0.040, corresponding to a 25.9% reduction, and reduced mean MAE from 0.038 to 0.028, corresponding to a 26.3% reduction. The reliability block of E13 reached RMSE = 0.029, MAE = 0.020, , F1 = 99.21%, and AUC = 99.53%, which is higher than E6 reliability performance of RMSE = 0.043, MAE = 0.030, , F1 = 98.81%, and AUC = 99.16%. These results indicate that LaRGO-Net’s superiority is not simply due to using a deeper or wider model, because strong alternatives such as GAT and FT-Transformer were included. Instead, the improvement is linked to the combination of pavement–layer graph representation, adaptive interlayer coupling, global scenario conditioning, task-aware prediction, physics-consistency loss, and coupling-consistency loss, which together allow the model to better capture structural, thermal, hydraulic, energy-harvesting, and reliability interactions in the renewable energy-integrated pavement system.
The final step from E8 to E9 provides the strongest evidence that the true novelty of the proposed framework lies not only in graph depth or latent width, but in the integration of physics-aware and coupling-aware optimization with the graph-operator backbone. E8 retains the strong architectural setting of , , attention pooling, and dropout 0.35, but it introduces the physics-consistency objective through with , trained using AdamW at for 320 epochs. This reduces the mean RMSE to 0.048 and raises the mean to 0.991. The structural block reaches RMSE = 0.051 and , the thermal block RMSE = 0.046 and , and the reliability block RMSE = 0.036, MAE = 0.025, , F1 = 99.08, and AUC = 99.37. The full E9 configuration adds the coupling-consistency objective, stronger dropout of 0.40, a lower learning rate of , smaller batch size 32, 350 training epochs, AdamW with weight decay , and early stopping, thereby producing the best results in every domain. Specifically, E9 achieves structural RMSE = 0.043 with , thermal RMSE = 0.038 with , hydraulic RMSE = 0.048 with , energy RMSE = 0.044 with , and reliability RMSE = 0.029, MAE = 0.020, , F1 = 99.21, and AUC = 99.53. Its mean RMSE of 0.040 and mean of 0.994 make it the most accurate and most stable configuration across all five output blocks. From a technical standpoint, this indicates that the physics-consistency term improves local plausibility of individual targets, while the coupling-consistency term strengthens coherence among mutually dependent outputs such as tensile strain, moisture penetration, energy-generation response, and reliability degradation.
Table 10 presents a fair comparison between the proposed LaRGO-Net and representative tabular, neural-network, and graph-neural-network baselines using the same training, validation, and test partitions. Among the tabular models, XGBoost achieved the strongest baseline performance, with RMSE = 0.091, MAE = 0.067,
, F1 = 96.82, and AUC = 98.34, showing that boosted trees can capture nonlinear interactions among the simulation-generated variables more effectively than Random Forest. However, both Random Forest and XGBoost still treat the pavement response as a flat tabular prediction problem and do not explicitly preserve the seven-layer pavement hierarchy or the directed transfer of load, heat, moisture, and energy effects across the pavement depth.
For the neural and graph-based baselines, the MLP produced the weakest performance because all variables were processed as an unstructured feature vector. The static GCN improved over the MLP, confirming that graph representation is useful for this problem, but its fixed adjacency and lack of adaptive coupling limited its ability to represent changing interlayer interactions under different loading and environmental conditions. GAT and GraphSAGE achieved stronger performance than GCN, with GAT reaching RMSE = 0.103 and , due to attention-based or neighborhood-aggregation mechanisms. Nevertheless, these graph baselines still remained less accurate than the proposed model because they do not incorporate the full layer-coupled operator structure, explicit node/global feature separation, physics-consistency loss, and coupling-consistency loss.
The full LaRGO-Net achieved the best overall performance, with RMSE = 0.040, MAE = 0.028, , F1 = 99.21, and AUC = 99.53. Compared with the strongest conventional baseline, XGBoost, LaRGO-Net reduced RMSE by approximately 56.0% and MAE by approximately 58.2%, while also improving , F1, and AUC. This confirms that the performance gain is not only due to nonlinear learning capacity, but also to the physically structured graph representation, adaptive interlayer coupling, attention-based readout, and physics/coupling-aware multi-task optimization.
Table 11 evaluates whether the final LaRGO-Net configuration remains stable under severe operating conditions rather than only under the random test split. The extreme-condition hold-out subset included cases combining high thermal loading, high moisture exposure, high wheel load or contact pressure, and low pavement modulus values. Compared with the random test subset, the hold-out subset produced a moderate increase in mean RMSE from 0.040 to 0.059 and a decrease in mean
from 0.994 to 0.987. The reliability block also decreased from F1 = 99.21 and AUC = 99.53 to F1 = 97.84 and AUC = 98.76. This reduction is expected because the hold-out subset represents more severe and less frequently occurring combinations of thermo-hydro-mechanical loading conditions. Nevertheless, the model retained high predictive consistency across structural, thermal, hydraulic, energy, and reliability outputs, indicating that its performance was not driven only by interpolation among similar randomly split samples.
Figure 5 shows the prediction quality of the E1 baseline multilayer perceptron for ten selected responses, and sets the baseline for the study. One of the notable visual features of this figure is the comparatively large spread of the predicted points away from the 45-degree line in the directions of both prediction over- and under-estimation, which signals that the E1 baseline model can capture the overall trend of the responses generated by the simulation model but is unable to sustain their subtle interdependencies and local nonlinearities. This is particularly evident in the structural and moisture hydrologic domains, where the Max Vertical Deflection
, Moisture Penetration Depth
, and Pore Pressure Max
show relatively large vertical scatter and deviation from the ideal line, meaning that dense flat learning cannot represent the coupled effects of layer stiffness, thermal and moisture transport, and embedded energy-functional behavior. Even the relatively good predictions, like Peak Surface Temperature
and Reliability Index
, still show some degree of scatter, so the model defines the general envelope of the system’s response but not the full physics-informed precision of the system at hand. The energy-related responses including TEG Voltage Output
and Piezoelectric Power
also indicate that the baseline lacks the inductive bias to represent the secondary electromechanical response to coupled loading. Hence,
Figure 5 reveals that the E1 baseline model can only give a coarse statistical representation of the target space and is inherently constrained by its lack of inductive bias to represent the layered multiphysics nature of the renewable energy-integrated pavement system.
Figure 6 shows the prediction quality of the E4 LaRGO-Net baseline configuration and represents a significant leap from tabular learning to physics-informed graph-based learning. Compared with the baseline plot, the points are now much closer to the identity line in almost all output families, which means that the model is now better calibrated, has lower variance and better preserves the causal relations in the simulation data. This is particularly evident in the structural family, where Max Vertical Deflection and Max Von Mises Stress make the leap to
and
, respectively, indicating that once the pavement is modeled as a layered graph with node- and graph-wise feature encoding, the model is significantly improved in terms of capturing of cross-layer stress transfer and deformation. This improvement is also seen in the thermal predictions, where Peak Surface Temperature
and Average Temperature Gradient
indicate a better learning of temperature gradients across the pavement depth to the thermoelectric layer. The improvement is also clear in the hydraulics, with Moisture Penetration Depth
and Pore Pressure Max
suggesting that the graph operator begins to learn how support effects from the lower layers and moisture-related transport processes interact with the pavement response. The energy and reliability targets show equally important gains, with TEG Voltage Output
, Piezoelectric Power
, Reliability Index
, and Failure Probability
. Overall,
Figure 6 shows that even the original LaRGO-Net configuration already offers a different learning regime to the MLP baseline, as it begins to capture the pavement not as an arbitrary set of variables, but in terms of an interacting multilayer infrastructure system whose responses are evolving through structured thermo-hydro-mechanical and energy-functional interactions.
Figure 7 represents the prediction performance of the final E9 configuration, which testifies to the final convergence of the proposed LaRGO-Net model, where multilayered graph propagation, latent space, attention-based aggregation, and physics-informed optimization, all contribute to an almost-perfect match between measured and predicted values for all of the representative outputs. Compared to
Figure 5 and
Figure 6, the point clouds are significantly tighter, more uniformly distributed around the 45-degree line and less scattered at higher output magnitudes, showing not only improved goodness-of-fit, but also greater stability and reduced heteroscedasticity over the full range of responses. This is evident in the structural response, where Max Vertical Deflection
and Max Von Mises Stress
demonstrate the full model captures the fine-scale load-induced response propagation through the multilayered pavement structure with great accuracy. The thermal responses also remain extremely strong, as seen with Peak Surface Temperature
and Average Temperature Gradient
, which is critical because these responses control thermoelectric activation. The “most uncertain” hydraulic outputs also become extremely stable, withMoisture Penetration Depth
and Pore Pressure Max
, showing that the model now captures support-layer and diffusion processes much more accurately than either the baseline or base graph. The energy outputs, TEG Voltage Output
and Piezoelectric Power
, show that the proposed model does not just predict traditional engineering responses, but can also capture the embedded harvesting mechanism with high precision. Finally, the reliability-related outputs perhaps show the greatest degree of success, with Reliability Index
at
and Failure Probability
, demonstrating that the full model can map pavement responses to highly precise safety-performance indicators. If
Figure 5,
Figure 6 and
Figure 7 are viewed as a progression from simple flat learning to informed graph learning, to full intelligent prediction with reliability awareness, then they visually support the central claim of the study that graph encoding and physics-based constrained optimization is critical for accurate multiphysics prediction in renewable energy-integrated pavement systems.
Table 12 presents a family-wise summary of regression performance of the best model (i.e., full E9 LaRGO-Net) and shows that the proposed model exhibits high accuracy across all five families of predictions rather than in just a single family. The model has RMSE = 0.043, MAE = 0.030, and
in the structural family (deflection, deformation, stress and strain) which suggests the precise modeling of the structural response and deformation of the layered pavement structure. The thermal family is even slightly better, with RMSE = 0.038, MAE = 0.026, and
, which suggests that the graph-based architecture is able to capture the temperature gradient with depth of the pavement and the internal gradients that activate the thermoelectric layer. The hydraulic family is still the most difficult to model with RMSE = 0.048, MAE = 0.034, and
, but still extremely good in terms of prediction, and this also shows that the moisture-sensitive effects, pore-pressure and penetration dynamics are still uncertain due to their higher dependence on support-layer interaction and environmental variations. The renewable-energy family of thermoelectric and piezoelectric voltage, power, and efficiency of the energy-harvesting system achieves RMSE = 0.044, MAE = 0.031, and
, which shows that the proposed architecture is capable of accurately capturing the functional energy-harvesting response while maintaining the prediction performance for mechanical and environmental responses. The continuous reliability family has the best performance with RMSE = 0.029, MAE = 0.020, and
for the prediction of
and
, which also features the smallest RMSE and the highest percentages of improvement over the baseline models, i.e., 83.24% versus E1 and 80.54% versus E2. These percentages are statistically meaningful in that they show that the full LaRGO-Net not only increases the prediction of response variables, but it actually increases in value as the prediction target moves towards decision variables for safety-critical design. In addition, the percentage improvement is quite high across all families, ranging from 76.12% to 83.24% against E1, and from 72.88% to 80.54% against E2, which also suggests that the benefits of the proposed graph-based, physics-aware learning framework are both widespread and even.
Figure 8 offers an aggregated quantitative assessment of regression prediction performance for the three major model versions, that is, the E1 baseline MLP, E4 LaRGO-Net base, and E9 full LaRGO-Net models; it thus provides a visual overview of the performance improvement for the proposed framework from traditional flat learning to physically structured graph-based learning. The monotonic rise of
in panel (a) for all ten representative targets confirms that each stage of architectural development enhances the variance explained by the model; this is particularly noticeable for targets that are sensitive to hydraulic interaction or reliability concerns such as Moisture Penetration Depth, Pore Pressure Max, and Failure Probability
, where the improvement from E1 to E4 is already significant, and the final step to E9 brings predictions close to unity. Panel (b) confirms this finding by displaying a monotonic decrease in RMSE across all outputs, meaning that the increased explanation of variance is coupled with a simultaneous decrease in absolute error. The largest drops are observed for variables that are the most difficult to learn due to their coupled multiphysics nature, namely, Moisture Penetration Depth, Pore Pressure Max, and Piezoelectric Power, which indicates that the graph-based architecture is increasingly superior as the target is more dependent on cross-layer interaction and support-condition sensitivity. Panel (c) reinforces this insight via the normalized MAE, which eliminates some of the scale bias and proves that the quality of E9 is not a mere effect of metric scale, but an indication of relative predictive accuracy with respect to diverse targets with different physical units and magnitudes. Interpreting the three panels in concert, we observe a uniform and technically meaningful trend: E1 has the lowest variance explanation and largest errors; E4 provides a significant improvement through graph encoding, residual operator learning, and shared latent encoding; and E9 strikes the best balance with a deeper propagation, attention pooling, and physical optimization.
Table 13 shows that the proposed full LaRGO-Net brings a decisive and very consistent gain in classification performance for reliability-oriented decision making, compared to the baseline dense MLP and the baseline shallow GCN, in all three decision targets, namely, Safety_Class, Performance_State, and Damage_Risk_Level. The flat MLP model for binary Safety_Class delivers 91.22% accuracy, 90.53% precision, 91.16% recall, 90.84 F1-score, 94.12 AUC, and 0.824 kappa, showing that the model has a fair separability but a substantial risk of misclassification in safety-critical applications. The graph structure in E2 leads to 93.04% accuracy, 92.41% precision, 92.88% recall, 92.63 F1-score, 95.84 AUC and 0.861 kappa, showing that even static graph structures are more effective in representing the interdependence of the pavement layers than dense learning. On the other hand, the complete E9 LaRGO-Net boosts this target to 99.37% accuracy, 99.18% precision, 99.25% recall, 99.21 F1-score, 99.53 AUC, and 0.987 kappa, which suggests near-perfect discrimination and an extremely high agreement beyond chance. This trend also holds for the multiclass Performance_State target, where the model improves from 89.74% accuracy and 89.92 F1-score in E1 to 91.86%, 91.98 in E2, and then to 98.94% accuracy and 98.95 F1-score in E9, while kappa improves from 0.792 to 0.972, which indicates that the proposed architecture is highly efficient for discriminating the safe, marginal, and unsafe operating states. The same monotonic improvement occurs to the Damage_Risk_Level target, from 90.16% accuracy, 90.27 F1-score and 0.801 kappa in E1 to 92.31%, 92.46 and 0.846 in E2, and finally to 99.08%, 99.07 and 0.975 in E9, showing that the model can perform accurate qualitative risk grouping based on coupled multiphysics response. These findings are technically important because they demonstrate that the improvement brought by LaRGO-Net is not limited to one particular reliability label, but is uniformly spread across binary safety classification, multiclass performance evaluation, and damage-based risk classification. From an engineering perspective, this confirms that graph-based encoding, interlayer adaptive coupling, attention-based aggregation and physics-informed training all contribute to enhance the model’s capacity to preserve causal links between structural damage, water vulnerability, thermal stress, renewable-energy response, and reliability assessment.
Figure 9 shows the distribution view of the prediction quality across the nine experiments, and it is important because it complements the summary statistics by offering a view of the distribution of the absolute prediction error over the entire test domain. A monotonic pattern is evident in how E1 through E9 distributions become progressively more concentrated around zero, their right tails become successively shorter, and the values of the mean and median (normalized) absolute errors drop in a highly systematic fashion. We start with the E1 baseline MLP which has the widest and most right-skewed distribution whose mean error
and median
, suggesting not only a higher absolute error, but also a large number of moderate to large errors. E2 and E3 drop the values to
and
, respectively, indicating that the addition of graph and then adaptive coupling already moves the error mass to lower values. The jumps to E4 and E5, with mean/median values of
and
, show that the introduction of residual learning of graph operators and deeper interactions between layers not only tighten the average error, but also the overall shape of the error distribution. E6 and E7 show a further concentration of the distributions toward zero, with the mean error reducing to 0.037 and 0.031, respectively, revealing the role of attention-based pooling and wider/deeper latent representations in reducing the number of medium errors. In the final two panels, E8 and E9, we see the most tightly concentrated and left-skewed distributions, with E8 having mean/median
and E9 giving the best overall results of
and
. These last panels are especially noteworthy because the very short right tails show that physics-consistency and coupling-consistency regularization not only reduce central tendency but also increase resilience to large, outlier-like errors. Analytically, the figure demonstrates that the proposed LaRGO-Net does not merely reduce mean error in a restricted set of cases, but progressively transforms the whole error landscape into higher concentration, lower spread, and greater stability as the network architecture better conforms to physics and as the optimization algorithm becomes more constrained by engineering principles.
Table 14 provides a component-level verification of LaRGO-Net by isolating the contribution of graph structure, adaptive interlayer coupling, global context injection, residual propagation, attention pooling, physics-aware losses, and multi-task prediction. The complete model in A1 achieved the strongest overall performance, with RMSE = 0.040,
, Macro-F1 = 99.21%, and a training time of 105.5 min. The most severe degradation occurred in A12, where both graph structure and adaptive interlayer coupling were removed simultaneously; RMSE increased to 0.226,
decreased to 0.812, and Macro-F1 dropped to 88.97%, corresponding to ΔRMSE = +0.186 and a Macro-F1 loss of 10.24 percentage points. Although this simplified configuration reduced training time to 34.8 min, the large accuracy loss confirms that the computational saving was achieved at the expense of removing the physical layer organization required to represent the seven-layer pavement system. The individual ablations support the same conclusion: removing graph encoding alone in A2 increased RMSE to 0.194 and reduced Macro-F1 to 90.84%, while removing adaptive interlayer coupling in A3 increased RMSE to 0.151 and reduced Macro-F1 to 94.52%. Therefore, the graph topology and learned
coefficients are both essential, because the model must preserve the ordered pavement hierarchy and learn how interlayer influence changes under different mechanical, thermal, hydraulic, and energy-harvesting conditions.
The results also demonstrate the importance of context-conditioned and multi-task learning. Removing global context injection in A8 increased RMSE from 0.040 to 0.083, reduced from 0.994 to 0.974, and decreased Macro-F1 from 99.21% to 97.62%, showing that pavement–layer interactions cannot be interpreted independently of scenario-level variables. This is important because the same pavement configuration may behave differently under wheel loads of 35–80 kN, contact pressures of 0.60–0.90 MPa, load speeds of 20–100 km/h, solar heat flux of 500–1000 W/m2, ambient temperatures of 15–45 °C, thermal gradients of 8–30 °C, rainfall infiltration of 0–20 mm/h, and moisture flux of 0.005–0.030. The multi-task ablations further verify that the proposed prediction strategy is not merely an architectural addition. In A9, replacing the separate structural, thermal, hydraulic, energy-harvesting, and reliability prediction heads with one shared output head increased RMSE to 0.073 and reduced Macro-F1 to 98.01%. In A10, reliability-only single-task training further increased RMSE to 0.088 and reduced Macro-F1 to 97.34%. These results indicate that reliability prediction benefits from simultaneously learning intermediate physical responses, including deformation, stress, temperature, moisture response, voltage, and power, because reliability states emerge from the interaction of these coupled response domains rather than from isolated reliability labels alone.
The remaining ablations clarify the roles of physics-aware regularization and architectural stabilization. Removing the physical-consistency loss in A6 increased RMSE to 0.052, reduced to 0.990, and lowered Macro-F1 to 98.96%, while removing the coupling-consistency loss in A7 increased RMSE to 0.047, reduced to 0.992, and lowered Macro-F1 to 99.08%. When both and were removed together in A11, the degradation became more evident, with RMSE = 0.064, , Macro-F1 = 98.23%, and ΔRMSE = +0.024. This confirms that the two loss terms are complementary: discourages physically inconsistent trends such as reduced deflection under higher wheel loading or reduced thermoelectric voltage under larger thermal gradients, whereas preserves agreement among related outputs such as tensile strain, failure probability, safety class, and damage-risk level. In addition, removing residual connections in A4 increased RMSE to 0.097 and reduced Macro-F1 to 97.03%, while replacing attention pooling with mean pooling in A5 increased RMSE to 0.061 and reduced Macro-F1 to 98.44%. Overall, the ablation results confirm that the final LaRGO-Net performance is produced by the combined effect of graph-based physical structuring, adaptive interlayer reasoning, scenario-conditioned message passing, task-specific multi-output prediction, and physics-aware regularization rather than by model size alone.
Figure 10 offers a family-wise residual-distribution view of the evolution from the most basic baseline models to the full LaRGO-Net and is particularly telling because it shows not only the error magnitude, but also the error centering, spread, and balance across the five output families. In the top row, representing E1, E2, and E3, the boxplots are clearly wider, the interquartile ranges are longer, and the whiskers and outlier ranges are stretched away from the zero-residual line, which shows there is much prediction variability, and lower calibration, in the structural, thermal, hydraulic, energy, and reliability families. The variability is especially high in the hydraulic and energy families in these early configurations, with greater residual spread and tail skew, which makes sense given the greater nonlinearity and multiphysics coupling of moisture-transport and energy-harvesting variables. In the middle row, E4 to E6, the distributions start to shrink around the zero line, the medians are more centered, and the boxes shrink, revealing that as the residual graph operator learning is added, the propagation depth is increased, and attention-based pooling is applied, there is a progressive reduction in systematic bias and diminishment in medium-scale error variability across all five response families. This constriction is particularly significant because it happens in all five families, which suggests that the graph-based architecture is not only improving some variables at the expense of others, but is delivering a coordinated improvement in the calibration of structurally and physically diverse targets. The bottom row, for E7, E8, and E9, shows the most compact, most centered, and most unbiased distributions of residuals, with very short interquartile ranges, whiskers, and noticeable contraction of the outlier distribution, especially in the structural, thermal and reliability families. The E9 full LaRGO-Net exhibits the most compact and most centered distributions, which suggests that the final combination of deeper and wider graph propagation, attention-based latent aggregation, physics-consistency regularization, and coupling-consistency constraints not only leads to lower average error across targets, but also to statistically more stable and less systematic predictions in the full target space. The figure is important from an analytical point of view because it shows that the evolution from E1 to E9 reflects a genuine improvement in the residual structure: the model’s average accuracy ascends while becoming more balanced in terms of prediction accuracy across the different families of responses, and it becomes more centered around a physically plausible prediction behavior.
Table 15 shows that the top-performing E9 configuration exhibits a strong degree of model stability under increasingly critical uncertainty-sensitive operating conditions and that the influence of different uncertainty sources on model performance is revealing. For all three types of conditions, the low-severity bands demonstrate the best overall performance as measured by RMSE (0.031, 0.033, and 0.030 for traffic growth, moisture variation, and temperature variation, respectively),
(0.996, 0.995, and 0.996, respectively) and F1 (greater than 99.27). This suggests that as long as the variability in operability is small, the LaRGO-Net model maintains near-perfect regression and classification accuracy for continuous reliability targets and discrete safety-related decision-making. From low to medium severity, a controlled and nearly uniform degradation pattern is observed, with RMSE values of 0.039, 0.041, and 0.038 and F1-scores of 99.08, 98.96, and 99.02 for traffic growth, moisture variation, and temperature variation, respectively, and the mean
error is maintained in the small range 0.021–0.024, while the mean
error is below
. This implies that the model maintains a high level of generalization under moderate uncertainty growth and that the graph-based latent space remains effective in capturing the predominant coupled relationships between the structural, thermal, hydraulic and energy-functional variables. The most interesting pattern emerges in the high-severity level, where there is a more significant loss of performance and more variability among different uncertainty sources. Under high traffic growth, the model records RMSE = 0.051, MAE = 0.036,
, F1 = 98.74, mean
error = 0.031, and mean
error =
, indicating that intensified loading remains challenging but still manageable for the model. High temperature variation induces a similar but less severe situation, with RMSE = 0.053, MAE = 0.037,
, F1 = 98.63, mean
error = 0.033, and mean
error =
, because increased thermal variability reduces model accuracy but only in a minor way. High moisture variation is undoubtedly the most challenging case, where RMSE increases to 0.057, MAE to 0.040,
decreases to 0.988, F1 to 98.48, mean
error to 0.036 and mean
error to
, which are the lowest values in the table. This is technically important because it suggests that moisture variation is still the most disruptive source of performance degradation, which is entirely consistent with the coupled thermo-hydro-mechanical (THM) nature of the proposed pavement system, where water ingress changes support conditions, affects stiffness redistribution, modifies pore-pressure evolution, and indirectly changes both structural response and reliability.
Figure 11 provides a dynamic perspective on the development of the baseline to the best-performing LaRGO-Net, and it shows that the convergence patterns reflect the improvements in the regression and reliability tables. In panel (a), the E1 baseline MLP has a large initial RMSE of both the training and validation curves, a relatively slow convergence, and a significant gap between the training and validation curves, with early stopping at epoch 107 of 120, which indicates low representational efficiency and moderate potential to further reduce the generalization error during training. In panel (b), the E4 LaRGO base model, we see a major shift towards well-behaved convergence: the RMSE values of both curves now converge more rapidly, the validation curves are smoother, and the gap between the curves has decreased at the early stopping point (around epoch 186 of 220), which suggests that not only does the graph encoding and the residual operator learning (which is also part of the latent fusion) improve the quality of the optimization but also the stability of the generalization behavior. In panel (c), the E8 network (LaRGO + Physics) reveals another shift towards desirable convergence where the RMSE values of the training and validation curves converge quickly to a low RMSE regime, and then converge strongly around a minimum value with early stopping at epoch 207 of 320; this is interesting because the physics-consistency loss not only improves the accuracy of the predictions but also stabilizes the optimization landscape early in the training process by preventing unphysical drift in late epochs. The optimal pattern is seen in panel (d) for the E9 LaRGO-Net, where the training and validation curves go down to the lowest RMSE range in the figure, have a very stable late-stage training trajectory, and a very small residual gap at early stopping (epoch 314 of 350). This final pattern is significant because it confirms that deeper graph propagation, attention-based aggregation, physics-consistency, coupling-consistency, AdamW optimization and early stopping not only result in the best final regression and reliability metrics but they also produce the most optimal convergence dynamics, which include fast error reduction, lack of jitter at late epochs and stability of the validation curve.
Figure 12 supplements the RMSE-based convergence analysis by displaying the evolution of the explanatory power of the models during training and validation for four representative configurations, which not only reflects the final prediction quality but also the stability and the level of maturity of the learning process across the model hierarchy. As shown in panel (a), the E1 baseline MLP has the lowest and slowest
evolution, with the training
curve starting from around 0.50 and progressing to around 0.83, while the validation
stabilizes to about 0.78 at early stopping (early stopping happens at around epoch 107), and there is a significant gap between the two curves; this suggests that the baseline model can learn a general response trend but lacks the ability to achieve strong explanatory power and consistency, as well as the capacity to effectively represent the pavement system. In panel (b), we observe a much better convergence for the E4 LaRGO base model, with both training and validation
rising more steeply and reaching the mid-0.94 range at early stopping around epoch 186, highlighting that graph encoding, residual graph operator learning, and shared latent fusion largely boost the explanatory power and validation stability. In panel (c), E8 LaRGO + Physics shows an even better convergence pattern in which the two curves rapidly enter the high-
regime and then remain very close to 0.99 after roughly half of the training time, implying that the physics-consistency objective improves not only the final predictive quality but also the convergence dynamics by encouraging the model to learn physically plausible and more stable generalization. Panel (d), corresponding to the full E9 LaRGO-Net, shows the best convergence behavior, where the training and validation
curves rapidly increase in the early epochs, are almost perfectly overlapped over a long training window, and converge very close to 0.99 at early stopping around epoch 314, revealing almost perfect explanatory power with negligible generalization gap. This last pattern is of analytical significance as it indicates that the combination of deeper graph propagation, attentional pooling, physics consistency, coupling consistency, and disciplined training not only achieves the best final prediction performance, but also the best overall learning dynamics, with fast convergence, high validation stability, and no overfitting.
Table 16 demonstrates that the benefits in predictive accuracy provided by LaRGO-Net are delivered with a moderate and acceptable increase in the computational cost of the model, hence demonstrating a good trade-off between predictive ability and implementation efficiency. E1, the baseline MLP, is the smallest model with only 0.31 million parameters, 412 MB of GPU memory, 4.8 s per epoch in training, and 0.19 ms per case in inference, but the performance score of 86.4 tells us that this is achieved at the cost of poor awareness of structural interactions and poor predictive fidelity. The baseline GCN in E2 slightly augments the model size to 0.58 million parameters and 568 MB of memory and training and inference times to 7.2 s and 0.32 ms, but the performance score increases to 89.7, which indicates that even a basic graph representation provides a substantial gain for a moderate increase in computational cost. The jump to the LaRGO-Net models is more informative: E4, the base graph-operator model, requires 1.42 million parameters and 1096 MB of memory, training time of 10.9 s per epoch, and inference time of 0.47 ms, but the performance score of 95.2 indicates that the added complexity is well justified by the significant improvement in predictive accuracy. E7, the deep-wide version, increases the backbone complexity to 3.86 million parameters and 1848 MB of memory, with training and inference costs of 16.4 s and 0.73 ms, respectively, but it improves the performance score to 98.1, indicating that the wider latent representation and deeper graph-operator backbone provide a strong fidelity gain without making the model prohibitively large. The full E9 model only slightly expands the E7 model in terms of model size (3.94 million parameters, 1926 MB memory), training time (18.1 s per epoch) and inference time (0.81 ms), yet it increases the performance score to 99.1. This modest increase in computational burden, coupled with the best possible predictive performance, is particularly significant because it demonstrates that the final gains of E9 are not due to simply increasing the model size, but the inclusion of physics-consistency, coupling-consistency, and more robust optimization schemes (AdamW, weight decay, and early stopping).
Figure 13 offers a spatially-distributed multiphysics interpretation of the proposed renewable energy-integrated pavement system by simultaneously depicting the temperature, moisture ingress, stress, strain, and vertical deformation fields against the final seven-layer pavement configuration, thus revealing the depth-wise evolution of the governing responses and their interaction with the functional layers. Panel (a) shows the temperature field and reveals a surface-dominated thermal field under solar heat flux, with the hottest region being in the upper asphalt and binder courses; the temperature field then progressively decreases with depth towards the subgrade, which demonstrates the existence of a persistent vertical thermal gradient across the pavement depth and directly motivates embedding the thermoelectric layer at an intermediate position where significant thermal gradients can be maintained. Panel (b) reveals the moisture-penetration field under rainfall infiltration, where the upper ingress zone gradually spreads through depth and lateral diffusion, resulting in a nonuniform hydraulic front that is more pronounced in the lower support courses; this is important because it indicates that the moisture-ingress zone is not only present on the surface but also penetrates toward the subbase and subgrade, thereby affecting support sensitivity, pore-pressure build up, and reliability-related damage. The stress–strain fields in panels (c) and (d) also show that the peak von Mises stress and principal strain occur in the immediate vicinity of the wheel-load zone and then decay with an increase in depth, but they still remain identifiable in the piezoelectric and thermoelectric insert zone, which suggests that the upper support courses absorb the direct load effect, whereas the intermediate functional zone remains sufficiently active to support energy harvesting, yet is not disconnected from the global stress transfer. This is complemented by the vertical deformation field in panel (e) where the maximum vertical displacement is concentrated below the tire contact zone and then attenuates through the base, subbase, and subgrade, revealing the role of the structural layers in spreading the load, and the downward flow of deformation-sensitive response towards the embedded insert region. Finally, panel (f) shows the layered structure of the pavement and enables the preceding field contours to be interpreted in terms of the physical structure, thus confirming that the thermoelectric layer and piezoelectric insert zone are not simply randomly positioned components, but they are instead embedded within the depth range where the thermal gradient, load-induced stress, and vertical deformation remain significant for coupled interactions. Overall, the figure confirms that the proposed pavement is indeed a fully thermo-hydro-mechanical infrastructure system, where heat conduction, moisture diffusion, loading-induced stress, strain localization and vertical deformation are all geometrically coupled and it is this field interdependency that justifies the use of simulation-based data generation and graph-based reliability-aware prediction in the rest of the investigation.
Nadeau–Bengio Corrected Resampled t-Test for Train–Test Generalization and Data-Leakage Assessment
The statistical generalization analysis in
Table 17 provides additional evidence that the high prediction performance of LaRGO-Net was not restricted to the training subset. Under the repeated 70/30 holdout evaluation, 70% of the 6000 simulation cases were used for training and 30% were reserved for testing in each repetition, while the Nadeau–Bengio corrected resampled paired
t-test was used to account for the dependence introduced by repeated data splits. The regression results show only minor train–test differences: RMSE increased from
during training to
during testing, with a mean gap of only
, corrected standard error of 0.00345, corrected
, and
. Similarly, MAE increased slightly from
to
, with a gap of
, corrected
, and
, while
decreased only from
to
, with a gap of
, corrected
, and
. Since all regression
p-values are greater than 0.05 and the 95% confidence intervals of the gaps include zero, the small differences between training and testing errors are not statistically significant. This supports the interpretation that the model retained its predictive accuracy on unseen simulation cases rather than simply memorizing the training records.
The classification results also demonstrate stable generalization between the training and testing stages. The testing accuracy remained very close to the training value, decreasing only from to , with a mean gap of , corrected , and . Precision decreased from to , recall decreased from to , and F1-score decreased from to . These changes correspond to small gaps of only , , and , respectively, with corrected p-values of 0.421, 0.428, and 0.379. The AUC also remained highly stable, decreasing only from to , with a gap of , corrected , and . The effect sizes remained small across all reported metrics, ranging approximately from 0.13 to 0.20, indicating that the observed train–test differences were not practically large. These findings provide statistical support that the reported classification performance was not the result of direct leakage from reliability-derived labels or finite element response outputs into the input feature matrix.
The high performance should be interpreted carefully because the dataset was generated from a controlled finite element simulation environment rather than from noisy field measurements. The leakage-control protocol reduces the risk of over-optimistic evaluation because each complete Abaqus simulation case was assigned to either the training or testing subset, normalization parameters were fitted using the training subset only, and target-derived quantities such as limit-state margins, reliability index , failure probability , safety class, performance state, and damage-risk level were excluded from the input features. However, the simulation setting still contains idealizations that may contribute to the strong performance, including controlled boundary conditions, structured parameter ranges, equivalent energy-output postprocessing, sequential rather than fully monolithic THM coupling, and limited representation of field uncertainties such as aging, construction defects, traffic wander, seasonal material degradation, and measurement noise. Therefore, the results demonstrate strong internal generalization within the 6000-case, 68-variable simulation-generated scenario space, but they should not be interpreted as unrestricted field-level validation. Future work should extend the evaluation using independent laboratory tests, field-monitoring data, or external finite element datasets generated under different material models, pavement geometries, climatic conditions, and boundary assumptions.