1. Introduction
Sustainable pavement engineering is evolving rapidly due to global environmental concerns, constraints on natural resources, and requirements for enhanced performance. Recent studies [
1,
2] have emphasized the necessity of virgin aggregates and bitumen consumption reduction and the better use of novel modifiers, additives, or recycling methods. However, soaring material prices, the consumption of natural resources, and global warming issues have catalyzed the push to employ bio-based materials, polymers, and industrial byproducts that can provide better mechanical properties together with longer pavement durability [
3,
4]. The use of recycled and waste-based materials embodies the principles of a circular economy. Fibers from textiles [
5], waste glass, plastics [
6], red mud [
7], and incineration ash [
8] have all been seen to improve the stiffness, durability, and rutting resistance of asphalt. Furthermore, because these materials can modify the microstructural characteristics of asphalt binders, functional additives and nanomaterials are becoming more and more popular. While Duan et al. [
9] exposed the aging retardation effects of zinc oxide/vermiculite composites, Hassan et al. [
10] showed the reinforcing potential of nano-CaCO
3 and basalt fibers. In a similar vein, Monticelli et al. [
11] investigated bio-oil-modified heavy polymer–asphalt systems, demonstrating enhanced fracture behavior for mixtures with high recycled asphalt concrete (RAC) content. With parallel developments in warm-mix asphalt technologies [
12], Çakı and Baş [
13] demonstrated notable reductions in emissions and energy consumption, reinforcing the importance of second-generation chemical and nanostructured additives in maximizing sustainability and performance.
AI-assisted pavement design has expanded, particularly for binder content prediction in mixtures containing RAC. Studies by Saleh et al. [
14,
15] and Ghafari et al. [
16] show that neural and hybrid models can accelerate mix design and improve performance forecasting for crumb rubber- and RAP-modified systems. Combining digital modeling with targeted laboratory testing can reduce trial-and-error iterations while delivering cost-effective, durable, and more sustainable pavements.
Predicting bitumen content is necessary because RAC mix design is a high-uncertainty decision problem: aged binder heterogeneity and strongly non-linear interactions among recycled constituents, recycling agents, and modifiers can shift the optimum binder dosage. Small dosing errors can propagate into volumetric imbalances, reduced durability, and variable field performance. In this context, machine learning prediction can reduce iterative laboratory trials and improve repeatability by learning coupled relationships between measured mixture descriptors and the target bitumen content.
Salalah, in the Governorate of Dhofar (Sultanate of Oman), provides a practical setting for recycled asphalt concrete (RAC) because routine pavement rehabilitation yields reclaimed asphalt streams that can be reincorporated into new mixtures. In this study, RAC was recovered from 260 deconstructed pavement samples in Salalah, screened to remove contaminants, and reintroduced at 30% and 50% replacement alongside locally sourced aggregates. However, aged binder heterogeneity and non-linear binder–aggregate interactions make the optimum bitumen content sensitive to blending uncertainty and rejuvenator effectiveness [
1]. This local context strengthens the case for coupling laboratory characterization with the hybrid AI-based prediction of bitumen content, consistent with AI-enabled asphalt design studies that address complex performance behavior and cracking resistance [
16].
This study positions the BAG-CARIMA-LGM hybrid as a layered prediction-and-decision pipeline for RAC mix design, where “modifier” variables represent only one subset of the control feature space. CARIMA provides a time-series-with-controls backbone by selecting model orders (P, d, q) using information criteria and residual diagnostics (AICc/BIC; Ljung–Box) and by applying cross-regime error checks to capture curing and aging drift [
17]. Bagging aggregates bootstrap-trained trees (approximately 100 trees; depth 5–10) to reduce variance and improve robustness under heterogeneous recycling regimes [
18]. Finally, the logistic probabilistic model (LGM) calibrates outputs to physically admissible bounds (0–1) via a logit link estimated by maximum likelihood, supporting uncertainty-aware decision making [
12]. This tripartite design targets temporal non-stationarity, recycling-level heterogeneity, and the need for constrained, decision-grade predictions.
Preliminary research on thiophene-modified binders [
19] and mining waste additives [
20] laid the groundwork for the sustainability-oriented design of mixes. Subsequent studies demonstrated that rejuvenated high-RAP mixtures can achieve the properties of virgin asphalt [
21] when facilitated by polymer or bio-based modification [
22]. This departure from empirical charts towards predictive modeling has allowed AI-based applications to assist with the uncertainties associated with sustainable mix design [
22,
23]. AI hybrids offer novel predictive insights into conventional ensembles, and the functions of rejuvenators and recycling agents in high-RAC mixtures have been well documented. Evaluations of rejuvenator effects on bitumen aging in hot recycled asphalt have shown temperature-dependent efficiencies that can be sequentially modeled by CARIMA components [
22]. The performance of recycled mixtures has been shown to be improved by polymer-modified binders, acting as rejuvenators to manage binder blending uncertainty [
24]. By using LGM localization to inform sustainable predictions, the analysis of bio-recycled asphalt fumes has revealed reductions in organic compounds. Blending charts used to forecast performance in high-RAP scenarios have shown significant improvements, but hybrids outperform them by combining ensemble robustness and data-driven autoregression [
25,
26]. All these studies support the novel advantage of hybrids in comprehensive bitumen prediction for environmentally friendly pavements. The superiority of AI hybrids in cost–benefit analyses is highlighted by technical and economic assessments of recycled asphalt mixtures.
Despite several studies performed regarding RAC, there are vital areas that remain unaddressed. Studies show that very little work has been conducted on the combined effects of bio-modifiers, nanomaterials, and recycled aggregates [
27], and AI predictions have yet to be fully integrated into mechanistic–empirical design methodologies. There is also little evidence on the effects of waste-based fillers in combination with bio-additives, since a balance between durability, environmental benefits, and structural performance is needed. Previous work demonstrates that these hybrid models (bagging (BAG), Controlled Autoregressive Integrated Moving Average (CARIMA), and large geospatial models (LGMs)) are superior to single-model methods in RAP optimization and foamed bitumen design because they can accurately represent non-linear responses [
9,
13,
28]. Unlike other research that has concentrated on discrete modifiers, our study integrates sustainability evaluation within a unified framework of predictive modeling and recycled materials. The current study posits that AI hybrids, such as BAG-CARIMA-LGM hybrids, are superior in predicting RAC samples due to their capacity to integrate ensemble variance reduction with autoregressive and localized modeling, thereby providing foundational elements for developing robust pavement solutions with a high degree of recyclability.
The central question of this research is whether this combination hybrid can successfully predict binder content through a range of RAC mixtures (0%, 30%, 50% RAC) that are themselves heterogeneous and where aging, variability, and non-linearity place strain on conventional models. Moreover, there is lack of integration of AI predictors with mechanistic–empirical thinking and understudied interactions between bio-/nanomodifiers and recycled constituents. Based on a laboratory dataset of 780 samples, including volumetric Marshall property tests, the hybrid consistently outperformed the baseline models by demonstrating novelty and practical necessity. The hybrid approach serves to combine the aging-capturing effects of CARIMA, the variance reduction capabilities of BAG, and the bounded mixture behavior modeling abilities of LGMs, filling the gap between experiment-based testing and intelligent ambitious sustainable-oriented pavement design. This holistic approach contributes to the development of sustainable, durable, high-performance asphalt systems that can meet both present and future infrastructure needs.
2. Methodology
This study follows an experimental–computational approach that uses traditional pavement engineering methods alongside modern machine learning (ML) algorithms to optimize the content of bitumen in asphalt mixtures with recycled asphalt concrete (RAC). The methodology is a hybrid approach between laboratory-based material characterization and data-driven predictive modeling to resolve binder aging, material heterogeneity, and non-linear interactions that are related to recycled materials. The experimental studies were performed based on the ASTM and AASHTO standards, and ML analysis was introduced into Python 3.12 with the help of existing and custom hybrid libraries.
2.1. Materials and Mix Design
RAC was sampled from 260 deconstructed pavement samples in Salalah, Oman and sieved and manually examined to eliminate contaminants. Virgin aggregates (crushed granite, basalt, and limestone) were obtained in local quarries, and the gradation, specific gravity, durability, and cleanliness were in accordance with the ASTM specifications. There were three mixtures of asphalt:
*Mix A: 0% RAC (control, fully virgin materials);
*Mix B: 30% RAC (blended with virgin aggregates);
*Mix C: 50% RAC (high-recycling formulation, incorporating rejuvenators to restore binder properties).
The range of bitumen was 4.552% and that of mineral filler was between 4 and 6% by aggregate weight. Mix C was the only mix that was reinforced with rejuvenators to increase binder compatibility and workability.
2.2. Extraction and Characterization of Binders
To avoid thermal degradation, the aged binder was separated by the centrifuge procedure in trichloroethylene solvent (ASTM D2172) and recovered by means of rotary evaporation (ASTM D5404). The recovered binder was described in terms of viscosity (at 135 °C), stiffness (at 20 °C), penetration, and softening point. These properties were used in the direction of blending modifications using virgin 60/70 penetration-grade bitumen to achieve the desired rheological performance with recycled mixes, as shown in
Table 1.
2.3. Marshall Testing and Preparation of Samples
Marshall specimens (101.6 mm diameter, 63.5 mm height) were prepared using 75 blows per face (ASTM D6926), with at least three specimens per blend. After conditioning for 24 h at 25 °C, the bulk density, air voids, VMA, VFB, Marshall stability, and flow were measured following ASTM D6927 and ASTM D3203. Additional testing included indirect tensile strength (ASTM D6931), moisture susceptibility (AASHTO T283), and thermal sensitivity testing from −12 °C to +40 °C using a universal testing machine. The dataset comprised 780 specimen records (260 per mixture—(0% RAC, 30% RAC, 50% RAC)), each including stability, flow, Gmb, Gmm, VMA, and ITS measurements for machine learning modeling.
2.4. Machine Learning Framework
The optimum content of bitumen at different levels of RAC was predicted using ML models. The applied models were Controlled Autoregressive Integrated Moving Average (CARIMA), Swapped Autoregressive Integrated Moving Average (SARIMA), radial basis function neural networks (RBFs), multilayer perceptron (MLP), bagging (BAG), boosting (BOT), and a hybrid BAG-CARIMA-LGM model. They were implemented with statsmodels (SARIMA/CARIMA), scikit-learn (BAG/BOT), TensorFlow/Keras (ANNs), and bespoke hybrid modules. To assess the model’s performance in terms of its robustness and ability to predict, cross-validation and measures of error were considered. Comprehensively, the combined approach allows the effective optimization of recycled asphalt mixtures and helps to address the goals of the circular economy by using more RAC when it does not negatively affect mechanical performance.
2.5. Machine Learning Modeling for Bitumen Content Prediction
A hybrid machine learning system that involved time-series, ensemble, and neural models was created to forecast the optimum content of bitumen under variability in RAC. While autoregressive modeling was used to characterize the aging of binders, ensemble learning was used to deal with the heterogeneity of materials, and neural networks were used to deal with non-linear interactions, leading to a hybrid BAG-CARIMA-LGM model. The data consisted of 780 samples whose input values were stability, flow, Gmb, VMA, and Gmm, and the target was the bitumen content. BC represents the bitumen content (%) of the paving mixture under the centrifuge method, MS is the Marshall stability (KN), MF represents the Marshall flow (mm), BSG stands for the bulk-specific gravity of compacted asphalt (Gmb), VMA stands for voids in mineral aggregates (%), and MSG represents the maximum specific gravity of the paving mixture (Gmm). Min-max normalization and 80:20 train tests were all performed as part of data preprocessing. Moreover, 5-fold cross-validation of 6 feature configurations was included (
Table 2 and
Supplementary Materials).
2.5.1. Controlled and Swapped Autoregressive Integrated Moving Average Models
An enhanced ARIMA-based model was applied to improve stability and robustness when modeling RAC-affected pavement datasets. Controlled ARIMA (CARIMA) extends classical ARIMA by incorporating exogenous control inputs (e.g., RAC level, compaction, curing, and modifier indicators), enabling the conditioning of time-ordered performance data. The model works on curing-/aging-indexed sequences, and it is formulated as
where
is the response variable,
denotes control inputs, and
represent autoregressive, differencing, and moving-average orders. The stabilization of forecasts in the case of variability caused by RAC and the simulation of rejuvenator effects as external inputs was achieved using Bayesian priors and iterative error correction. CARIMA was also applied in Python with a train test split of 75:25 and showed increased resistance to non-stationarity compared to standard ARIMA. Simultaneously, a SARIMA model was fitted with a swap-based cross-regime validation procedure, in which the parameters obtained at one level of RAC were tested at another to determine how robust they were to changes in composition. The formulation of the underlying SARIMA was not altered, and seasonal extensions allowed modeling of the compaction cycles and aging effects. When used on 780 samples of data, SARIMA selection and probabilistic validation led to an improvement in the accuracy of the prediction of volumetric indicators (Gmm and VMA).
Table 3 summarizes the CARIMA input–output combinations. Bitumen content (BC, %) was measured by the centrifuge method. Marshall stability (MS, kN), Marshall flow (MF, mm), the bulk-specific gravity of compacted mixtures (Gmb), voids in mineral aggregates (VMA, %), and maximum specific gravity (Gmm) were used as predictors.
2.5.2. Radial Basis Function (RBF) Neural Network
The radial basis function (RBF) neural network uses a three-layered neural network that relies on unsupervised and supervised learning to predict non-linear behavior in RAC-modified asphalt mixtures. The Marshall stability, flow, VMA, and Gmm are features in the input layer used to identify localized data clusters. The hidden layer uses unsupervised learning to estimate Gaussian kernel centers and spreads, which allows for reduced overfitting when there is variability due to RAC. The fixed post-clustering first-layer weights are similar to those in k-means-based center selection. Supervised least-squares optimization was used to train the output layer to regress the content of bitumen. This local approximation can be used to treat material heterogeneity [
29]. The model was used in TensorFlow/Keras with 50–100 centers with 1.0–0.1, which supports the sustainable design of asphalt.
2.5.3. Multilayer Perceptron (MLP) Artificial Neural Network Model
The multilayer perceptron (MLP) is a type of feed-forward neural network that can be applied to address non-linear interdependent relationships between RAC-modified asphalt mixtures and optimize the content of bitumen. The model uses the back-propagation of learning under supervision to reduce the mean squared error of the output between predictions and measurements. Marshall testing volumetric and mechanical inputs were used as inputs to one or more hidden layers, which were processed using ReLU or hyperbolic tangent activation functions. Regression was performed using linear output neurons. In TensorFlow/Keras, the MLP was applied with the Adam optimizer (500 epochs, dropout = 0.2) to achieve stabilized and reproducible predictions.
2.5.4. Bagging Ensemble (BAG) Model
Introduced by Breiman [
30], the bagging ensemble (BAG) model eliminates variance in heterogeneous RAC-modified asphalt datasets by boosting samples with the help of bootstrap sampling and the aggregation of numerous base learners. The parallelization of training on resampled data decreases forecasting and overfitting, especially with decision tree models, as described in a previous study [
31] and the
Supplementary Materials. The method enables the more precise optimization of the bitumen content when RAC variability occurs. BAG was also combined with localized granular models (LGM) to identify RAC-specific sub-patterns. The scikit-learn package was used to implement the model on 100 decision trees with a maximum depth of 5–10, and this supported sustainable asphalt mix design.
2.5.5. Boosting Ensemble (BOT) Model
The boosting ensemble (BOT) model, which was first introduced by Freund and Schapire in 1996 as a sequential learning technique, creatively transforms weak learners into robust predictors by repeatedly concentrating on cases that are misclassified or have high residuals. This minimizes the loss function by using gradient descent and adaptive reinforcement. This framework improves predictive associations among features like Marshall stability, flow, and volumetric parameters in recycled asphalt concrete (RAC) mixtures by building a base regression tree with options for optimal subtree pruning and surrogate splits to handle missing data. An aggregated strong model that lowers bias and variance for non-linear bitumen content forecasting was produced by initializing uniform probabilities across training samples (for example, 75% of the 780-sample dataset), creating bootstraps for sequential predictors, and updating weights based on iteration-specific loss calculations.
2.5.6. Hybridization of CARIMA with BAG and a Logistic Probabilistic Model (LGM)
As a flexible framework for modeling binary or bounded outcomes in pavement engineering, the logistic probabilistic model (LGM), which is based on the continuous logistic distribution as a member of the exponential family, converts linear predictor combinations into S-shaped cumulative probabilities between 0 and 1 using the logit link function estimated through maximum likelihood to determine variable significance.
Table 4 below describes a transparent, reproducible hyperparameter-tuning protocol that documents the grid structure, selection rules, and final choices per model layer in the tripartite architecture. Details of the model setup can be found in the studies of Agbor et al. and Nwokolo [
32,
33,
34,
35,
36] and the
Supplementary Materials.
To handle heterogeneous datasets like those for recycled asphalt concrete (RAC) mixtures, the LGM creatively goes beyond traditional logistic regression by incorporating localized generalizations, such as spatially varying coefficients or kernel-based adaptations.
Figure 1 highlights the overall division of labor across stages—the time-series structure in CARIMA, robustness and uncertainty quantification in bagging, and probability calibration in the LGM—making the rationale and benefits of the proposed hybrid architecture obvious at a glance.
Table 5 illustrates the complete derivation of the hybridization of BAG and CARIMA with the logistic probabilistic model (BAG-CARIMA-LGM) to estimate bitumen content under 0%, 30%, and 50% recycled asphalt concrete (RAC) combinations.
2.6. Rationale for the Hybrid BAG-CARIMA-LGM Model’s Selection Among Several Prevalent Machine Learning Models
The BAG-CARIMA-LGM architecture was selected to address three RAC mix design challenges: temporal non-stationarity during curing and aging, heterogeneity across recycling levels, and the need for physically constrained, decision-grade outputs. CARIMA captures sequential drift and provides forecasts and a residual structure; bagging reduces variance using a lightweight ensemble that remains stable on specimen-limited laboratory datasets; and the LGM stage calibrates predictions to admissible binder ranges for practical decision making. The use of the tripartite BAG-CARIMA-LGM (advanced) prediction method is justifiable by the realities of recycled asphalt concrete (RAC) mix design, where parameters are simultaneously time-varying (curing/aging drift), non-homogeneous across recycling levels, and constrained by the admissible binder ranges needed for decision-grade outputs—conditions under which single learners tend to be either unstable, overfit, or physically unconstrained. The method demonstrates its advantages by explicitly allocating the “division of labor” across layers: CARIMA captures sequential drift and residual signatures, bagging suppresses variance/overfitting under specimen-limited laboratory data, and LGM calibration enforces bounded outputs while addressing uncertainty limits, avoiding the computational burden of deep sequence models under the same data budget. These advantages are evidenced in the study’s comparative results: ensemble-only and classical time-series baselines remain around R
2 ≈ 0.72–0.76 for the control mix, whereas the hybrid achieves R
2 > 0.96 with low error metrics and sustained generalization as RAC increases (reported training R
2 = 0.983, testing R
2 ≈ 0.970, and RMSE < 0.40), including robust performance in the demanding 50% RAC case, where SARIMA/CARIMA and MLP/RBF degrade sharply. This layered hybrid logic is consistent with prior reports stating that hybridizing ARIMA-family predictors can improve stability via adaptive error correction [
29] and that bagging-style variance reduction strengthens robustness in heterogeneous engineering datasets [
31].
CARIMA models sequential drift and exposes forecast and residual characteristics, which are bagged with a lightweight bagging ensemble (approximately 100 shallow trees) that is stable on small laboratory data. The LGM stage implements admissible binder ranges through a logistic connection and implements uncertainty limits through optional Beta-likelihood modeling. This is superior to high-capacity deep sequence models of specimen-limited workflows because it is a more efficient design. A comparison with SARIMA, CARIMA, BAG, BOT, MLP, and RBF proves high accuracy in 0, 30, and 50% RAC regimes. It is worth noting that the hybrid has close-to-perfect fits at 0% RAC and large gains at high RAC (R 2 = 0.967 at 50% RAC), with a low RMSE and MaxAPE.
2.7. Analytical Tools and Performance Evaluation
Models were evaluated using the coefficient of determination (R
2), mean absolute percentage error (MAPE), root mean square error (RMSE), and maximum absolute percentage error (MaxAPE), with hyperparameter tuning via grid search, as shown in Equations (2)–(5):
The root mean square error (RMSE) evaluates the absolute predictive accuracy, defined as
which is further normalized to the maximum absolute percentage error (MaxAPE) for cross-dataset comparability:
Additionally, the mean absolute percentage error (MAPE) is crucial for understanding relative deviations:
4. Limitations of the Study
Notwithstanding the study’s innovative contributions, a number of limitations must be noted in order to place the results in context.
Experimental Scope: Most of the performance assessments were carried out in a laboratory setting. While these offer valuable insights, field-scale conditions like fluctuating traffic loads, extreme weather, moisture damage, and long-term degradation mechanisms might not be accurately replicated. This limits the ability to directly translate the findings into extensive applications.
Modeling Scope: While AI-based predictive models greatly improved the mix design and performance forecasting accuracy, their success is heavily reliant on the quality, diversity, and size of the input dataset. The generalizability of the developed models across various regions and traffic scenarios may be limited by the lack of extensive field data.
Practical Scalability: Practical viability is hindered by the explicit constraints of the three-stage pipeline: CARIMA’s temporal layer presumes curing-/aging-ordered sequences; in field deployments, where timestamps and batching metadata are incomplete, the temporal signal export that feeds the ensemble diminishes, elevating uncertainty and necessitating periodic recalibration. End-to-end complexity can propagate specification errors across layers, even if the integration logic is fixed (CARIMA → residual/metafeatures → bagging mean/variance → LGM calibration), which necessitates strict versioning and documented interface checks during transfer to new plants or RAC regimes. Domain shift exposure is partially addressed through “swapped” evaluation across RAC regimes, although broader covariate shifts—materials sources, rejuvenators, ambient climates—require site-specific validation before routine adoption. Evidence of superiority is established on a modest corpus (~780 samples) spanning 0%, 30%, and 50% RAC, which supports small-data applicability yet limits claims beyond the represented mixtures and testing protocols. The reported accuracy at 30% and 50% RAC underscores the method’s robustness under heterogeneity, but it should be stress-tested against additional plants and gradations to ensure portability. Finally, while comparisons included classical and ML baselines (SARIMA, CARIMA, BAG, BOT, MLP, RBF), practical deployment would benefit from periodic head-to-head evaluations against contemporary hybrid/deep architectures as data resources evolve, seeking to guard against performance regression under changing operational regimes.
4.1. Conclusions
By combining recycled materials, cutting-edge modifiers, and AI-based optimization, this study provides a novel framework for improving the sustainability and performance of asphalt mixtures. The results demonstrate that mechanical strength, aging resistance, and environmental performance can be considerably enhanced by nano-reinforcement, bio-additives, and industrial by-product fillers. These developments, when combined with predictive digital tools, speed up mix design, lessen the reliance on resources, and match pavement engineering goals with those of the circular economy. The study provides useful avenues for the development of resilient and carbon-conscious infrastructure by demonstrating that intelligent and eco-efficient pavements can be accomplished without losing structural dependability. The hybrid BAG-CARIMA-LGM framework transcends mere result recitation by implementing an intelligent mix design paradigm that integrates temporally aware decomposition (CARIMA), variance-reduced learning (bagging), and bounded probabilistic calibration (LGM) to produce decision-grade predictions that explicitly facilitate circular, performance-oriented pavement design. Bounded logistic calibration converts ensemble values into permissible binder ranges with clear intervals, rendering the estimates practical for specification development and field risk management. The architecture effectively mitigates error proliferation in high-recycling combinations, which is crucial in maximizing sustainable value, thereby facilitating increased RAC utilization without compromising mechanical dependability and guiding cost-conscious maintenance strategies. The comparative data among virgin, mid-RAC, and high-RAC regimes demonstrates a significant transition in methodologies, from empirical charts to AI-assisted optimization, with the hybrid model surpassing baseline learners in contexts most pertinent to sustainable deployment. Robustness in the face of data scarcity is tackled via “swapped” evaluation across compositional regimes and a modeling framework intentionally designed to be compatible with limited laboratory-to-plant datasets, enhancing the method’s practical applicability while integrating AI methodologies with materials engineering limitations.
The theoretical contribution resides in formalizing a multi-stage, bounded probabilistic learning paradigm for asphalt mix design: CARIMA exposes curing/aging dynamics and exports residuals/metasignals, bagging aggregates over heterogeneity to stabilize estimation, and the LGM calibrates predictions to physically admissible binder ranges with interpretable uncertainty—an integrated construct that advances beyond single-stage regressors and unbounded outputs while institutionalizing swapped cross-regime validation for robustness under composition shifts. Practically, the pipeline demonstrates superior accuracy across RAC levels and translates ensemble dispersion into decision-grade intervals, supporting specification setting, budgeting, and sustainable, high-RAC deployment under modest data regimes and plant-level computing. Methodological transparency and reproducibility are reinforced through explicit metrics (R2, MAPE, RMSE, MaxAPE) and a documented training/selection protocol that can be ported to routine QC/QA contexts.
The tripartite BAG-CARIMA-LGM is optimally suited for specimen-limited laboratory workflows that require the prediction of the ideal bitumen content for recycled asphalt concrete across varying recycling levels (0–50% RAC) in environments like Salalah, where the mixture response may vary with curing/aging and the objective is to achieve a physically plausible, decision-grade binder estimate instead of an unrestricted point forecast. The CARIMA layer is suitable when measurements can be organized into a significant sequence (curing/aging order, batch/condition progression), allowing for the formal management of non-stationarity through KPSS-triggered differencing, AICc/BIC order selection, and Ljung–Box residual diagnostics prior to validating the predictive accuracy. The BAG layer is particularly advantageous when RAC regimes generate significant variation and non-linear feature interactions, since bootstrap aggregation with about 100 shallow trees (depths of 5 to 10) mitigates overfitting and enhances the prediction stability amid diverse laboratory variability. The LGM layer is necessary when outputs must be constrained and interpretable as calibrated probabilities, using maximum likelihood estimation and calibration residual assessments, with the whole tuning methodology fully detailed in
Table 4 for repeatability and transferability. Practical deployment prerequisites necessitate a uniform feature schema based on Marshall stability/flow and volumetric parameters, along with sufficient representation of the target RAC regimes (780 specimens) to ensure that the cross-validated RMSE/MAPE and hold-out validation can be used to determine model acceptance prior to utilizing predictions for binder dosage decisions.
4.2. Suggestions for Future Research
Subsequent studies ought to concentrate on verifying lab findings in actual field settings with varying traffic and climate conditions. To better capture the intricate physicochemical interactions between recycled aggregates, nanomaterials, and bio-based additives, multi-scale experimental and computational methods are advised. Adaptive pavement management systems and improved predictive robustness can be achieved by combining AI models with mechanistic–empirical design frameworks. To create comprehensive sustainability and cost-effectiveness profiles, life cycle assessments and techno-economic evaluations should also be carried out. It is also necessary to expand the swapped evaluation to broader covariate shifts (aggregates, rejuvenators, climates) and consider multi-plant transfer to stress-test portability. Other suggestions for future research are as follows:
Benchmark against contemporary hybrids/deep architectures under identical data budgets to quantify gains beyond classical baselines;
Develop physics-informed or microstructure-aware feature generators that fuse CARIMA outputs with mechanistic film thickness/adhesion descriptors;
Introduce active learning and domain adaptation for data-sparse sites, prioritizing experiments that maximally reduce the predictive uncertainty captured by the LGM;
Conduct longitudinal field validations linking predictive intervals to life cycle performance and costs, embedding QC/QA triggers for model recalibration;
Integrate environmental and cost modules to co-optimize binder content with circularity targets and maintenance budgets under bounded risk criteria.