Next Article in Journal
Residual Mechanical Properties of Printable Concrete Subjected to Elevated Temperatures
Previous Article in Journal
Deformation Prediction of Deep Foundation Pit Support Piles Based on a CNN-LSTM-Transformer Model with Spatiotemporal Feature Fusion
Previous Article in Special Issue
Integrated Risk Priority Assessment of Engineering and Non-Engineering Factors Influencing Saudi Arabian Construction Projects
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Construction Input Price Forecasting for Probabilistic Contingency Estimation in a Road Infrastructure Bridge Case Study

1
Facultad de Ingeniería Civil, Universidad Nacional Federico Villarreal, Lima 15001, Peru
2
Facultad de Ingeniería y Arquitectura, Universidad San Martín de Porres, Lima 15084, Peru
3
Facultad de Ingeniería, Universidad Privada del Norte, Trujillo 13001, Peru
*
Author to whom correspondence should be addressed.
Buildings 2026, 16(11), 2124; https://doi.org/10.3390/buildings16112124
Submission received: 21 April 2026 / Revised: 22 May 2026 / Accepted: 23 May 2026 / Published: 26 May 2026

Abstract

Road infrastructure projects are frequently affected by cost overruns driven by volatility in critical construction inputs and by the uneven association between external market shocks and material price movements. However, existing studies still provide limited evidence on how comparative forecasting, temporal price-signal diagnostics and probabilistic simulation can be integrated into a contingency-oriented decision framework. This study examines how construction input price forecasting and probabilistic simulation can inform contingency estimation in a road infrastructure case study. The empirical application is based on a Peruvian bridge project and combines benchmark-oriented forecasting using Bi-GRU and Random Walk models, descriptive temporal diagnostics based on lead–lag assessment and rolling-correlation analysis, and Monte Carlo simulation. Monthly series for structural steel, construction steel, cement, and diesel were transformed into log-returns and evaluated under a strict chronological design, while oil, the exchange rate, and the consumer price index were incorporated as exogenous variables. The Random Walk model produced lower forecasting errors for most inputs, achieving lower RMSE values in seven of the eight input-period comparisons; Bi-GRU outperformed it only for diesel in the test subset, with a 7.24% lower RMSE. From a project cost-risk perspective, the P95 contingency was estimated at 3.92% under Bi-GRU and 3.96% under Random Walk, indicating a similar upper-percentile contingency envelope under both forecasting specifications. The findings support contingency as a confidence-based budgeting decision rather than a fixed percentage.

1. Introduction

Globally, it has become increasingly common for road infrastructure projects to be affected by cost overruns. In road infrastructure projects, the volatility of critical inputs has become a relevant source of uncertainty for cost estimation, budget updating, and contingency definition. In the specific case of road projects, Herrera et al. [1] identified 38 cost-overrun factors and ranked material price variation as the second most important factor according to their Influence Index, second only to design failures. In addition, in recent years, the construction sector has faced successive escalation pressures associated with global disruptions in markets and supply chains, which has reinforced project exposure to abrupt changes in strategic inputs [2]. In bridges and highways, this effect may be amplified by the size of the investment, by the high share of certain materials in the cost structure, and by the need to commit resources before market uncertainty has fully dissipated. In this context, the problem is not only to anticipate price variations, but also how to translate that uncertainty into more realistic budgeting and contingency decisions.
A related practical limitation is that contingency in construction is still often treated as a fixed reserve or as a reference percentage, even though project exposure is inherently dynamic and depends on the interaction between project-specific characteristics and changing market conditions. Recent studies have emphasized that contingency should be updated and interpreted according to actual cost behavior rather than assumed as a static margin from the outset [3,4]. Likewise, it has been shown that cost performance in road projects varies according to project characteristics, bidding conditions, and contracting processes, which suggests that cost-risk interpretation should be context-sensitive rather than uniform across projects [5,6]. From this perspective, contingency estimation becomes a problem of quantitative risk interpretation under uncertainty, rather than simply a deterministic add-on to a base estimate.
Research shows that useful tools have been implemented to address this challenge; however, these depend on the availability of data associated with the project variables that generate cost overruns. Studies on construction cost indices and material prices have shown that macroeconomic variables and other related market factors can improve the analysis and prediction of construction input behavior compared with purely univariate approaches [7,8,9,10,11]. More recently, both traditional machine learning and deep learning models have been applied to construction cost and cost-index forecasting. Tree-based and ensemble approaches, including Random Forest, XGBoost, LightGBM, and natural-gradient boosting variants, have shown practical value in construction cost prediction, particularly when the available data are limited and interpretability remains relevant for decision-making [12,13]. Deep learning models such as LSTM, GRU, and hybrid recurrent architectures have also been applied to construction costs and highway cost indices under volatile conditions [14,15,16,17]. However, the same literature also shows that greater model complexity does not automatically lead to better predictive performance, especially in persistent series where transparent benchmarks remain highly competitive [18,19]. Therefore, the value of forecasting in this field should not be judged only by whether a complex model outperforms a simple benchmark, but also by whether forecast-derived uncertainty can be translated into useful information to improve project-level decision-making.
At the same time, the quantitative risk analysis literature has increasingly emphasized the importance of moving from deterministic estimates to probabilistic cost interpretations. In infrastructure projects, probabilistic methods such as Monte Carlo simulation have been used to represent cost exposure, contingency ranges, and the effect of correlated risk factors under uncertainty [4,20,21]. Nevertheless, integrated applications remain comparatively limited in studies that connect comparative forecasting, price-signal interpretation, and contingency estimation within the same empirical analytical workflow. This limitation appears especially relevant in infrastructure practice and remains relatively underdocumented in Latin American contexts, where the literature has indeed documented cost and schedule performance problems in road projects [1,6]. In this context, available applications have tended to rely mainly on classical stochastic models, with less explicit incorporation of historical construction material price series as a direct input for probabilistic contingency estimation [22,23].
An additional element that deserves attention is that construction input prices do not evolve in isolation. Fuel prices, inflationary conditions, exchange-rate effects, and pressures in commodity markets may influence construction materials with different intensities and different time lags, implying that cost signals can be associated unevenly across inputs [24,25,26]. For this reason, descriptive temporal diagnostics such as lead–lag analysis and rolling-correlation analysis may add value even when they are not used as formal causal models. Their contribution is interpretive: they help identify whether key inputs co-move with external pressure synchronously or with delays, thereby improving the timing of procurement review, cost monitoring, and contingency reassessment in infrastructure projects.
Therefore, the central gap addressed in this study is not the isolated prediction of construction input prices, but the integration of benchmark-aware forecasting outputs into a probabilistic simulation framework that supports contingency estimation at the project level.
Against this background, the present study is framed as a case-based empirical quantitative cost-risk analysis in road infrastructure. Using the Carrasquillo Bridge in Peru as the case study, the research combines comparative forecasting through Bi-GRU and Random Walk models, descriptive temporal diagnostics, Monte Carlo simulation, and sensitivity analysis to examine how uncertainty in construction input prices can support probabilistic contingency estimation. The study addresses three specific aims: to compare the performance of nonlinear models and benchmarks for critical project inputs; to document the temporal behavior of selected relationships among input prices under external market pressure; and to translate forecast uncertainty into project-level contingency information. In this way, the contribution of the article lies in showing how comparative forecasting can serve as a practical input for quantitative risk analysis and contingency-oriented decision-making in infrastructure projects exposed to construction input volatility.

2. Materials and Methods

2.1. Research Design and Workflow

This study adopted a quantitative, longitudinal, and case-based research design to analyze the dynamic behavior of critical construction input prices and to assess their implications for cost risk in road infrastructure projects. The empirical application was developed using the Carrasquillo Bridge in Peru as the case study. The workflow was used to organize the empirical analysis of construction input price behavior and project-level probabilistic contingency estimation under material price volatility.
The workflow comprised five sequential stages, as shown in Figure 1. First, the case-study project was characterized, and the most cost-relevant construction inputs were identified from the bridge cost structure. Second, monthly time-series data for construction input indices and related macroeconomic drivers were compiled and preprocessed. Third, a comparative forecasting framework was implemented using a Bi-GRU model and a Random Walk benchmark under strict chronological partitioning. Fourth, dynamic dependence diagnostics were conducted through rolling correlations, lead–lag analysis, and a descriptive volatility–correlation assessment in order to identify time-varying co-movement patterns between oil and selected construction inputs. Fifth, the forecasting outputs were integrated into a Monte Carlo simulation scheme to estimate probabilistic cost-risk metrics, contingency levels at different percentiles, and the relative influence of key inputs on overall cost exposure.
This workflow was designed to connect time-series forecasting with project-level cost-risk interpretation in a case-based infrastructure setting. Such integration is consistent with recent construction management research emphasizing that forecasting tools become more useful when they support planning, control, and risk-informed decision-making rather than being assessed only through point-prediction accuracy [20,27,28].

2.2. Case Study, Variables, and Data Preparation

The empirical application focused on the Carrasquillo Bridge project in Peru. The selection of variables was guided by the cost structure of the case study and by the relevance of the inputs to bridge and road infrastructure projects. To clarify the scope of the case study, Table 1 summarizes the main characteristics of the selected bridge cost basket and the role of each input in the probabilistic analysis.
Accordingly, the findings are most transferable to road infrastructure components with steel-intensive material baskets and exposure to input price volatility. They should not be generalized without adjustment to projects in which asphalt, earthworks, pavement layers, or other cost components dominate the risk structure. The analysis considered the most influential construction inputs in the selected cost basket, including structural steel, construction steel, cement, and diesel-related cost signals, together with macroeconomic and external drivers such as oil, exchange rate, and CPI. This selection is consistent with the construction cost forecasting literature, which has shown that multivariate specifications provide a better representation of construction cost dynamics than purely univariate approaches when material prices are affected by broader economic conditions [7,8,10,24,29]. The construction input indices and the exchange rate and CPI series were obtained from [30]. The oil price series was obtained from [31]. The project quantities, unit prices, and material weights used to define the selected bridge cost basket were obtained from the case-study cost estimate. Further details on the input variables and the Monte Carlo cost basket are provided in Appendix B and Appendix C, respectively.
The dataset was organized as a strictly monthly level time series covering the period from January 2013 to December 2025. Because the variables were measured on different scales and the interest of the study was centered on temporal changes rather than absolute levels, the series were transformed into monthly log-returns before model estimation, as shown in (1). This transformation yielded an effective return series from February 2013 onward and supported the construction of 12-month input sequences for one-step-ahead forecasting. Log-return transformation is commonly used in time-series and price-dynamics studies because it improves comparability across variables and facilitates the analysis of relative variations over time [24,32]. Only strictly positive observations were retained in the original level series to preserve the validity of the logarithmic transformation.
r i , t = ln ( X i , t ) ln ( X i , t 1 )
where r i : log-return of input i in period t and X i : index value or price of input i at time t. Therefore, the original level series were not used directly for model training; they were first transformed into monthly log-returns and later reconstructed into index levels for forecasting evaluation and cost simulation.
To preserve the temporal structure of the dataset and prevent information leakage, the forecasting design used a strict chronological split without shuffling. After constructing the 12-month input sequences, the target dates were divided into training (February 2014 to December 2021), validation (January 2022 to December 2023), and test (January 2024 to December 2025) subsets. This partitioning yielded 95 training samples, 24 validation samples, and 24 test samples. In addition, the scaling procedure was fitted only on the training log-return data and then applied unchanged to the validation and test subsets using a MinMaxScaler with range (−1, 1). This design follows best practices in forecasting evaluation, which recommend temporally ordered partitions, explicit benchmark comparison, and leakage-free preprocessing in order to obtain credible out-of-sample results [18,19]. This scaling was used only as a numerical preprocessing step for model estimation on log-returns; forecasting evaluation and cost simulation were conducted after reconstructing the series into index levels.

2.3. Comparative Forecasting Framework

The forecasting component was designed as a benchmark-aware comparative exercise. The main forecasting specification was a bidirectional gated recurrent unit (Bi-GRU) model. GRU architectures are frequently described as computationally lighter alternatives to LSTM because they require fewer parameters while retaining the ability to capture nonlinear temporal dependencies in sequential data [15,33,34]. This characteristic is particularly relevant for monthly construction-related time series of limited length, where parsimony and training stability are important considerations. Recent studies in construction and infrastructure forecasting have shown that GRU-based and hybrid recurrent architectures can achieve competitive performance in construction cost and highway cost-index prediction tasks [35].
The bidirectional structure was used to improve feature extraction within each historical input window. Importantly, the model did not use information beyond the forecast origin; rather, the sequence was processed in both directions only within the observed input window. The forecasting task was formulated as a one-step-ahead monthly prediction problem. The implementation adopted sliding windows, early stopping, and learning-rate reduction to limit overfitting and improve model generalization under a relatively short sample.
Therefore, the Bi-GRU model was not used under the assumption that deep learning would necessarily outperform simpler alternatives. Rather, it was included to test whether a recurrent multivariate architecture could provide practical gains over a transparent benchmark under a limited monthly dataset. The selected Bi-GRU configuration should be interpreted as a parsimonious and regularized recurrent specification for a short monthly dataset, rather than as the outcome of an exhaustive neural architecture search. The complete Bi-GRU model specification is provided in Appendix A.
The rationale for selecting Bi-GRU was that the forecasting task was formulated as a multivariate sequential problem, in which 12-month input windows combined four target input series with three exogenous variables. GRU-based architectures provide a recurrent structure for learning temporal dependencies while using fewer parameters than LSTM-based alternatives, which is relevant when the available monthly sample is limited [15,33,34]. However, given the small effective training size, the Bi-GRU specification should be interpreted as an exploratory recurrent benchmark rather than as a fully optimized deep learning architecture.
To place the performance of the Bi-GRU model in context, a Random Walk benchmark was also implemented. Benchmark comparison was considered essential because in persistent economic and price-index series, simple no-change forecasts often provide strong short-horizon performance and may be difficult to outperform consistently [18,19]. Accordingly, the objective of the comparison was not only to identify the model with the lowest forecasting error, but also to determine whether the additional complexity of a deep learning model yielded a meaningful practical gain over a transparent baseline.
More conventional machine learning methods, particularly tree-based models such as Random Forest, XGBoost, and related boosting approaches, may be more stable and competitive in small-sample settings when the problem is reformulated with lagged tabular features [12,13]. They were not included as additional benchmarks in this version because the objective was to evaluate the decision value of a recurrent multivariate specification against a transparent no-change benchmark, rather than to conduct an extensive model competition. Nevertheless, their inclusion is an important direction for future research.
The Random Walk benchmark was implemented as a univariate no-change forecast for each target input series. Therefore, it did not incorporate the exogenous variables used in the Bi-GRU specification. This choice is consistent with the role of Random Walk as a transparent benchmark. The benchmark was not intended to reproduce the same information set as the neural network, but to provide a transparent operational reference against which the practical value of the richer multivariate deep learning specification could be assessed. Consequently, the comparison should be interpreted as a benchmark-aware decision test rather than as a same-feature model competition.
Forecast accuracy was evaluated on reconstructed level series, since decision-making in construction budgeting and procurement is ultimately performed in the original cost-index space rather than in transformed returns. The evaluation relied on standard error metrics, including mean absolute error (MAE), root mean square error (RMSE), and mean absolute percentage error (MAPE), which are widely used in forecasting studies to compare predictive performance from complementary perspectives [18,19].

2.4. Dynamic Dependence Diagnostics

A lead–lag analysis was conducted by shifting the oil return series from 0 to 6 months in order to identify the lag associated with the maximum absolute correlation with the Cement_Index, Construction-Steel_Index, and Diesel_Index. This procedure was used to approximate the temporal association structure between oil and the selected construction-related price indices. The optimal lag for each target variable was then incorporated into a second rolling-correlation analysis to examine the time-varying behavior of the lag-adjusted relationships.
Additionally, rolling oil-price volatility was computed as the 12-month standard deviation of oil log-returns and descriptively contrasted with the lag-adjusted rolling correlations between oil and the selected construction input indices. Specifically, the analysis considered Structural-Steel_Index, Diesel_Index, Cement_Index, and Construction-Steel_Index, using in each case the optimal lag identified in the lead–lag assessment. This comparison was intended to examine whether periods of higher oil volatility tended to coincide with stronger or weaker time-varying co-movements across inputs. The volatility–correlation analysis was interpreted as a descriptive diagnostic rather than as a formal causal test.

2.5. Probabilistic Cost-Risk Simulation and Sensitivity Analysis

To translate the forecasting results and descriptive dependence diagnostics into project-level cost-risk evidence, a Monte Carlo simulation procedure was implemented. Recent construction research has shown that probabilistic modeling is particularly useful when material price volatility must be propagated into project-level cost risk, since point forecasts alone do not provide information on uncertainty ranges, percentile exposure, or contingency requirements [20,27,28]. Following this logic, the present study used simulated scenarios of critical input behavior to estimate the distribution of cost outcomes associated with the selected bridge cost basket.
For each simulation run, the project-level material cost was estimated by aggregating the simulated contribution of the selected critical inputs according to their participation in the case-study cost structure. In conceptual terms, the simulated cost outcome for scenario s can be expressed as shown in (2):
C ( s ) = i = 1 n w i   P i ( s )
where wi represents the relative participation of input i in the selected project cost basket, and Pi(s) denotes the simulated price level or cost factor of that input under scenario s. Repeated sampling produced an empirical distribution of possible cost outcomes for the selected bridge materials.
In the Monte Carlo stage, uncertainty was propagated from the forecasting layer to project-level cost outcomes by centering each simulation on the deterministic one-step-ahead return forecast for each material and period, and then adding stochastic shocks drawn from a model-specific walk-forward residual pool. For each forecasting model (Bi-GRU and Random Walk), 10,000 joint shock vectors were generated using a smoothed bootstrap procedure in which residual rows were resampled jointly across materials to preserve their empirical contemporaneous cross-dependence, and a small multivariate normal jitter based on the residual covariance matrix, estimated from the same model-specific walk-forward residual pool, was added to avoid repeated draws and smooth the simulated distribution. This procedure was intended to preserve empirical contemporaneous cross-dependence among residuals, but it was not designed as a formal nonlinear dependence or tail-dependence model. Also, in similar studies conducted in comparable contexts, a similar number of 10,000 iterations has also been used [22,23]. The simulated returns were then transformed into index levels using the previous observed month as the reconstruction base, after which index levels were converted into unit prices through proportional scaling relative to the selected cost base date and finally multiplied by material quantities to obtain simulated material costs and total project material cost. This design allowed forecast uncertainty to be translated explicitly into a probabilistic cost distribution while maintaining consistency with the observed cost structure of the case-study project.
The simulation results were summarized using central tendency and risk percentiles, particularly P50, P80, P90, and P95, together with maximum simulated exposure. These percentiles were used to quantify contingency requirements under different confidence levels, which is consistent with probabilistic cost-risk assessment practices in construction under material price uncertainty [20]. In this study, contingency was defined consistently as the difference between a target percentile of the simulated cost distribution and P50. Thus, the P90 and P95 contingencies correspond to P90 − P50 and P95 − P50, respectively.
In addition, a sensitivity analysis was performed to identify the relative importance of each input in the simulated cost distribution. The results were summarized through a tornado diagram, which provides a visual ranking of the dominant drivers of cost variability. This step was included because, from a project-management perspective, understanding which inputs contribute most strongly to overall uncertainty is essential for procurement review, monitoring priorities, and risk response planning [20,27].

3. Results

3.1. Comparative Forecasting Performance of Critical Construction Inputs

The forecasting stage was first assessed at the level of the individual input indices used in the project-level cost-risk framework. Table 2 and Table 3 summarize the one-step-ahead forecasting performance of the Random Walk and Bi-GRU models, respectively, on reconstructed level series for the validation and test subsets. This comparison is relevant because the subsequent probabilistic cost analysis depends on the quality of the underlying price projections. Table 4 complements this comparison by reporting the relative RMSE variation in Bi-GRU with respect to the Random Walk benchmark, thereby showing where model complexity translated into practical gains and where it did not.
The results indicate that the forecasting behavior was heterogeneous across materials. The Random Walk benchmark provided lower errors for most inputs and periods, confirming the persistence of several material price indices. However, the Bi-GRU model showed a relative advantage for diesel in the test period, suggesting that its nonlinear structure captured part of the short-term dynamics of this input more effectively than the benchmark. Cement exhibited the lowest forecasting errors under both models, whereas diesel and the steel-related indices showed larger deviations, particularly in the validation subset. These results indicate that forecasting gains were material-specific rather than uniform across the selected input basket.
Taken together, Table 2, Table 3 and Table 4 show that the benchmark remained competitive in most cases, while the Bi-GRU added value only selectively. This finding is important for the rest of the analysis because it suggests that model complexity should be judged not only by average forecasting performance, but also by whether it materially changes the project-level uncertainty envelope.

3.2. Lead–Lag Structure and Dynamic Dependence of Selected Inputs

The second stage of the results examined the temporal interaction between oil and the main construction inputs included in the selected bridge cost basket. Table 5 reports the optimal lead–lag structure obtained from the correlation analysis.
Because the estimated correlations are low to moderate, these results should be interpreted as descriptive association patterns rather than as evidence of causal price transmission. Table 6 complements this analysis by presenting the descriptive association between rolling oil-price volatility and the lag-adjusted rolling correlation for each input. These results are relevant because they show whether the selected materials show associations with oil-related signals with similar timing or whether they exhibit differentiated co-movement patterns that may affect cost planning.
The lead–lag results indicate that diesel showed the strongest and most immediate association, while cement showed a delayed association and the steel-related indices reached their highest correlations at longer lags but with lower magnitude. This pattern suggests that the selected inputs did not exhibit synchronous associations with oil-related movements. The evidence from the lag-adjusted Oil–Cement series also shows that the relationship was regime-dependent rather than stable over time, with both sustained negative phases and one clearly sustained positive phase. Finally, the volatility–correlation assessment indicates that stronger oil volatility tended to coincide with stronger lag-adjusted co-movement, particularly for cement, although the strength of this association remained moderate.
These results indicate that the descriptive temporal relationships were both input-specific and time-varying. From a project-management perspective, this means that oil-related market signals should not be interpreted as affecting all inputs at the same horizon. Instead, procurement timing and contingency review may benefit from a differentiated view of how cement, diesel, and steel-related costs co-move with external market pressure over different horizons.

3.3. Probabilistic Cost-Risk Results from Monte Carlo Simulation

After forecasting the input indices and analyzing their dynamic interactions, the next stage translated the projected input behavior into project-level cost-risk metrics. Table 7 summarizes the Monte Carlo results for the validation and test subsets under both forecasting models. This table is relevant because it moves the analysis from individual input forecasts toward the distribution of total material cost for the selected bridge cost basket. In addition to P50, P90, P95, and maximum simulated values, the table reports percentile-based contingencies measured as the difference between the upper percentile and P50.
The simulation results show that the uncertainty band was wider in validation than in test under both forecasting models, which is consistent with the larger forecast errors observed in the validation subset. At the aggregate cost level, both models produced very similar central and upper-tail values, suggesting that the project-level probabilistic envelope was relatively robust to the forecasting specification. In the test subset, the Random Walk model generated slightly higher central and upper percentile estimates, while the Bi-GRU model produced a slightly higher maximum simulated exposure. Thus, the main difference between models was not the central range of the distribution, but the behavior of the extreme right tail.
To provide a more detailed view of the latest test period, Table 8 reports the deterministic total cost, and simulated percentile outcomes for December 2025. This table is relevant because it translates the model outputs into a single decision point that can be interpreted directly in budgeting and contingency terms. In both models, the deterministic estimate remained very close to the observed total cost, while the probabilistic percentiles shifted the expected exposure upward. The two models again yielded highly similar central values, although the Random Walk specification produced a slightly wider extreme tail in the latest period.
The maximum simulated value should be interpreted only as an exploratory extreme-tail indicator, since it is sensitive to the number of simulations and distributional assumptions. For contingency estimation, the main decision indicators in this study are P90 and P95. Figure 2 shows that the Bi-GRU-based distribution is centered slightly above the observed cost and exhibits a moderate right tail. The mean and P50 are close to each other, indicating that the simulated distribution remains reasonably concentrated in its central range. The separation between P50, P90, and P95 is visible but controlled, which suggests a moderate upward risk band for the latest test period. The maximum simulated value lies notably farther to the right, indicating that extreme upside cost outcomes remain possible even though they are located in the low-probability tail.
Figure 3 presents a very similar central pattern for the Random Walk-based simulation, confirming that both forecasting models generate comparable median and upper-percentile cost levels. However, the right tail extends further than in the Bi-GRU case, and the maximum simulated value is higher. This suggests that the Random Walk specification preserves a slightly heavier extreme-cost tail in the latest period, even though its central percentile estimates remain close to those obtained under Bi-GRU.
Taken together, Table 7 and Figure 2 and Figure 3 show that the probabilistic interpretation is more informative than the deterministic estimate alone. Although the deterministic totals nearly match the observed cost, the simulation reveals a non-negligible contingency requirement if the decision-maker wishes to operate at P90 or P95 confidence levels. In practical terms, the latest-period results indicate that the additional cost allowance required to move from the central estimate to a higher-confidence estimate is material and should not be ignored in cost planning.

3.4. Sensitivity of Total Cost to Critical Inputs

The final stage of the results assessed which inputs most strongly influenced the simulated total material cost. Table 9 demonstrates the Spearman correlations between each input and total cost for the latest test-period simulation under both forecasting models. This table is relevant because it ranks the dominant drivers of uncertainty within the bridge material basket and therefore supports prioritization in procurement monitoring and contingency management. Figure 4 complements this evidence by presenting the tornado sensitivity diagrams for both models in a single two-panel figure.
The sensitivity results are highly consistent across the two forecasting specifications. In both cases, Construction Steel was the most influential driver of total cost variability, followed by Structural Steel. Diesel showed a moderate contribution, whereas Cement exhibited the weakest sensitivity. This ranking indicates that steel-related items dominated the propagation of cost uncertainty in the selected bridge basket. However, this does not imply that the remaining inputs should be ignored. Rather, it indicates that marginal forecast errors or price shocks in steel-related inputs are more likely to shift the contingency envelope than comparable disturbances in less influential materials such as cement. The stability of this ordering across both models strengthens the managerial relevance of the result, since the prioritization of key inputs does not depend materially on the selected forecasting approach.
Figure 4 confirms the dominance of steel-related inputs in the simulated cost-risk structure. In panel (a), corresponding to the Bi-GRU simulation, Construction Steel exhibits the highest association with total cost, followed by Structural Steel, while Diesel and Cement have clearly smaller effects. Panel (b), corresponding to the Random Walk simulation, shows the same ranking with only marginal differences in magnitude. The visual agreement between both panels indicates that the principal cost-risk drivers are structurally stable and not an artifact of a specific forecasting model. From an engineering standpoint, this means that risk monitoring and contingency review should focus first on steel-related inputs, with diesel as a secondary driver and cement as a comparatively less influential contributor to the variability of the selected material basket.
Results show that the Random Walk benchmark remained competitive at the forecasting stage, that the dynamic dependence structure among oil and the selected inputs was heterogeneous and lag-sensitive, and that the project-level probabilistic cost envelope was driven primarily by steel-related inputs. The Monte Carlo framework further showed that, even when deterministic totals remain close to observed costs, the upper-tail exposure and contingency requirements can still be material for engineering decision-making.

4. Discussion

Rather than proposing a new dependence model, this study provides case-based empirical evidence on how comparative forecasting and probabilistic simulation can support contingency-oriented cost interpretation in bridge infrastructure under input price volatility.

4.1. Analysis of Forecasting Performance and Cost-Risk Implications

Within the two-model comparison conducted in this case study, the Random Walk benchmark provided the stronger forecasting reference for most inputs. The model achieved lower forecasting errors for most inputs and, in the latest-period simulation, produced a slightly wider extreme right tail than Bi-GRU. This combination is relevant for engineering practice because the decision depends not only on average fit, but also on the model’s ability to represent low-frequency, high-impact cost overrun scenarios.
The superior performance of the Random Walk benchmark in most comparisons may be explained by the persistence and noise observed in monthly construction input price indices. At short horizons, the latest observed value may provide a stronger signal than complex nonlinear patterns, especially under abrupt and non-stationary market shocks. Although Bi-GRU incorporated a richer 12-month multivariate structure, the limited training sample likely reduced its capacity to learn stable relationships between exogenous variables and material prices. Thus, the finding does not weaken the forecasting–simulation framework; rather, it shows that model complexity does not necessarily improve cost-risk information under persistent, noisy, and short-horizon conditions. This behavior is consistent with near-random-walk dynamics at the monthly forecasting horizon, where persistence and short-term noise may dominate over learnable nonlinear structure.
The simulation confirms that both models produced similar central and upper-percentile values. Expressed relative to the selected central reference, the Bi-GRU model produced increases of 2.90% at P90 and 3.92% at P95, while the Random Walk model produced 2.79% and 3.96%, respectively. Therefore, the main decision indicators remain P90 and P95. Maximum simulated values were retained only as exploratory extreme-tail references and were not used to define contingency levels.
The sensitivity diagrams reinforce this interpretation. In Bi-GRU, construction steel reached a 90.1% correlation with total cost, followed by structural steel with 79.1%, diesel with 54.4%, and cement with 29.9%. In Random Walk, the values were 89.4%, 81.4%, 54.9%, and 22.7%, respectively. Both models preserve the same ranking of importance. Therefore, the study clearly identifies that total cost variability is concentrated in steel-related inputs, while cement contributes a significantly smaller effect.
The results also suggest that the additional exogenous information used by the Bi-GRU did not consistently improve out-of-sample accuracy. This may indicate that the oil, exchange-rate, and CPI signals were either weakly associated, temporally unstable, or difficult to learn reliably from the available sample size.
This evidence has direct implications for practice. The project should not distribute contingency uniformly across materials. Management should focus market monitoring, price review, procurement strategy, and contractual discussion on construction steel and structural steel.

4.2. Comparison with the Relevant Literature

The results are consistent with the forecasting literature, which cautions against assuming the superiority of complex models over benchmark statistical approaches a priori. Makridakis et al. argue that machine learning methods do not necessarily outperform traditional approaches, particularly in highly persistent time series [18]. Similarly, Hewamalage et al. emphasize the importance of benchmarking against reference models and applying rigorous temporal validation schemes in forecasting studies [19]. The present study follows this approach by explicitly comparing Bi-GRU and Random Walk under a chronological evaluation framework, allowing the results to be interpreted within a robust methodological context. In this setting, the findings confirm that greater model complexity does not automatically translate into better performance for decision-making.
However, artificial intelligence and deep learning models remain valuable tools for cost engineering. Previous research has shown that models such as GRU, LSTM, and their hybrid variants can achieve strong performance in forecasting construction costs and indices [35,36,37]. Nevertheless, these studies also highlight that performance depends on data structure, forecasting horizon, and evaluation design. In the present case study, the deep learning model demonstrates competitive performance but does not surpass the benchmark statistical model as the preferred alternative for decision-making in a contingency analysis context.
The behavior of construction inputs is also aligned with the literature on dynamic relationships between energy and material prices. Several studies have documented time-varying dependencies and lag structures between oil and metal markets, as well as associations that vary across economic regimes [38]. The results of this study follow this pattern, showing that cost variability is not evenly distributed across inputs and that the association with external shocks differs according to the specific material considered.
The sensitivity analysis is consistent with empirical evidence in construction economics. Previous studies have shown that material inflation and macroeconomic variables affect cost components differently, with inputs characterized by higher energy intensity and logistical exposure concentrating a larger share of risk [25,39]. In this study, the predominance of steel-related inputs as the main drivers of cost variability is consistent with this literature and with the technical nature of infrastructure projects such as bridges.

4.3. Implications for Engineering Decision-Making in Construction Projects

From a cost management perspective, contingency should be defined as a budgetary decision based on confidence levels rather than as a fixed add-on. The results show that the contingency at the P95 level can reach 3.96%, while the maximum simulated exposure is reported only as an exploratory extreme-tail reference. Despite this, each project must estimate its own contingency, as there is no fixed range for its calculation.
The P95 contingency of approximately 4% should not be interpreted as evidence that conventional 5–10% buffers are necessarily excessive. The result is specific to the selected bridge material basket, the modeled inputs, and the observed price conditions of the case study. Therefore, it should be understood as a project-specific probabilistic estimate rather than as a general replacement for standard contingency ranges. It is important to clarify that the contingency percentages reported in this study are calculated with respect to the simulated cost distribution of the selected bridge material basket. They do not represent a contingency applied to the entire contract budget or a general emergency reserve for the whole project. Therefore, comparisons with contingency ranges reported in other infrastructure studies should be interpreted as contextual benchmarks rather than as strictly equivalent calculation bases.
Other studies, such as one in water infrastructure, reported that a contingency of 8.46% was added but 13.58% was required, while another study in hydropower projects found contingency ranges between 8.28% and 17.97% [4]. The values obtained in this study are generally consistent with those reported in the literature; however, it is important to emphasize that each project should independently determine its contingency. Previous research in road infrastructure has reported contingency ranges between 1.34% and 11% [22]. Moreover, the results of this study complement previous stochastic applications by linking comparative forecasting outputs and residual-based simulation to probabilistic contingency estimation under input price volatility [23]. Consequently, infrastructure budgeting should link contingency reserves to explicit confidence targets and actual risk conditions, rather than to historically comfortable percentages. In this case study, the longer empirical association horizon observed for steel-related inputs suggests that procurement reviews and contingency updates should not be limited to immediate price movements but should also consider delayed co-movement windows when evaluating steel-intensive packages.

4.4. Limitations and Future Research

This study has two main limitations. First, the analysis relies on a relatively short monthly sample and on aggregated indices, which restricts the granularity of project-level decision rules. After constructing 12-month input sequences, the Bi-GRU model was trained with only 95 samples. This small effective training size increases the risk of overfitting and limits the reliability, stability, and generalizability of the neural-network results. Although early stopping, dropout, learning-rate reduction, and chronological validation were used to reduce this risk, no repeated-run analysis, random-seed sensitivity test, or exhaustive hyperparameter search was conducted. Therefore, the Bi-GRU results should be interpreted as indicative evidence for this case-study dataset rather than as a definitive validation of the architecture or as evidence of a stable architecture-level advantage.
Although this level of aggregation is useful for identifying broad cost-risk patterns, it does not fully capture the heterogeneity that may exist across specific materials, suppliers, procurement packages, or contract items. As a result, the practical translation of the results into detailed operational rules for purchasing, escalation clauses, or contingency release remains limited.
Second, the dynamic analysis adopted in this study is descriptive and does not establish formal causal transmission mechanisms among oil, macroeconomic drivers, and construction materials. The rolling correlations, lag diagnostics, and volatility-based comparisons are useful for identifying temporal patterns and possible association structures, but they do not provide a structural explanation of how shocks are associated across markets. Therefore, the results should be interpreted as decision-support evidence rather than as causal proof of market behavior. A related limitation is that the residual-based simulation does not explicitly test nonlinear dependence or alternative tail-dependence structures.
These limitations suggest several methodological recommendations for future research. A first recommendation is to test more formal time-varying dependence models capable of representing changing interaction structures across market regimes. In that regard, copula-based approaches [40,41] offer a promising extension because they allow the analyst to move beyond linear correlation and to represent asymmetric and upper-tail dependence more explicitly. Future studies could therefore compare Gaussian, Student-t, and Archimedean copulas, as well as dynamic copula specifications, to examine whether tail behavior changes over time and whether contingency estimates remain stable under alternative dependence assumptions.
A second recommendation concerns the forecasting component. Although the present study used Bi-GRU as the deep learning specification, the literature has shown that LSTM-based models remain one of the most robust alternatives for modeling nonlinear temporal dynamics in construction cost indices and related infrastructure signals. For that reason, future research should evaluate whether LSTM or hybrid LSTM-based architectures provide a better balance between predictive accuracy, temporal stability, and downstream risk simulation performance [35,36,42]. This recommendation does not imply that the current modeling choice was inappropriate, but rather that LSTM deserves explicit examination as a complementary benchmark within the same forecasting-plus-simulation framework.
A third recommendation is to strengthen the treatment of exogenous variables. Construction material prices are not driven only by their own past behavior, but also by fuel prices, exchange rates, inflation, and other macroeconomic signals [26]. In addition, the exogenous variable set was limited to oil, exchange rate, and CPI. Other macroeconomic, political, and logistics-related factors, such as local political stability, international freight conditions, or global logistics indices, may also affect construction material prices in emerging markets and should be considered in future extensions.
Future studies should compare recurrent neural architectures with conventional machine learning models such as Random Forest, XGBoost, and LightGBM using lagged-feature formulations, repeated temporal validation, and random-seed sensitivity analysis, particularly in small or moderate construction cost datasets. They should also extend the benchmark set to include classical time-series models such as ARIMA/ARIMAX, VAR/VECM, and exponential smoothing, particularly when the objective is to compare forecasting suitability rather than to test the decision value of a transparent baseline.

5. Conclusions

This study examined how construction input price forecasting, descriptive temporal diagnostics, and Monte Carlo simulation can inform probabilistic contingency estimation in a road infrastructure case study, using the Carrasquillo Bridge in Peru as the empirical application. The framework combined benchmark-aware forecasting, dynamic dependence diagnostics, Monte Carlo simulation, and sensitivity analysis to evaluate how critical construction input prices affect project-level contingency requirements.
Among the two tested models, the Random Walk benchmark provided lower forecasting errors for most inputs in the case-study dataset. Although the Bi-GRU model achieved competitive performance, it did not consistently outperform the statistical benchmark. Random Walk provided lower forecasting errors for most inputs and preserved a more severe upper tail in the simulated cost distribution, which is especially relevant for contingency-oriented decision-making.
The dynamic analysis also showed that the relationship between oil and the selected inputs was heterogeneous and time-varying. Diesel exhibited the strongest short-term association, cement showed a delayed association, and steel-related inputs presented differentiated lag structures. These findings indicate that cost-related associations do not occur uniformly across materials and that contingency review should consider material-specific timing.
At the project level, the Monte Carlo results confirmed that both models produced similar central estimates but differed more clearly in the upper tail. The P95 contingency reached 3.92% under Bi-GRU and 3.96% under Random Walk, confirming that both forecasting specifications produced similar upper-percentile contingency estimates. These results support the interpretation of contingency as a confidence-based budgeting decision rather than as a fixed percentage in projects exposed to construction input price volatility. The sensitivity analysis further identified construction steel and structural steel as the main drivers of total material-cost variability.
Beyond these empirical results, the study contributes to construction cost-risk management in three ways. First, it provides an integrated forecasting–simulation workflow that converts construction input price uncertainty into percentile-based contingency information for project-level decision-making. Second, it shows that benchmark models should not be treated merely as secondary comparators, since a transparent Random Walk specification may provide more robust decision support than a more complex neural architecture when price indices are persistent and the available sample is limited. Third, it translates probabilistic outputs into managerial priorities by identifying which inputs dominate the simulated cost-risk envelope. These contributions reinforce the need to define contingency as a confidence-based, project-specific decision rather than as a fixed percentage applied uniformly across infrastructure projects.

Author Contributions

Conceptualization, V.A.A.F.; methodology, V.A.A.F.; software, V.A.A.F. and D.P.; validation, V.A.A.F., D.P. and A.O.; formal analysis, V.A.A.F.; investigation, V.A.A.F. and A.P.; resources, D.P., A.O., and A.P.; data curation, V.A.A.F. and A.O.; writing—original draft preparation, V.A.A.F.; writing—review and editing, V.A.A.F., D.P., A.O. and A.P.; visualization, V.A.A.F.; supervision, A.P.; project administration, V.A.A.F.; funding acquisition, A.O. and A.P. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The dataset and source code supporting the findings of this study are openly available in the following GitHub repository: https://github.com/varizaciv-ai/BiGRU-and-RandomWalk-Forecasting.git, accessed on 20 April 2026. The computational analysis was performed using Python 3.9.3 in a Jupyter Notebook environment, with the main packages including pandas, NumPy, scikit-learn, Matplotlib.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A

Appendix A presents the complete Bi-GRU model specification, including the network architecture, training hyperparameters, callback configuration, temporal data partitioning scheme, scaling procedure, and index-level reconstruction settings used in the forecasting framework. These details are included to enhance reproducibility and to facilitate a transparent interpretation of the comparative forecasting results.
Table A1. Complete Bi-GRU model specification used in the comparative forecasting framework.
Table A1. Complete Bi-GRU model specification used in the comparative forecasting framework.
SectionItemValue
GeneralModelBi-GRU multi-output
GeneralTarget variablesStructural-Steel_Index, Diesel_Index, Cement_Index, Construction-Steel_Index
GeneralNumber of target variables4
GeneralInput sequence length (months)12
GeneralInput features7
GeneralInput feature namesStructural-Steel_Index, Diesel_Index, Cement_Index, Construction-Steel_Index, Oil_Index, FX_Rate, CPI
GeneralOutput1-step-ahead multi-output forecast of target log-returns
ArchitectureLayer 1Bidirectional GRU, 32 units, return_sequences = True
ArchitectureLayer 2Dropout, rate = 0.20
ArchitectureLayer 3GRU, 16 units, return_sequences = False
ArchitectureLayer 4Dropout, rate = 0.10
ArchitectureLayer 5Dense, 4 units, linear activation
TrainingOptimizerAdam
TrainingLearning rate0.00025
TrainingLoss functionMean squared error
TrainingBatch size8
TrainingMaximum epochs500
TrainingValidation checkpoint selectionEarly stopping with best validation checkpoint restored
TrainingBest validation loss0.185476
TrainingRandom seed42
CallbacksEarlyStopping patience8
CallbacksEarlyStopping restore_best_weightsTrue
CallbacksReduceLROnPlateau factor0.5
CallbacksReduceLROnPlateau patience4
CallbacksReduceLROnPlateau min_lr0.00001
DataTemporal split—train endDecember 2021
DataTemporal split—validation endDecember 2023
DataTemporal split—test endDecember 2025
DataTrain target datesFebruary 2014 to December 2021
DataValidation target datesJanuary 2022 to December 2023
DataTest target datesJanuary 2024 to December 2025
DataTrain samples95
DataValidation samples24
DataTest samples24
DataScalingMinMaxScaler fitted on training log-returns only, range = (−1, 1)
DataInput transformationMonthly log-returns
DataForecast reconstructionPredicted returns compounded to reconstructed index levels

Appendix B

Appendix B summarizes the variables used in the forecasting framework, including their unit or form, data source, transformation, and role in the model. All construction input indices and macroeconomic variables were transformed into monthly log-returns before model estimation, while the original levels were later reconstructed for forecasting evaluation and cost simulation.
Table A2. Variables used in the forecasting framework and their modeling roles.
Table A2. Variables used in the forecasting framework and their modeling roles.
VariableUnit/FormSourceTransformationRole in the Model
Structural-Steel_IndexConstruction input indexINEIMonthly log-returnTarget variable
Construction-Steel_IndexConstruction input indexINEIMonthly log-returnTarget variable
Cement_IndexConstruction input indexINEIMonthly log-returnTarget variable
Diesel_IndexConstruction input indexINEIMonthly log-returnTarget variable
Oil_IndexBrent crude oil price indexFREDMonthly log-returnExogenous variable/diagnostic variable
FX_RateExchange rateINEI/official sourceMonthly log-returnExogenous variable
CPIConsumer price indexINEIMonthly log-returnExogenous variable

Appendix C

Appendix C reports the cost basket used in the Monte Carlo simulation, including quantities, base unit prices, cost base date, and relative cost shares. The cost shares wi were calculated from the deterministic base cost of each input and were kept fixed during the simulation, so that uncertainty was propagated through simulated price movements rather than through changes in quantities. The weights define the exposure structure of the selected case-study basket and are not intended to represent a universal cost structure for all bridge or road infrastructure projects.
Table A3. Cost basket and fixed input shares used in the Monte Carlo simulation.
Table A3. Cost basket and fixed input shares used in the Monte Carlo simulation.
InputQuantityUnitBase Unit Price (PEN)Base DateCost Share wi
Structural steel3445.92ton5522.40December 20210.2393
Construction steel6,642,198.16kg4.14December 20210.3459
Cement23,897,597.51kg0.63December 20210.1894
Diesel-related input910,798.28gal19.68December 20210.2254

References

  1. Herrera, R.F.; Sánchez, O.; Castañeda, K.; Porras, H. Cost Overrun Causative Factors in Road Infrastructure Projects: A Frequency and Importance Analysis. Appl. Sci. 2020, 10, 5506. [Google Scholar] [CrossRef]
  2. Chammout, B.; El-adaway, I.H.; Nabi, M.A.; Assaad, R.H. Price Escalation in Construction Projects: Examining National and International Contracts. J. Constr. Eng. Manag. 2024, 150, 04024109. [Google Scholar] [CrossRef]
  3. Stachoń, T.; Szóstak, M.; Konior, J. Updating Financial Contingency in Execution of Typologically Diverse Construction Projects. Appl. Sci. 2025, 15, 4445. [Google Scholar] [CrossRef]
  4. Baccarini, D.; Love, P.E.D. Statistical Characteristics of Cost Contingency in Water Infrastructure Projects. J. Constr. Eng. Manag. 2014, 140, 04013063. [Google Scholar] [CrossRef]
  5. Lee, K.-W. Cost Performance Comparison of Road Construction Projects Considering Bidding Condition and Project Characteristics. Sustainability 2024, 16, 10083. [Google Scholar] [CrossRef]
  6. Gómez-Cabrera, A.; Cortés, S.; Rojas, J.; Sánchez, O.; Torres, A. Data-Driven Analysis of Contracting Process Impact on Schedule and Cost Performance in Road Infrastructure Projects in Colombia. Buildings 2025, 15, 3739. [Google Scholar] [CrossRef]
  7. Ashuri, B.; Lu, J. Time Series Analysis of ENR Construction Cost Index. J. Constr. Eng. Manag. 2010, 136, 1227–1237. [Google Scholar] [CrossRef]
  8. Ashuri, B.; Shahandashti, S.M.; Lu, J. Empirical tests for identifying leading indicators of ENR Construction Cost Index. Constr. Manag. Econ. 2012, 30, 917–927. [Google Scholar] [CrossRef]
  9. Shahandashti, S.M.; Ashuri, B. Forecasting Engineering News-Record Construction Cost Index Using Multivariate Time Series Models. J. Constr. Eng. Manag. 2013, 139, 1237–1243. [Google Scholar] [CrossRef]
  10. Shiha, A.; El-adaway, I.H. Forecasting Construction Material Prices Using Macroeconomic Indicators of Trading Partners. J. Manag. Eng. 2024, 40, 04024036. [Google Scholar] [CrossRef]
  11. Passek, M.; Nübel, K. Forecasting Office Construction Price Indices for Cost Planning in Germany Using Regularized VARX Models. Buildings 2025, 16, 103. [Google Scholar] [CrossRef]
  12. Chakraborty, D.; Elhegazy, H.; Elzarka, H.; Gutierrez, L. A novel construction cost prediction model using hybrid natural and light gradient boosting. Adv. Eng. Inform. 2020, 46, 101201. [Google Scholar] [CrossRef]
  13. Zhang, J.; Yuan, J.; Mahmoudi, A.; Ji, W.; Fang, Q. A data-driven framework for conceptual cost estimation of infrastructure projects using XGBoost and Bayesian optimization. J. Asian Archit. Build. Eng. 2025, 24, 751–774. [Google Scholar] [CrossRef]
  14. Cao, Y.; Ashuri, B. Predicting the Volatility of Highway Construction Cost Index Using Long Short-Term Memory. J. Manag. Eng. 2020, 36, 04020020. [Google Scholar] [CrossRef]
  15. Alzara, M.; Gihad, N.; Abdou, H.; Soltan, A.; Hilal, A.; Ehab, A. Macroeconomic-aware forecasting of construction costs in developing countries: Using gated recurrent unit and long short-term memory deep learning framework. PLoS ONE 2025, 20, e0333189. [Google Scholar] [CrossRef]
  16. Wang, J.; Qu, Z.; Lee, C.-Y.; Skitmore, M. Highway construction cost index forecasting: A hybrid VMD–LSTM–GRU method. Constr. Manag. Econ. 2025, 43, 849–863. [Google Scholar] [CrossRef]
  17. Kim, J.-S. AI-Powered Forecasting of Environmental Impacts and Construction Costs to Enhance Project Management in Highway Projects. Buildings 2025, 15, 2546. [Google Scholar] [CrossRef]
  18. Makridakis, S.; Spiliotis, E.; Assimakopoulos, V. Statistical and Machine Learning forecasting methods: Concerns and ways forward. PLoS ONE 2018, 13, e0194889. [Google Scholar] [CrossRef]
  19. Hewamalage, H.; Ackermann, K.; Bergmeir, C. Forecast evaluation for data scientists: Common pitfalls and best practices. Data Min. Knowl. Discov. 2023, 37, 788–832. [Google Scholar] [CrossRef] [PubMed]
  20. Jezzini, Y.; Assaad, R.H.; El-adaway, I.H. Modeling Framework to Quantify and Gauge Project Cost Risks due to Construction Material Price Volatilities Using Predictive Probabilistic Deep-Learning Algorithms and Stochastic Risk Modeling. J. Constr. Eng. Manag. 2025, 151, 04025071. [Google Scholar] [CrossRef]
  21. Senić, A.; Dobrodolac, M.; Stojadinović, Z. Development of Risk Quantification Models in Road Infrastructure Projects. Sustainability 2024, 16, 7694. [Google Scholar] [CrossRef]
  22. Ariza, V.; Zavala, G. Quantitative Risk Analysis Framework for Cost and Time Estimation in Road Infrastructure Projects. Infrastructures 2025, 10, 139. [Google Scholar] [CrossRef]
  23. Zavala, G.; Flores, V.A.; Santos, R.; Cano, J.B. Stochastic Cost Estimation in Transportation Infrastructure Projects Using Monte Carlo Simulation and Correlated Risk Variables. Future Transp. 2025, 5, 176. [Google Scholar] [CrossRef]
  24. Faghih, S.A.M.; Kashani, H. Forecasting Construction Material Prices Using Vector Error Correction Model. J. Constr. Eng. Manag. 2018, 144, 04018075. [Google Scholar] [CrossRef]
  25. Khalfaoui, R.; Tiwari, A.K.; Kablan, S.; Hammoudeh, S. Interdependence and lead-lag relationships between the oil price and metal markets: Fresh insights from the wavelet and quantile coherency approaches. Energy Econ. 2021, 101, 105421. [Google Scholar] [CrossRef]
  26. Castelblanco, G.; Fenoaltea, E.M.; De Marco, A.; Chiaia, B. Influence of macroeconomic factors on construction costs: An analysis of project cases. Constr. Manag. Econ. 2025, 43, 196–212. [Google Scholar] [CrossRef]
  27. Sheikhkhoshkar, M.; El-Haouzi, H.B.; Aubry, A.; Hamzeh, F.; Rahimian, F. A data-driven and knowledge-based decision support system for optimized construction planning and control. Autom. Constr. 2025, 173, 106066. [Google Scholar] [CrossRef]
  28. Cheng, M.-Y.; Khasani, R.R. Least Square Moment Balanced Machine: A New Approach to Estimating Cost to Completion for Construction Projects. J. Inf. Technol. Constr. 2024, 29, 503–524. [Google Scholar] [CrossRef]
  29. Shahandashti, S.M.; Ashuri, B. Highway Construction Cost Forecasting Using Vector Error Correction Models. J. Manag. Eng. 2016, 32, 04015040. [Google Scholar] [CrossRef]
  30. INEI. Reporte de Precios al Consumidor e Indices Unificados de Construcción. 2025. Available online: https://m.inei.gob.pe/estadisticas/indice-tematico/price-indexes/ (accessed on 1 April 2026).
  31. Federal Reserve Bank. Crude Oil Prices: Brent—Europe (DCOILBRENTEU). 2026. Available online: https://fred.stlouisfed.org/series/DCOILBRENTEU (accessed on 2 April 2026).
  32. Pandit, S.; Luo, X. A novel prediction model to evaluate the dynamic interrelationship between gold and crude oil. Int. J. Data Sci. Anal. 2025, 20, 1161–1182. [Google Scholar] [CrossRef]
  33. Jeong, M.-H.; Lee, T.-Y.; Jeon, S.-B.; Youm, M. Highway Speed Prediction Using Gated Recurrent Unit Neural Networks. Appl. Sci. 2021, 11, 3059. [Google Scholar] [CrossRef]
  34. Batur Dinler, Ö.; Aydin, N. An Optimal Feature Parameter Set Based on Gated Recurrent Unit Recurrent Neural Networks for Speech Segment Detection. Appl. Sci. 2020, 10, 1273. [Google Scholar] [CrossRef]
  35. Shi, T.; Shide, K. A comparative analysis of LSTM, GRU, and Transformer models for construction cost prediction with multidimensional feature integration. J. Asian Archit. Build. Eng. 2026, 25, 634–649. [Google Scholar] [CrossRef]
  36. Zafar, N.; Haq, I.U.; Chughtai, J.-R.; Shafiq, O. Applying Hybrid Lstm-Gru Model Based on Heterogeneous Data Sources for Traffic Speed Prediction in Urban Areas. Sensors 2022, 22, 3348. [Google Scholar] [CrossRef] [PubMed]
  37. Zarzycki, K.; Ławryńczuk, M. LSTM and GRU Neural Networks as Models of Dynamical Processes Used in Predictive Control: A Comparison of Models Developed for Two Chemical Reactors. Sensors 2021, 21, 5625. [Google Scholar] [CrossRef] [PubMed]
  38. Zhao, J. Time-varying impact of geopolitical risk on natural resources prices: Evidence from the hybrid TVP-VAR model with large system. Resour. Policy 2023, 82, 103467. [Google Scholar] [CrossRef]
  39. Le, T.-H.; Boubaker, S.; Bui, M.T.; Park, D. On the volatility of WTI crude oil prices: A time-varying approach with stochastic volatility. Energy Econ. 2023, 117, 106474. [Google Scholar] [CrossRef]
  40. Ghorbany, S.; Yousefi, S.; Noorzai, E. Evaluating and Optimizing Performance of Public-Private Partnership Projects Using Copula Bayesian Network. Eng. Constr. Archit. Manag. 2024, 31, 290–323. [Google Scholar] [CrossRef]
  41. Li, Y.; Dong, Y.; Guo, H. Copula-based multivariate renewal model for life-cycle analysis of civil infrastructure considering multiple dependent deterioration processes. Reliab. Eng. Syst. Saf. 2023, 231, 108992. [Google Scholar] [CrossRef]
  42. Dong, J.; Chen, Y.; Guan, G. Cost Index Predictions for Construction Engineering Based on LSTM Neural Networks. Adv. Civ. Eng. 2020, 2020, 6518147. [Google Scholar] [CrossRef]
Figure 1. Methodological workflow for comparative forecasting and dynamic relationship analysis of construction material price indices.
Figure 1. Methodological workflow for comparative forecasting and dynamic relationship analysis of construction material price indices.
Buildings 16 02124 g001
Figure 2. Monte Carlo Simulation for the latest test period using the Bi-GRU-based simulation.
Figure 2. Monte Carlo Simulation for the latest test period using the Bi-GRU-based simulation.
Buildings 16 02124 g002
Figure 3. Monte Carlo Simulation for the latest test period using the Random Walk-based simulation.
Figure 3. Monte Carlo Simulation for the latest test period using the Random Walk-based simulation.
Buildings 16 02124 g003
Figure 4. Tornado sensitivity diagrams for the latest test-period simulation: (a) Bi-GRU; (b) Random Walk.
Figure 4. Tornado sensitivity diagrams for the latest test-period simulation: (a) Bi-GRU; (b) Random Walk.
Buildings 16 02124 g004
Table 1. Case-study scope and material basket used in the probabilistic analysis.
Table 1. Case-study scope and material basket used in the probabilistic analysis.
ComponentDescription
Project typeRoad infrastructure bridge project
Case-study locationPeru
Main analytical scopeSelected critical material basket
Included inputsStructural steel, construction steel, cement, and diesel-related cost signals
Excluded inputsAsphalt/bitumen and other materials not dominant in the selected bridge basket
Main cost-risk driversConstruction steel and structural steel
Transferability conditionMore applicable to bridge or road infrastructure components with steel-intensive cost structures
Main limitation for extrapolationNot directly generalizable to road projects dominated by asphalt, earthworks, or pavement layers
Table 2. Forecasting performance of the Random Walk model on reconstructed input index levels.
Table 2. Forecasting performance of the Random Walk model on reconstructed input index levels.
DatasetMaterialMAERMSEMAPE (%)
ValidationStructural-Steel_Index21.441327.16822.600
ValidationDiesel_Index37.805446.43283.294
ValidationCement_Index4.09926.03650.759
ValidationConstruction-Steel_Index25.264235.09492.858
TestStructural-Steel_Index17.777924.95342.135
TestDiesel_Index15.950020.23141.455
TestCement_Index1.37252.78900.230
TestConstruction-Steel_Index18.386725.80682.135
Table 3. Forecasting performance of the Bi-GRU model on reconstructed input index levels.
Table 3. Forecasting performance of the Bi-GRU model on reconstructed input index levels.
DatasetMaterialMAERMSEMAPE (%)
ValidationStructural-Steel_Index23.049128.12822.797
ValidationDiesel_Index38.589347.74143.372
ValidationCement_Index4.43956.06490.822
ValidationConstruction-Steel_Index26.101835.68632.959
TestStructural-Steel_Index19.604726.44882.358
TestDiesel_Index14.724018.76581.338
TestCement_Index1.71502.82110.286
TestConstruction-Steel_Index20.137426.80242.351
Table 4. Relative RMSE variation with respect to the baseline model.
Table 4. Relative RMSE variation with respect to the baseline model.
DatasetMaterialRelative RMSE Variation (%)
ValidationStructural-Steel_Index3.53
ValidationDiesel_Index2.82
ValidationCement_Index0.47
ValidationConstruction-Steel_Index1.69
TestStructural-Steel_Index5.99
TestDiesel_Index−7.24
TestCement_Index1.15
TestConstruction-Steel_Index3.86
Table 5. Optimal lead–lag structure between oil and selected construction inputs.
Table 5. Optimal lead–lag structure between oil and selected construction inputs.
Dependent VariableOptimal Lag (Months)Correlation *
Structural-Steel_Index60.165
Diesel_Index20.290
Cement_Index50.205
Construction-Steel_Index60.187
* These lag values should be interpreted as exploratory empirical monitoring horizons rather than as statistically validated lag structures or structural supply-chain delay estimates.
Table 6. Volatility–correlation summary for lag-adjusted rolling relationships.
Table 6. Volatility–correlation summary for lag-adjusted rolling relationships.
Window (Months)MaterialLag Used (Months)SeriesCorrelation
12Structural-Steel_Index6Oil volatility vs. rolling Oil(t − 6)-Structural-Steel_Index(t) correlation0.158
12Diesel_Index2Oil volatility vs. rolling Oil(t − 2)-Diesel_Index(t) correlation0.224
12Cement_Index5Oil volatility vs. rolling Oil(t − 5)-Cement_Index(t) correlation0.331
12Construction-Steel_Index6Oil volatility vs. rolling Oil(t − 6)-Construction-Steel_Index(t) correlation0.160
Table 7. Monte Carlo summary of total material cost by split and forecasting model.
Table 7. Monte Carlo summary of total material cost by split and forecasting model.
SplitModelP50 Total Cost (PEN)P90 Total Cost (PEN)P95 Total Cost (PEN)Max Total Cost (PEN)Contingency
P90 − P50 (PEN)
Contingency
P95 − P50 (PEN)
TestBi-GRU84,880,373.9188,460,274.8189,631,531.2896,784,156.503,579,900.904,751,157.37
TestRandom Walk85,009,564.1288,552,641.7389,720,742.9196,530,033.953,543,077.624,711,178.79
ValidationBi-GRU85,012,537.5189,573,001.9590,638,927.6598,097,476.304,560,464.445,626,390.14
ValidationRandom Walk85,318,620.4589,855,126.1790,916,549.6997,966,338.554,536,505.725,597,929.24
Table 8. Probabilistic cost results for the latest test period (December 2025).
Table 8. Probabilistic cost results for the latest test period (December 2025).
PeriodModelDeterministic Total Cost (PEN)P50 Total Cost (PEN)P90 Total Cost (PEN)P95 Total Cost (PEN)Max Total Cost (PEN)Contingency P90 − P50 (PEN)Contingency P95 − P50 (PEN)
December 2025Bi-GRU80,767,416.0181,384,837.5683,654,970.3684,478,312.0187,733,194.612,270,132.803,093,474.46
December 2025Random Walk80,684,156.7881,466,666.5483,665,782.3384,616,037.6388,667,436.712,199,115.803,149,371.09
Table 9. Sensitivity ranking of critical inputs in the latest test-period simulation.
Table 9. Sensitivity ranking of critical inputs in the latest test-period simulation.
ModelMaterialSpearman Correlation with Total CostAbsolute Spearman
Bi-GRUCement_Index0.29920.2992
Bi-GRUDiesel_Index0.54390.5439
Bi-GRUStructural-Steel_Index0.79090.7909
Bi-GRUConstruction-Steel_Index0.90150.9015
Random WalkCement_Index0.22650.2265
Random WalkDiesel_Index0.54940.5494
Random WalkStructural-Steel_Index *0.81400.8140
Random WalkConstruction-Steel_Index *0.89450.8945
* The dominance of steel-related inputs reflects the combined effect of their cost participation in the selected bridge basket and their simulated price variability. Therefore, their influence should not be interpreted as a volatility effect alone.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Ariza Flores, V.A.; Pinedo, D.; Orellana, A.; Pinedo, A. Construction Input Price Forecasting for Probabilistic Contingency Estimation in a Road Infrastructure Bridge Case Study. Buildings 2026, 16, 2124. https://doi.org/10.3390/buildings16112124

AMA Style

Ariza Flores VA, Pinedo D, Orellana A, Pinedo A. Construction Input Price Forecasting for Probabilistic Contingency Estimation in a Road Infrastructure Bridge Case Study. Buildings. 2026; 16(11):2124. https://doi.org/10.3390/buildings16112124

Chicago/Turabian Style

Ariza Flores, Victor Andre, Diego Pinedo, Alan Orellana, and Amador Pinedo. 2026. "Construction Input Price Forecasting for Probabilistic Contingency Estimation in a Road Infrastructure Bridge Case Study" Buildings 16, no. 11: 2124. https://doi.org/10.3390/buildings16112124

APA Style

Ariza Flores, V. A., Pinedo, D., Orellana, A., & Pinedo, A. (2026). Construction Input Price Forecasting for Probabilistic Contingency Estimation in a Road Infrastructure Bridge Case Study. Buildings, 16(11), 2124. https://doi.org/10.3390/buildings16112124

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop