Abstract
Rapid and reliable prediction of embodied carbon emissions is essential for supporting sustainable design decision-making and reducing the environmental impacts of building engineering projects. However, existing studies have mainly focused on single buildings, with limited attention to project-level prediction and variations in information availability across design stages. To address this gap, this study developed a machine learning framework for project-level embodied carbon prediction based on a dataset of 78 projects involving 426 individual buildings. Using project attributes, scale indicators, structural characteristics, material quantities, and construction-related information, nine machine learning models were developed for the schematic design stage and the construction drawing design stage. Two residual-corrected weighted ensemble models were further introduced to improve predictive performance. The results show that the Extra Trees–KNN residual-corrected weighted ensemble model achieved the best performance at the construction drawing design stage, with a test-set R2 of 0.949. SHAP analysis further revealed a stage-dependent shift in dominant drivers: gross floor area and land area dominated at the schematic design stage, whereas concrete and reinforcement quantities became the leading predictors at the construction drawing design stage. The proposed framework provides interpretable and stage-specific quantitative support for low-carbon design decision-making, thereby facilitating embodied carbon reduction and the transition toward a more sustainable built environment.
1. Introduction
1.1. Research Background
With the continued advancement of global climate governance and the deepening green and low-carbon transition of the construction industry, the building sector has become a critical contributor to carbon mitigation efforts [1,2]. Building activities are characterized by intensive resource consumption, long supply chains, and multiple emission-related processes. Their carbon emissions arise not only from operational energy use, but also to a large extent from upstream stages, including building material production, material transportation, and on-site construction. According to the 2024 edition of the Research Report on Carbon Emissions in China’s Urban and Rural Development Sector, which reports national data for 2022, carbon emissions from building operation and building-industry construction in China totaled 5.13 billion tCO2, accounting for 48.3% of the nation’s energy-related carbon emissions. Of this total, building-industry construction contributed 2.82 billion tCO2, equivalent to approximately 54.9% of the combined emissions from building operation and building-industry construction [3]. The 2025 edition reported that building-industry construction emissions were 2.78 billion tCO2 in 2024, of which 1.81 billion tCO2 originated from house-building construction [4]. As improvements in operational energy efficiency continue, the focus of carbon mitigation across the building life cycle is gradually shifting from the operational stage to these upstream processes [5,6]. Reducing embodied carbon emissions is therefore an important component of sustainable building development. Because embodied carbon is generated early, becomes locked in rapidly, and is costly to modify afterward, its mitigation potential depends heavily on decisions made during the design stage [7,8]. Therefore, rapid and reliable prediction of building embodied carbon under incomplete design information has become a key issue for supporting low-carbon design and scheme optimization [9].
At present, most building carbon assessments are grounded in the life cycle assessment (LCA) framework, which provides a systematic basis for evaluating the environmental impacts of building activities [2,10]. However, conventional detailed accounting methods generally rely on relatively complete bills of quantities, material consumption records, and construction activity data. Their efficiency and responsiveness are therefore constrained when project information is still limited in the early stages [9,11,12]. By contrast, the stages of project initiation, conceptual design, and design development require a technical approach capable of rapidly estimating carbon emission levels under limited information, thereby supporting low-carbon scheme selection, structural optimization, and material-related decision-making [13,14,15].
In engineering practice, low-carbon decisions are typically made for projects composed of multiple buildings and ancillary works rather than for isolated buildings. Compared with a single building, a construction project exhibits stronger overall integration in terms of site conditions, development intensity, building configuration, spatial organization, and construction coordination. Its embodied carbon formation mechanism is therefore more complex and cannot be regarded as a simple aggregation of emissions from individual buildings. Accordingly, project-level embodied carbon research is more consistent with real engineering decision-making scenarios and is better suited to supporting project-wide low-carbon planning, scheme comparison, and implementation control. Existing prediction studies have still focused predominantly on individual residential or office buildings, while project-level research remains limited [13,16,17]. At the same time, the information available differs substantially across design stages [18,19]. Therefore, developing a high-quality embodied carbon dataset for building engineering projects and conducting stage-specific prediction studies are of considerable importance for improving the scientific basis and practical applicability of early-stage low-carbon decision-making [20].
1.2. Literature Review
Systematic research has been conducted worldwide on the accounting, rapid estimation, and data-driven prediction of building embodied carbon emissions. Overall, the existing literature can be broadly grouped into three research streams:
- (1)
- Embodied carbon accounting based on life cycle assessment (LCA);
- (2)
- Rapid estimation methods for design-stage applications; and
- (3)
- Machine learning-based prediction and interpretation of carbon emissions. These three research streams are discussed sequentially below.
The first research stream focuses on embodied carbon accounting, with most studies grounded in the LCA framework [21,22]. From the perspectives of material production, transportation, construction, and even the full life cycle, researchers have quantified carbon emissions for residential buildings, office buildings, prefabricated buildings, and other building types, while also examining the effects of structural form, building scale, and material input on embodied carbon emissions [19,23,24]. In addition, some studies have developed automated estimation tools and database-supported systems based on LCA to improve the efficiency and practical usability of building carbon assessment [25]. Existing evidence indicates that major materials, particularly steel and concrete, are the dominant sources of embodied carbon in buildings, and that early design decisions exert a substantial influence on subsequent emission levels [26]. Multi-sample statistical analyses have further shown relatively stable associations between embodied carbon emissions and variables such as building height, structural type, and major material input [27]. Other studies have emphasized that embodied carbon is jointly influenced by multiple factors, including floor area, use characteristics, envelope systems, and building service configurations, reflecting marked complexity and multi-variable coupling [28]. However, such studies usually rely heavily on detailed bills of quantities, material inventories, and construction activity data. They are therefore more suitable for detailed design stages or data-rich scenarios, but remain constrained by limited timelines when design information is still incomplete.
The second research stream concerns rapid estimation methods for early design-stage applications. In response to these limitations, some scholars have explored rapid estimation approaches for early design stages. Victoria and Perera [9] proposed a parametric prediction model for embodied carbon at the early design stage, demonstrating that rapid estimation can still be achieved using key design parameters even in the absence of detailed quantity information. Cang et al. [18] developed a regression-based prediction model using samples of residential buildings with different structural forms, confirming the feasibility of early-stage prediction based on major material intensity or basic building characteristics. These studies provide an important foundation for the early identification of building embodied carbon emissions. However, most existing models include relatively few variables, and their applicability and generalization remain limited.
The third research stream involves machine learning-based prediction and interpretation. With the continued advancement of data-driven methods and artificial intelligence, machine learning has gradually expanded into LCA-related research [29] and has demonstrated strong potential for classification, prediction, and pattern discovery in complex systems. In the building sector, machine learning has been applied to the prediction of operational carbon emissions, construction-stage carbon emissions, and life cycle carbon emissions. Studies [30,31] showed that support vector regression and random forest models have considerable potential for building carbon prediction. Using residential building samples from China, Zhang et al. [13] developed machine learning models for different design stages and showed that predictive accuracy can improve substantially as more design information becomes available. Starting from key building materials, Su et al. [32] developed machine learning-based prediction models and tools for design-stage applications, further extending the practical scope of rapid embodied carbon prediction. Zhang et al. [16] generated a dataset of concrete frame structures through parametric design, trained multiple machine learning models to predict embodied carbon emissions, and combined simplified models with a genetic algorithm for practical application. Luo et al. [20] further proposed a progressive prediction strategy across design stages, emphasizing that differences in information completeness should be explicitly reflected in the prediction framework. Meanwhile, some studies have introduced interpretable methods such as SHAP to identify key influencing factors and thereby enhance the engineering interpretability of prediction results [13,33]. Nevertheless, most of these studies focus on individual residential or office buildings, or on specific structural systems. Their variable systems are relatively simplified, and their analytical scale usually remains at the level of a single building or local structural component. As a result, insufficient attention has been paid to the integrated object of projects composed of multiple buildings and ancillary works, which limits the direct transferability of their findings to project-level decision-making contexts.
Overall, existing studies have generated substantial insights into embodied carbon accounting, rapid prediction, and key factor identification. However, several limitations remain.
Most existing studies use individual buildings as the unit of analysis, and project-level research is still scarce [34,35,36]. As a result, they cannot adequately capture the overall carbon emission characteristics arising from the combined effects of construction coordination in real engineering projects, making current models difficult to apply directly to project-level scenarios.
Although many studies emphasize the importance of early-stage prediction, they still lack a systematic treatment of differences in information completeness across different levels of design development. A stage-adapted prediction framework aligned with actual engineering design workflows is still lacking [35].
Existing sample datasets remain limited in terms of boundary consistency and project-level integrity. Existing datasets are typically developed around the main structural system or a small number of key materials [20,34], and their coverage of emissions from civil works, mechanical and electrical installations, decoration and finishing, waterproofing and roofing, and ancillary works remains relatively limited. This weakens both the integrity of the training basis and the comparability of model outputs [13,16]. For project-level studies, insufficient coverage of carbon emission factors or inconsistencies in accounting scope may still yield high goodness of fit, but will substantially undermine the engineering interpretability and transferability of the resulting models [37].
Based on the above review, the central challenge is to develop a rapid embodied carbon estimation framework that is aligned with real-world project-level decision-making. Such a framework should satisfy four key requirements:
- (1)
- Use complete building engineering projects as the unit of analysis;
- (2)
- Adopt a sufficiently comprehensive and consistent accounting boundary;
- (3)
- Adapt the predictor set to the information available at different design stages; and
- (4)
- Provide interpretable outputs that can directly support engineering decision-making.
1.3. Research Scope and Contributions
To address the limitations of existing studies in terms of project-level analysis, stage-adapted modeling, and the integrity of embodied carbon labels, this study takes building engineering projects as the unit of analysis and develops an integrated research framework centered on accounting, database development, estimation, and interpretation, thereby systematically investigating project-level embodied carbon emissions. The main contributions of this study are threefold.
The analytical scope is extended from individual buildings to building engineering projects, enabling project-level data to be examined as an integrated whole.
A more comprehensive embodied carbon labeling system is developed for project-level research. The dataset comprises real completed projects in China over the past five years, covering residential, public, and industrial buildings. For carbon emission factors, we comprehensively utilize national standards, industry standards, and an expanded factor database compiled by our team based on literature reviews (containing a total of over 2000 carbon emission factors).
A project-level embodied carbon prediction framework is developed for different design stages. Combined with SHAP analysis, the framework further identifies key influencing factors and their effect patterns across design stages, thereby providing quantitative support for early-stage low-carbon scheme selection and design optimization in building engineering projects.
2. Methodology
2.1. Research Scope
To enable rapid prediction of embodied carbon emissions and identification of key drivers at the project level, this study develops an integrated research framework consisting of carbon accounting, sample database construction, stage-specific modeling, and result interpretation. First, based on the life cycle assessment (LCA) concept and the characteristics of data acquisition in building engineering projects, an embodied carbon accounting method is established for project-level applications [22]. Second, drawing on bills of quantities, construction drawings, budget documents, and related technical materials from completed projects, information on basic project attributes, project scale, structural characteristics, material consumption, and construction organization is extracted to build a project-level sample database for model training [32]. Furthermore, given the substantial differences in information completeness between the schematic design stage and the construction drawing design stage, corresponding feature sets are developed for each stage, and multiple machine learning algorithms are employed for stage-specific prediction and performance comparison [13,35]. Building on the single-model results, a residual-corrected weighted ensemble strategy is further introduced to enhance the ability of the models to capture complex nonlinear relationships and local prediction errors. Finally, SHAP is applied to the optimal model to interpret the underlying formation mechanism and stage-dependent evolution of embodied carbon emissions in building engineering projects from the perspectives of feature contribution, effect direction, and nonlinear response. The overall research framework is illustrated in Figure 1.
Figure 1.
Overall research framework.
2.2. Embodied Carbon Accounting Method and Dataset Construction
2.2.1. Accounting Method
In this study, building engineering projects are taken as the unit of analysis, and embodied carbon emissions during the construction phase are quantified at the project level. Based on the life cycle assessment (LCA) approach and the characteristics of project data acquisition [2], the system boundary includes material production (A1–A3), transportation to the construction site (A4), and on-site construction and installation processes (A5), which together constitute the construction/materialization phase of buildings [38]. Operational stages (B1–B7), end-of-life stages (C1–C4), and benefits and loads beyond the system boundary (Module D) are excluded from the present analysis. The spatial boundary covers all building units and ancillary works involved in the project construction process, and total greenhouse gas emissions are consistently expressed in carbon dioxide equivalent (CO2e). Unlike simplified studies that focus only on individual buildings or major structural materials, this study incorporates, to the extent possible, emissions from civil works, mechanical and electrical installations, decoration and finishing works, waterproofing and roofing works, and ancillary works, so as to provide a more realistic representation of the overall embodied carbon level during project construction.
Methods for quantifying building carbon emissions mainly include direct measurement, the carbon emission factor method, input–output analysis, and hybrid methods. Given the research scope and the available data, the carbon emission factor method is adopted in this study to estimate the embodied carbon emissions of engineering projects [39].
The carbon emission factor method can be expressed as:
where is the data of the th activity, and is the corresponding carbon emission factor. In the context of building engineering projects, activity data may include material quantities, transport mass and distance, construction machinery shifts, and energy consumption, while the corresponding factors represent unit carbon emission coefficients for different materials, transportation modes, machinery types, or energy sources.
This study integrates the Standard for building carbon emission calculation (GB/T 51366–2019) [40], the Code for carbon reduction construction of building engineering (T/CCIAT 0104–2025) [41], and an extended factor database compiled by the research team containing more than 2000 carbon emission factors. Carbon emission factors associated with material production, material transportation, and construction are uniformly organized and standardized. Factor selection follows the principles of standard priority, engineering applicability, and stage-specific matching. The consistency of accounting results is further improved through the standardization of factor names, units, and system boundaries.
2.2.2. Case Study
To verify the applicability of the embodied carbon accounting boundary, calculation model, and carbon emission factor system developed for building engineering projects, this study selected a cultural and sports center project as a case study. The project has a total land area of 42,781 m2 and a gross floor area of approximately 14,974 m2, including an underground fire water tank of about 219 m2, a cultural and sports center building of about 14,663 m2, and ancillary buildings of about 92 m2. Following the accounting boundary defined above, the project’s embodied carbon emissions were divided into three stages: building material production, material transportation, and construction. Based on the carbon emission factor method, detailed carbon calculations were carried out for 1085 types of construction materials in the production and transportation stages and for 148 types of construction machinery in the construction stage. The results are summarized in Table 1. Overall, building material production was the dominant source of embodied carbon emissions, accounting for 89.00% of the total, followed by the construction stage at 9.63%, while material transportation contributed only 1.37%. This case study not only demonstrates the applicability of the proposed accounting framework in a real engineering project, but also provides a concrete example for subsequent sample label construction.
Table 1.
Stage-wise embodied carbon emissions of the case project.
The dominance of the production stage in the present case is broadly consistent with previous materialization-phase LCA studies. Teng and Pan [42] reported that material production and prefabrication, transportation, and on-site construction accounted for 82.9%, 9.3%, and 7.8%, respectively, in a prefabricated high-rise residential building. Zhan et al. [43] reported that the production and transportation stages accounted for 85.73% and 8.55%, respectively, in a prefabricated residential building, while the construction stage accounted for approximately 5.72%, as calculated from the emission values reported in their study. Compared with these building-level cases, the present project exhibits a higher production-stage share (89.00%), a substantially lower transportation-stage share (1.37%), and a somewhat higher construction-stage share (9.63%). These differences may be related to project type, prefabrication level, transport distance, construction methods, system-boundary assumptions, and the unit of analysis. In particular, the previous studies assessed individual residential buildings, whereas the present study evaluates an entire building engineering project comprising multiple buildings and associated ancillary works.
2.2.3. Data Acquisition and Construction of the Project-Level Carbon Emission Sample Set
The raw data used in this study were primarily collected from bills of quantities, construction drawings, and related technical documents of completed building engineering projects in China over the past five years. For some projects, additional information was supplemented from budget documents, material statistics, and construction organization records to further compile data on basic project attributes, major quantities, and construction-related characteristics. The sample covers multiple project types, including residential, public, and industrial buildings. Based on the embodied carbon accounting boundary, calculation method, and carbon emission factor system established in the previous section, carbon emissions were calculated for each project across three stages: building material production, material transportation, and construction. These results were initially aggregated to form a project-level embodied carbon dataset comprising 79 samples. Following the data-cleaning procedure described in Section 3.2, one anomalous sample was removed, resulting in a final modeling dataset of 78 valid projects. To characterize the key factors affecting project-level embodied carbon emissions, 15 explanatory variables were selected. Their definitions, abbreviations, units or data types, and availability at each design stage are summarized in Table 2. The total embodied carbon emissions of each project were used as the target variable.
Table 2.
Definitions, units, and stage availability of the explanatory variables. ✓ indicates that the variable is available at the corresponding design stage, whereas – indicates that it is unavailable.
To ensure the reliability of the subsequent modeling results, data completeness checks and outlier identification were further conducted after label construction. The detailed data-cleaning procedure and corresponding results are presented in Section 3.
2.3. Prediction Model Development
To systematically compare the applicability of different machine learning algorithms for predicting embodied carbon emissions in building engineering projects, model development was conducted separately based on two feature scenarios: the schematic design stage and the construction drawing design stage. Under each feature scenario, nine machine learning models were developed. Accordingly, a total of 18 individual prediction models and two ensemble models were constructed to compare predictive performance across different design stages and algorithmic settings. The nine standalone algorithms were Linear Regression (LR), K-Nearest Neighbors (KNN), Support Vector Regression (SVR), Random Forest (RF), Extra Trees (ET), Gradient Boosting Regression (GB), XGBoost (XGB), LightGBM (LGBM), and CatBoost. Bayesian Ridge was additionally employed as the residual learner in the two residual-corrected ensemble models.
2.3.1. Prediction Framework for Different Design Stages
Separate feature sets were developed for the schematic design stage and the construction drawing design stage to reflect differences in information availability across stages. As summarized in Table 2, the schematic design-stage model uses ten project- and building-level variables, whereas the construction drawing design-stage model additionally incorporates total construction period, total machinery shifts, concrete consumption, reinforcement consumption, and mortar consumption.
2.3.2. Machine Learning Algorithms
Given that embodied carbon emissions in building engineering projects are jointly influenced by multiple factors, including project scale, structural type, material input, and construction organization, the relationships among variables may exhibit substantial nonlinearity and interaction effects. Several nonlinear algorithms evaluated in this study, particularly tree-based ensemble models and SVR, can implicitly represent such nonlinear joint relationships and possible conditional interactions among the included predictors. In addition, the sample size in this study is relatively limited, while both continuous and categorical variables are involved. Therefore, it was necessary to select algorithms that balance generalization ability, modeling stability, and nonlinear fitting capability for comparative analysis. Based on these considerations, nine representative regression algorithms were selected to develop the prediction models: (1) a linear model, Linear Regression (LR); (2) an instance-based learning method, K-Nearest Neighbors (KNN) [44]; (3) a kernel-based method, Support Vector Regression (SVR) [45]; (4) bagging-based ensemble tree models, namely Random Forest (RF) [46] and Extra Trees (ET) [47]; and (5) boosting-based ensemble tree models, including Gradient Boosting Regression (GB) [48], XGBoost [49], LightGBM [50], and CatBoost [51]. ET and RF both belong to the family of tree-based ensemble models, but ET introduces stronger randomness in node splitting, thereby increasing diversity among base learners. GB improves predictive accuracy by iteratively fitting the residuals of the previous model and progressively combining weak learners. XGBoost further incorporates regularization terms into the objective function, thereby enhancing both fitting capability and generalization performance. LightGBM improves training efficiency and memory usage through histogram-based splitting and a leaf-wise growth strategy. CatBoost is particularly effective in handling categorical variables and reduces prediction bias and overfitting risk through ordered boosting.
To evaluate predictive performance, the coefficient of determination (R2), mean absolute error (MAE), and root mean square error (RMSE) were adopted as the primary evaluation metrics. These metrics provide a comprehensive assessment of model performance from the perspectives of goodness of fit, average prediction error, and error dispersion.
Building on the individual model predictions, a residual-corrected weighted ensemble model was further developed to improve the prediction accuracy of project-level embodied carbon emissions [52]. After the standalone-model comparison, Extra Trees and SVR were combined at the schematic design stage, whereas Extra Trees and KNN were combined at the construction drawing design stage. For each stage, the predictions of the two base models were integrated through a nonnegative weighted average, with the weights constrained to sum to one. Bayesian Ridge was then trained to estimate the residuals between the observed values and the weighted predictions, and the predicted residuals were added to the weighted predictions to obtain the final ensemble outputs.
2.3.3. Data Processing and Hyperparameter Optimization
During data preprocessing, numerical variables were standardized to reduce the influence of differences in scale and value range on model training, while categorical variables were transformed into model-readable inputs using one-hot encoding [53].
For dataset partitioning, the original dataset was split into a training set and a test set at a ratio of 8:2 [54]. Given the relatively small sample size and the evident heterogeneity in the distribution of the target variable, the target variable was first grouped into quintile-based bins, after which stratified sampling was performed according to the binning results [55]. This procedure was intended to maintain consistency in the target-variable distribution between the training and test sets as much as possible, thereby improving the stability of model evaluation.
Compared with a single random split, k-fold cross-validation makes more efficient use of limited data [56] and helps reduce the incidental bias introduced by a one-time partition; it is therefore widely used in small-sample modeling problems [57]. In this study, ten-fold cross-validation was applied within the training set. The average performance across the 10 folds was then used as an important basis for hyperparameter tuning and model performance evaluation.
To further improve predictive performance, hyperparameter optimization was conducted for the candidate models within the training set. Among the common hyperparameter tuning methods, grid search is computationally expensive, random search is more sensitive to stochastic variation, and evolutionary algorithms are relatively complex to implement. By contrast, Bayesian optimization can continuously update the parameter space based on previous trials and prioritize the exploration of potentially better regions, thereby generally achieving higher search efficiency under a limited number of evaluations [58]. Accordingly, Bayesian optimization was adopted in this study for hyperparameter tuning, with the optimal parameter combinations determined in conjunction with ten-fold cross-validation on the training set. Optimization hyperparameters for different feature combination scenarios are presented in the Supplementary Materials. All data preprocessing, machine learning modeling, hyperparameter optimization, and model interpretation were performed using Python 3.13.5.
2.3.4. SHAP-Based Interpretability Analysis
To further reveal the underlying mechanism of machine learning predictions and enhance the engineering interpretability of model outputs, SHAP (Shapley Additive Explanations) was employed to interpret the optimal prediction models [33,59]. Derived from the concept of Shapley values in cooperative game theory, SHAP explains model predictions by decomposing the final output into the sum of the marginal contributions of individual input features, thereby providing a unified interpretation framework for complex black-box models.
3. Dataset Construction
3.1. Data Collection
After data compilation, an initial dataset of 79 building engineering project samples was obtained. Following the data-cleaning procedure described in Section 3.2, one anomalous sample was removed, leaving 78 valid projects involving 426 individual buildings for model development and evaluation. The dataset covers information on basic project attributes, project scale, structural characteristics, construction organization, and major material consumption. The sample size distribution was controlled to ensure representativeness, excluding micro-scale projects with a gross floor area of less than 1000 m2 and exceptionally large special projects exceeding 500,000 m2. The regional distribution was also balanced, covering both highly active construction regions, such as East China and North China, and relatively less active regions, such as Northwest China and Southwest China, so as to reduce the risk of limited model generalizability caused by regional homogeneity. Overall, the dataset reasonably captures the project-level characteristics of embodied carbon emissions in building engineering projects. A coverage analysis of the samples is presented in Figure 2.
Figure 2.
Sample coverage analysis: (a) Land-use category; (b) Regional composition; (c) Seismic fortification intensity; (d) Structural types.
3.2. Data Cleaning
Data cleaning was conducted in three steps. First, the initial dataset of 79 projects was examined for missing values, duplicate records, unit consistency, and logical consistency. No missing values or duplicate project records were identified.
Second, potential abnormal samples were screened using two complementary methods. Total embodied carbon emissions were examined using the mean ± 3 standard deviations (μ ± 3σ) criterion, which identified one potential abnormal sample. In parallel, the Isolation Forest algorithm was applied to the numerical explanatory variables to identify multivariate anomalies, resulting in three potentially abnormal samples. The detection results are presented in Figure 3.
Figure 3.
Schematic illustration of outlier identification: (a) μ ± 3σ rule; (b) Isolation Forest algorithm.
Third, all samples identified by either method were retrospectively checked against the original engineering documents, including bills of quantities, construction drawings, budget documents, material statistics, and construction-organization records. The review focused on possible data-entry errors, unit inconsistencies, abnormal material quantities, and logical contradictions between project scale and resource consumption. Sample 35 was identified as anomalous by both methods and showed a pronounced deviation that could substantially affect model training. Therefore, this sample was removed. The other potentially abnormal samples were retained because their relatively high values were supported by their actual project scale, material consumption, or machinery input, and no data-entry or unit errors were found during document verification. The final modeling dataset therefore comprised 78 valid project samples, 15 explanatory variables, and one target variable. The distributions of the categorical variables are presented in Figure 2, while Table 3 summarizes the descriptive statistics of the numerical explanatory variables and the target variable in the final cleaned dataset.
Table 3.
Descriptive statistics of the numerical predictor features and the target variable.
3.3. Correlation Analysis
To reveal the correlation structure among numerical features and to provide a reference for subsequent model development and result interpretation, Spearman’s rank correlation analysis was performed, and the results are presented in Figure 4. Overall, features related to project scale, material consumption, and construction intensity exhibited strong correlations. In particular, gross floor area was highly correlated with concrete consumption and reinforcement consumption, with correlation coefficients of 0.91 and 0.88, respectively, while the correlation coefficient between concrete consumption and reinforcement consumption reached 0.92. These results indicate that an increase in project scale is typically accompanied by a concurrent increase in the input of major structural materials. The total number of machinery shifts also showed strong positive correlations with gross floor area, concrete consumption, and reinforcement consumption, suggesting that the intensity of construction machinery input is closely associated with project scale and resource consumption. It should be noted that the purpose of the correlation analysis in this study was primarily to identify the intrinsic relationships among variables and potential collinearity patterns, rather than to use correlation as the sole basis for feature selection. Given that each feature has clear engineering significance and represents a distinct aspect of project characteristics, highly correlated variables were not directly removed. Instead, their roles were comprehensively examined in the subsequent model comparison and interpretability analysis.
Figure 4.
Heatmap of Spearman’s rank correlation coefficients.
4. Results and Discussion
4.1. Predictive Performance at the Schematic Design Stage
To evaluate the predictive performance of different models for embodied carbon emissions in building engineering projects at the schematic design stage, a multi-model comparison was conducted based on the corresponding feature set, and the results are shown in Figure 5. Overall, substantial differences in predictive performance were observed among the models. Tree-based ensemble models performed relatively well, followed by kernel-based methods, whereas linear models showed comparatively weaker performance. According to the test-set results, CatBoost, Extra Trees, and Gradient Boosting achieved the best predictive performance, with R2 values of 0.760, 0.754, and 0.746, respectively. The prediction scatter plots further show that the better-performing tree-based ensemble models produced sample points that were more closely clustered around the ideal line, with most test samples falling within the ±10% error band. This indicates good agreement between predicted and actual values and suggests that tree-based ensemble models are better able to capture the nonlinear relationships between project-level embodied carbon emissions and early-stage design features. Overall, even in the absence of detailed quantity and material information, preliminary prediction of project-level embodied carbon emissions can still be achieved using only general project attributes, scale-related indicators, and basic structural information. Such predictions are therefore suitable for trend identification and scheme comparison at the early design stage [35,36].
Figure 5.
Performance comparison of different models at the schematic design stage: (a) LR; (b) KNN; (c) SVR; (d) LGBM; (e) ET; (f) RF; (g) GB; (h) XGB; and (i) CatBoost.
4.2. Predictive Performance at the Construction Drawing Design Stage
The predictive results of the models at the construction drawing design stage are presented in Figure 6. Overall, model performance was markedly better than that at the schematic design stage, indicating that the correspondence between input features and project-level embodied carbon emissions becomes substantially stronger as design information becomes more detailed [13,20]. Among all models, Extra Trees achieved the best performance, with a test-set R2 of 0.944 and corresponding MAE and RMSE values of 7.50 and 10.00, respectively. SVR and Gradient Boosting also demonstrated strong predictive capability, with test-set R2 values of 0.909 and 0.906, respectively. Overall, after incorporating variables such as concrete, reinforcement, and mortar consumption, as well as total construction period and total machinery shifts, the models were able to characterize the predictive relationship between detailed project information and project-level embodied carbon emissions more accurately. The improvement in predictive performance reflects not only the nonlinear modeling capability of the algorithms, but also the greater completeness of construction-drawing-stage information and the more direct physical relationship between the material- and machinery-related inputs and a subset of the activity-data terms used in the carbon accounting equation. As a result, the predictions progressed from preliminary trend identification to relatively high-accuracy quantitative estimation, thereby providing more direct quantitative support for construction drawing refinement, quantity verification, and emission reduction optimization.
Figure 6.
Performance comparison of different models at the construction drawing design stage: (a) LR; (b) KNN; (c) SVR; (d) LGBM; (e) ET; (f) RF; (g) GB; (h) XGB; and (i) CatBoost.
4.3. Ensemble Models
Building on the comparison of individual models, a residual-corrected weighted ensemble model was further developed to improve the prediction accuracy of embodied carbon emissions in building engineering projects. At the schematic design stage, the ensemble combined Extra Trees and SVR, whereas at the construction drawing design stage, Extra Trees and KNN were combined. In both cases, Bayesian Ridge was adopted as the residual learner. The predictive results of the two residual-corrected weighted ensemble models are presented in Figure 7. At the schematic design stage, the Extra Trees-SVR ensemble improved the test-set R2 to 0.778, while both MAE and RMSE decreased further. This indicates that, under conditions of relatively limited early-stage information, the ensemble strategy can effectively enhance the predictive performance for project-level embodied carbon emissions. By contrast, at the construction drawing design stage, the Extra Trees-KNN ensemble achieved a test-set R2 of 0.949, representing only a marginal improvement over the standalone Extra Trees model. This suggests that when major material quantities and construction-related information are already sufficiently available, a single model is capable of adequately capturing the formation mechanism of total project embodied carbon emissions, and the additional benefit of ensemble modeling becomes relatively limited. Therefore, the ensemble approach appears to offer greater practical value at earlier project stages [16]. Table 4 summarizes the best-performing individual models and ensemble models under the two prediction scenarios.
Figure 7.
Ensemble models under the two prediction scenarios: (a) Extra Trees–SVR ensemble at the schematic design stage; (b) Extra Trees–KNN ensemble at the construction drawing design stage.
Table 4.
Best-performing individual and ensemble models under the two prediction scenarios.
4.4. Model Interpretability
To elucidate the engineering implications of the model predictions, SHAP analysis was conducted for the best-performing ensemble model at each design stage, namely the Extra Trees-SVR model for the schematic design stage and the Extra Trees-KNN model for the construction drawing design stage. The analysis focused on the contribution magnitude, effect direction, and nonlinear response patterns of key features [33,36].
Figure 8 presents the SHAP feature importance bar plots for the ensemble models under the two prediction scenarios. At the schematic design stage, gross floor area had the highest relative importance, accounting for 45.2%, which was substantially higher than that of all other variables. It was followed by land area (16.1%), seismic intensity (8.4%), number of building units (5.6%), and structural type (5.4%), with the top five variables contributing a cumulative 80.7%. These results indicate that, at the schematic design stage, the identification of project-level embodied carbon emissions still depends primarily on overall scale, structural constraints, and building configuration characteristics. By contrast, at the construction drawing design stage, feature importance shifted markedly toward material input and construction intensity. The relative importance values of concrete consumption, reinforcement consumption, gross floor area, total machinery shifts, and mortar consumption were 26.1%, 21.8%, 14.1%, 11.2%, and 8.2%, respectively, and the cumulative contribution of the top five variables reached 81.4%. This suggests that, once major quantities and construction-related information become available, the explanatory focus of project-level embodied carbon emissions shifts from “scale and form” to “material input and construction intensity” [33]. Notably, total machinery shifts ranked fourth, exceeding variables such as total construction period, land area, number of building units, and maximum building height. This indicates that machinery input can effectively reflect the intensity of resource use during construction and serves as an important supplement to the interpretation of project-level carbon emissions. In addition, mortar consumption ranked among the top five variables, suggesting that, beyond major structural materials, work items related to enclosure systems, masonry, and plastering also make a non-negligible contribution to total project embodied carbon emissions.
Figure 8.
SHAP feature importance bar plots: (a) schematic design stage; (b) construction drawing design stage.
Figure 9 presents the SHAP summary beeswarm plots for the ensemble models under the two prediction scenarios. At the schematic design stage, samples with high values of gross floor area and land area were mainly distributed in the positive SHAP region, whereas samples with low values were more concentrated in the negative region. This indicates that an increase in project scale generally drives up predicted embodied carbon emissions [60]. At the same time, the number of building units, seismic intensity, and structural type also exhibited relatively clear directional effects, suggesting that, at the project level, carbon emissions are influenced not only by overall scale, but also by building organization, differences in structural systems, and seismic design requirements [61]. At the construction drawing design stage, samples with high values of concrete consumption, reinforcement consumption, total machinery shifts, and mortar consumption were more concentrated in the positive SHAP region, indicating that material quantities and construction resource inputs are direct drivers of increases in total project embodied carbon emissions. By contrast, the SHAP distributions of morphological indicators such as floor area ratio, average number of floors, and maximum building height contracted markedly toward zero at this stage. This suggests that once detailed quantity information becomes available, the marginal explanatory power of early-stage morphological indicators is substantially reduced [20].
Figure 9.
SHAP summary beeswarm plots: (a) schematic design stage; (b) construction drawing design stage.
Figure 10 presents the SHAP dependence plots of key numerical features for the ensemble models under the two prediction scenarios. At the schematic design stage, the SHAP values of gross floor area and land area increased rapidly from the low to medium ranges and then gradually leveled off in the high-value range. This indicates that expansions in overall project scale and site scale tend to continuously increase total project embodied carbon emissions, but with a clear pattern of diminishing marginal effects. This finding is consistent with previous research results [12,35,60]. The number of building units showed pronounced discreteness and step-like variation, suggesting that its effect is more susceptible to project organization and interactions with other variables. Although floor area ratio maintained an overall positive relationship with SHAP values, its response magnitude was relatively small, indicating that it serves more as a supplementary explanatory variable for development intensity. By contrast, at the construction drawing design stage, the dependence patterns of concrete consumption, reinforcement consumption, total machinery shifts, and gross floor area more directly reflect the formation mechanism associated with material and construction inputs. The SHAP values of concrete and reinforcement consumption increased continuously with increasing feature values and gradually flattened in the high-value range, indicating that major structural material input is a core driver of increasing total project embodied carbon emissions, while also exhibiting a certain marginal attenuation effect [62]. Total machinery shifts showed a smoother and more stable positive growth pattern, suggesting that this variable has relatively strong independence and stability in its contribution to model output and is more representative of a process-related driving factor at the construction stage. Gross floor area still maintained a positive effect, but its response magnitude was clearly weaker than that of concrete and reinforcement consumption, indicating that once detailed quantity information becomes available, its role shifts from a direct driving factor to a background explanatory variable reflecting project scale.
Figure 10.
SHAP dependence plots: (a) schematic design stage; (b) construction drawing design stage.
Overall, the SHAP results are consistent with established engineering understanding and previous embodied carbon studies [26,60,61,62]. At the schematic design stage, gross floor area and land area mainly function as aggregate indicators of project scale and development intensity. Larger projects generally require greater quantities of structural and non-structural materials, as well as more extensive construction inputs, which explains their positive contributions to predicted embodied carbon emissions. At the construction drawing design stage, concrete and reinforcement become the dominant predictors because they are used in large quantities in common structural systems and their production represents a substantial source of embodied carbon. Total machinery shifts provide an indicator of on-site equipment use and construction intensity, while mortar consumption reflects additional material inputs associated with masonry, enclosure, plastering, and finishing-related works. Therefore, the observed transition from scale- and form-related indicators to material- and construction-related indicators is consistent with the progressive refinement of project information across design stages. This stage-dependent consistency between the SHAP results and established engineering knowledge strengthens the practical interpretability of the proposed framework and its potential to support the identification of carbon-intensive design decisions, the prioritization of material-efficiency measures, and more informed sustainable decision-making across building design stages.
5. Conclusions
5.1. Main Findings
This study investigated the accounting, dataset construction, stage-adapted prediction, and model interpretation of embodied carbon emissions in building engineering projects. In response to current research gaps, including the lack of high-quality project-level labels, insufficient consideration of stage-specific differences, and limited engineering interpretability of prediction results, an integrated research framework of “accounting–database construction–prediction–interpretation” was established.
- (1)
- An embodied carbon accounting framework was established for building engineering projects. Using building engineering projects rather than individual buildings as the unit of analysis, the framework incorporated building material production, material transportation, and construction within a unified system boundary, while covering, to the extent possible, civil works, mechanical and electrical installations, decoration and finishing works, waterproofing and roofing works, and ancillary works.
- (2)
- A project-level sample database was developed, extending the research object from individual buildings to the project level. By integrating GB/T 51366–2019 [40], T/CCIAT 0104–2025 [41], and an extended database of more than 2000 carbon emission factors compiled by the research team, and based on bills of quantities, construction drawings, budget documents, material statistics, and construction organization records, a total of 78 building engineering project samples involving 426 individual buildings were compiled, providing a reliable data foundation for subsequent machine learning modeling.
- (3)
- A total of 20 machine learning models for project-level embodied carbon prediction were developed, including ensemble models, and the results demonstrated clear stage adaptability. At the schematic design stage, CatBoost achieved the best performance among the individual models, whereas Extra Trees performed best at the construction drawing design stage. After the ensemble models were introduced, the test-set R2 of the Extra Trees-SVR model at the schematic design stage increased to 0.778, while that of the Extra Trees-KNN model at the construction drawing design stage reached 0.949. These results indicate that machine learning can effectively support project-level embodied carbon prediction, and that the improvement brought by ensemble modeling is particularly valuable at early stages when information is relatively limited. The reported R2 values should therefore be interpreted as aggregate measures of predictive performance based on the complete feature sets, which may incorporate nonlinear and joint relationships, including possible interactions among the included predictors, rather than as direct estimates of any specific interaction effect.
- (4)
- SHAP analysis showed that the key driving factors of embodied carbon emissions in building engineering projects exhibit a clear stage-dependent transition, shifting from a “scale- and form-dominated” pattern at the schematic design stage to a “material- and construction-dominated” pattern at the construction drawing design stage. The former mainly depends on overall project scale, structural constraints, and building organization characteristics to identify carbon emission levels, whereas the latter more directly reflects the influence of material input and construction resource intensity on project-level carbon emissions. Notably, total machinery shifts still made a relatively strong contribution at the construction drawing design stage, indicating that construction-stage resource input intensity and organizational complexity are important factors affecting project-level embodied carbon emissions. The inclusion of this feature also represents an important extension of the present study beyond previous research.
- (5)
- The results further indicate that when the upstream accounting boundary is sufficiently comprehensive and the carbon emission factors are adequately covered, the downstream prediction models can not only achieve strong predictive accuracy, but also reveal patterns that are highly consistent with engineering knowledge and the SHAP-based interpretation results. Overall, the proposed framework contributes to sustainable building development by providing a rapid, stage-adapted, and interpretable approach for quantifying project-level embodied carbon emissions. It provides a unified decision-support basis for low-carbon scheme comparison at the schematic design stage, quantity verification and targeted emission reduction at the construction drawing design stage, and resource-efficient control during project implementation, thereby facilitating the integration of embodied carbon management into building design and project delivery.
5.2. Limitations and Future Work
Although this study systematically investigated project-level embodied carbon accounting, dataset construction, and stage-adapted prediction for building engineering projects, several limitations remain. First, the effective sample size of this study is limited at the project level. Although the dataset contains 426 individual buildings, these buildings are nested within 78 complete building engineering projects, and the project is the independent unit of analysis. This project-level setting was adopted to reflect actual engineering decision-making scenarios, in which carbon accounting and emission-reduction decisions are commonly implemented for complete projects comprising multiple buildings and ancillary works. Nevertheless, the use of 78 independent project samples results in a relatively small held-out test set. Consequently, the reported performance metrics may remain sensitive to sample composition and data partitioning, and the current results should be regarded as preliminary evidence of predictive feasibility rather than definitive estimates of model performance across the broader building stock.
Second, the nature and geographical composition of the dataset constrain the external validity of the findings. The samples were collected primarily from completed building engineering projects in China over the past five years and mainly comprise residential, public, and industrial buildings within the predefined size range. The project-level carbon labels were manually constructed from detailed engineering documents under a consistent A1–A5 boundary. The carbon emission factor system integrated Chinese national and industry standards with an extended database containing more than 2000 factors, which was systematically compiled, standardized, and curated by our research team. Although this procedure improves the completeness and internal consistency of the carbon labels, the resulting models are still closely linked to the Chinese construction, regulatory, technological, and carbon-accounting context. Their applicability to projects in other countries and regions, infrastructure projects, exceptionally large or special projects, and projects adopting substantially different structural systems, construction technologies, energy mixes, or carbon emission factor systems remains to be validated.
Future research should therefore focus on expanding both the number and diversity of independent project samples. Multi-regional and cross-country project datasets should be developed through standardized data-collection and carbon-accounting protocols, followed by independent external validation on projects not involved in model development. Repeated resampling, nested cross-validation, and uncertainty analysis could also be introduced to evaluate the stability of model performance more rigorously. In addition, future studies could investigate transfer learning, regional recalibration, and dynamic updating of carbon emission factors to improve the adaptability of the framework across different construction and accounting contexts.
Supplementary Materials
The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/su18157723/s1, Table S1: Main hyperparameters of machine learning models.
Author Contributions
Conceptualization, Z.W. and L.Z.; methodology, Z.W. and L.Z.; software, Z.W.; validation, L.Z. and M.P.; formal analysis, Z.W.; investigation, Z.W. and M.P.; resources, M.P.; data curation, Z.W. and M.P.; writing—original draft preparation, Z.W.; writing—review and editing, L.Z.; visualization, Z.W.; supervision, L.Z. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
The data presented in this study are available on request from the corresponding author.
Conflicts of Interest
Author Mengmeng Pu was employed by Zhejiang Senxin Low-Carbon Technology Co., Ltd. The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
References
- United Nations Environment Programme. Global Status Report for Buildings and Construction 2024/2025. Available online: https://www.unep.org/resources/report/global-status-report-buildings-and-construction-20242025 (accessed on 30 March 2026).
- Huang, Z.; Zhou, H.; Miao, Z.; Tang, H.; Lin, B.; Zhuang, W. Life-Cycle Carbon Emissions (LCCE) of Buildings: Implications, Calculations, and Reductions. Engineering 2024, 35, 115–139. [Google Scholar] [CrossRef] [Scilit]
- Building Energy Consumption and Carbon Emissions Data Committee of the China Association of Building Energy Efficiency. Research Report on Carbon Emissions in China’s Urban and Rural Development Sector (2024 Edition); China Association of Building Energy Efficiency: Chongqing, China, 2024; Available online: https://cmsfiles.zhongkefu.com.cn/cmsossjianzhujieneng/upload/file/20250120/1737310863233732.pdf (accessed on 31 March 2026).
- Building Energy Consumption and Carbon Emissions Data Committee of the China Association of Building Energy Efficiency. Research Report on Carbon Emissions in China’s Urban and Rural Development Sector (2025 Edition); China Association of Building Energy Efficiency: Beijing, China, 2026; Available online: https://cmsfiles.zhongkefu.com.cn/cmsossjianzhujieneng/upload/file/20260204/1770187228435856.pdf (accessed on 31 March 2026).
- Röck, M.; Saade, M.R.M.; Balouktsi, M.; Rasmussen, F.N.; Birgisdottir, H.; Frischknecht, R.; Habert, G.; Lützkendorf, T.; Passer, A. Embodied GHG Emissions of Buildings—The Hidden Challenge for Effective Climate Change Mitigation. Appl. Energy 2020, 258, 114107. [Google Scholar] [CrossRef] [Scilit]
- Dunant, C.F.; Drewniok, M.P.; Orr, J.J.; Allwood, J.M. Good Early Stage Design Decisions Can Halve Embodied CO2 and Lower Structural Frames’ Cost. Structures 2021, 33, 343–354. [Google Scholar] [CrossRef] [Scilit]
- Luo, L.; Cai, W. An Innovative Method for Rapid Estimation and Optimization of Building Embodied Carbon Emissions. Sustain. Cities Soc. 2025, 130, 106556. [Google Scholar] [CrossRef] [Scilit]
- Häkkinen, T.; Kuittinen, M.; Ruuska, A.; Jung, N. Reducing Embodied Carbon during the Design Process of Buildings. J. Build. Eng. 2015, 4, 1–13. [Google Scholar] [CrossRef] [Scilit]
- Victoria, M.F.; Perera, S. Parametric Embodied Carbon Prediction Model for Early Stage Estimating. Energy Build. 2018, 168, 106–119. [Google Scholar] [CrossRef] [Scilit]
- Abanda, F.H.; Tah, J.H.M.; Cheung, F.K.T. Mathematical Modelling of Embodied Energy, Greenhouse Gases, Waste, Time–Cost Parameters of Building Projects: A Review. Build. Environ. 2013, 59, 23–37. [Google Scholar] [CrossRef] [Scilit]
- Santos, A.K.; Ferreira, V.M.; Dias, A.C. Promoting Decarbonisation in the Construction of New Buildings: A Strategy to Calculate the Embodied Carbon Footprint. J. Build. Eng. 2025, 103, 112037. [Google Scholar] [CrossRef] [Scilit]
- Deng, X.; Lu, K. Multi-Level Assessment for Embodied Carbon of Buildings Using Multi-Source Industry Foundation Classes. J. Build. Eng. 2023, 72, 106705. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Chen, H.; Sun, J.; Zhang, X. Predictive Models of Embodied Carbon Emissions in Building Design Phases: Machine Learning Approaches Based on Residential Buildings in China. Build. Environ. 2024, 258, 111595. [Google Scholar] [CrossRef] [Scilit]
- Alwan, Z.; Ilhan Jones, B. IFC-Based Embodied Carbon Benchmarking for Early Design Analysis. Autom. Constr. 2022, 142, 104505. [Google Scholar] [CrossRef] [Scilit]
- D’Amico, B.; Arehart, J.H. Probabilistic Inference of Material Quantities and Embodied Carbon in Building Structures. J. Build. Eng. 2024, 94, 109891. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Guo, X.; Zhang, J.; Huang, R.; Zhang, X. Ensemble Learning-Based Prediction and Optimization of Embodied Carbon Emissions from Concrete Building Structures in Early Design Phases. J. Build. Eng. 2026, 119, 115343. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Wang, Y.; Zhao, L.; Wang, W.; Luo, Z.; Wang, Z.; Luo, J.; Lv, Y. Integrating BIM and Machine Learning to Predict Carbon Emissions under Foundation Materialization Stage: Case Study of China’s 35 Public Buildings. Front. Archit. Res. 2024, 13, 876–894. [Google Scholar] [CrossRef] [Scilit]
- Cang, Y.; Yang, L.; Luo, Z.; Zhang, N. Prediction of Embodied Carbon Emissions from Residential Buildings with Different Structural Forms. Sustain. Cities Soc. 2020, 54, 101946. [Google Scholar] [CrossRef] [Scilit]
- Zheng, L.; Mueller, M.; Luo, C.; Yan, X. Predicting Whole-Life Carbon Emissions for Buildings Using Different Machine Learning Algorithms: A Case Study on Typical Residential Properties in Cornwall, UK. Appl. Energy 2024, 357, 122472. [Google Scholar] [CrossRef] [Scilit]
- Luo, L.; Chen, Y.; Feng, W.; Wu, J.; Cai, W. Progressive Prediction of Embodied Carbon Emissions across Stages of Schematic Design with Machine Learning. J. Clean. Prod. 2025, 517, 145799. [Google Scholar] [CrossRef] [Scilit]
- Shi, Y.; Cao, X.; Yang, X. Assessment and Reduction of Embodied Carbon Emissions in Buildings: A Systematic Literature Review of Recent Advances. Energy Build. 2025, 345, 116058. [Google Scholar] [CrossRef] [Scilit]
- Huang, B.; Zhang, H.; Ullah, H.; Lv, Y. BIM-Based Embodied Carbon Evaluation during Building Early-Design Stage: A Systematic Literature Review. Environ. Impact Assess. Rev. 2025, 112, 107768. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Xie, W.; Xu, L.; Li, L.; Jim, C.Y.; Wei, T. Holistic Life-Cycle Accounting of Carbon Emissions of Prefabricated Buildings Using LCA and BIM. Energy Build. 2022, 266, 112136. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Li, Y.; Chen, H.; Yan, X.; Liu, K. Characteristics of Embodied Carbon Emissions for High-Rise Building Construction: A Statistical Study on 403 Residential Buildings in China. Resour. Conserv. Recycl. 2023, 198, 107200. [Google Scholar] [CrossRef] [Scilit]
- Li, D.; Cui, P.; Lu, Y. Development of an Automated Estimator of Life-Cycle Carbon Emissions for Residential Buildings: A Case Study in Nanjing, China. Habitat Int. 2016, 57, 154–163. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Zheng, R. Reducing Building Embodied Emissions in the Design Phase: A Comparative Study on Structural Alternatives. J. Clean. Prod. 2020, 243, 118656. [Google Scholar] [CrossRef] [Scilit]
- Luo, Z.; Yang, L.; Liu, J. Embodied Carbon Emissions of Office Building: A Case Study of China’s 78 Office Buildings. Build. Environ. 2016, 95, 365–371. [Google Scholar] [CrossRef] [Scilit]
- Zheng, L.; Mueller, M.; Luo, C.; Menneer, T.; Yan, X. Variations in Whole-Life Carbon Emissions of Similar Buildings in Proximity: An Analysis of 145 Residential Properties in Cornwall, UK. Energy Build. 2023, 296, 113387. [Google Scholar] [CrossRef] [Scilit]
- Neupane, B.; Belkadi, F.; Formentini, M.; Rozière, E.; Hilloulin, B.; Abdolmaleki, S.F.; Mensah, M. Machine Learning Algorithms for Supporting Life Cycle Assessment Studies: An Analytical Review. Sustain. Prod. Consum. 2025, 56, 37–53. [Google Scholar] [CrossRef] [Scilit]
- Fang, Y.; Lu, X.; Li, H. A Random Forest-Based Model for the Prediction of Construction-Stage Carbon Emissions at the Early Design Stage. J. Clean. Prod. 2021, 328, 129657. [Google Scholar] [CrossRef] [Scilit]
- Xikai, M.; Lixiong, W.; Jiwei, L.; Xiaoli, Q.; Tongyao, W. Comparison of Regression Models for Estimation of Carbon Emissions during Building’s Lifecycle Using Designing Factors: A Case Study of Residential Buildings in Tianjin, China. Energy Build. 2019, 204, 109519. [Google Scholar] [CrossRef] [Scilit]
- Su, S.; Zang, Z.; Yuan, J.; Pan, X.; Shan, M. Considering Critical Building Materials for Embodied Carbon Emissions in Buildings: A Machine Learning-Based Prediction Model and Tool. Case Stud. Constr. Mater. 2024, 20, e02887. [Google Scholar] [CrossRef] [Scilit]
- Zhou, C.; Wang, Z.; Wang, X.; Guo, R.; Zhang, Z.; Xiang, X.; Wu, Y. Deciphering the Nonlinear and Synergistic Role of Building Energy Variables in Shaping Carbon Emissions: A LightGBM- SHAP Framework in Office Buildings. Build. Environ. 2024, 266, 112035. [Google Scholar] [CrossRef] [Scilit]
- Xiong, X.; Li, X.; Chen, S.; Chen, D.; Lin, J. Review and Prediction: Carbon Emissions from the Materialization of Residential Buildings in China. Sustain. Cities Soc. 2025, 121, 106211. [Google Scholar] [CrossRef] [Scilit]
- Cang, Y.; Luo, Z.; Yang, L.; Wang, W.; Si, Q.; Zhang, J.; Tong, Y. Prediction Method of Building Embodied Carbon Emissions for Conceptual Design Stage: Based on the Overall Design Parameters by Machine Learning. Build. Environ. 2026, 289, 114053. [Google Scholar] [CrossRef] [Scilit]
- Fenton, S.K.; Munteanu, A.; De Rycke, K.; De Laet, L. Embodied Greenhouse Gas Emissions of Buildings—Machine Learning Approach for Early Stage Prediction. Build. Environ. 2024, 257, 111523. [Google Scholar] [CrossRef] [Scilit]
- Marsh, E.; Orr, J.; Ibell, T. Quantification of Uncertainty in Product Stage Embodied Carbon Calculations for Buildings. Energy Build. 2021, 251, 111340. [Google Scholar] [CrossRef] [Scilit]
- Ding, Z.; Liu, S.; Luo, L.; Liao, L. A Building Information Modeling-Based Carbon Emission Measurement System for Prefabricated Residential Buildings during the Materialization Phase. J. Clean. Prod. 2020, 264, 121728. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Zhang, X. A Subproject-Based Quota Approach for Life Cycle Carbon Assessment at the Building Design and Construction Stage in China. Build. Environ. 2020, 185, 107258. [Google Scholar] [CrossRef] [Scilit]
- GB/T 51366–2019; Standard for Building Carbon Emission Calculation. Ministry of Housing and Urban-Rural Development of the People’s Republic of China: Beijing, China; China Architecture & Building Press: Beijing, China, 2019. (In Chinese)
- T/CCIAT 0104–2025; Code for Carbon Reduction Construction of Building Engineering. China Construction Industry Association: Beijing, China; China Construction Industry Association: Beijing, China, 2025. (In Chinese)
- Teng, Y.; Pan, W. Estimating and Minimizing Embodied Carbon of Prefabricated High-Rise Residential Buildings Considering Parameter, Scenario and Model Uncertainties. Build. Environ. 2020, 180, 106951. [Google Scholar] [CrossRef] [Scilit]
- Zhan, Z.; Xia, P.; Xia, D. Study on Carbon Emission Measurement and Influencing Factors for Prefabricated Buildings at the Materialization Stage Based on LCA. Sustainability 2023, 15, 13648. [Google Scholar] [CrossRef] [Scilit]
- Cover, T.M.; Hart, P.E. Nearest Neighbor Pattern Classification. IEEE Trans. Inf. Theory 1967, 13, 21–27. [Google Scholar] [CrossRef] [Scilit]
- Smola, A.J.; Schölkopf, B. A Tutorial on Support Vector Regression. Stat. Comput. 2004, 14, 199–222. [Google Scholar] [CrossRef] [Scilit]
- Breiman, L. Random Forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef] [Scilit]
- Geurts, P.; Ernst, D.; Wehenkel, L. Extremely Randomized Trees. Mach. Learn 2006, 63, 3–42. [Google Scholar] [CrossRef] [Scilit]
- Friedman, J.H. Greedy Function Approximation: A Gradient Boosting Machine. Ann. Stat. 2001, 29, 1189–1232. [Google Scholar] [CrossRef] [Scilit]
- Chen, T.; Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; Association for Computing Machinery (ACM): New York, NY, USA, 2016; pp. 785–794. [Google Scholar] [CrossRef] [Scilit]
- Ke, G.; Meng, Q.; Finley, T.; Wang, T.; Chen, W.; Ma, W.; Ye, Q.; Liu, T.-Y. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
- Prokhorenkova, L.; Gusev, G.; Vorobev, A.; Dorogush, A.V.; Gulin, A. CatBoost: Unbiased Boosting with Categorical Features. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2018; Volume 31. [Google Scholar]
- Wu, Y.; Wang, S.; Gao, C.; Sun, R.; Xia, C. Residual-Corrected Stacking Ensemble Learning for Concrete Strength Prediction and Optimization of Low-Carbon Mix Design. AIP Adv. 2025, 15, 115237. [Google Scholar] [CrossRef] [Scilit]
- Mu, W.; Cardelli, R.; Ferrari, S. Data Preprocessing Techniques for Machine Learning Towards Improving Building Energy Performance: A Systematic Review. Energies 2026, 19, 1561. [Google Scholar] [CrossRef] [Scilit]
- Anastasiadou, M.; dos Santos, V.D.; Dias, M.S. An Explainable Deep Learning Model for Energy Performance Classification and Retrofitting Recommendations. Energy Build. 2025, 349, 116522. [Google Scholar] [CrossRef] [Scilit]
- Baset, A.; Jradi, M. Prioritizing Building Envelope Retrofits through Data-Driven Heat Flux Prediction. J. Build. Eng. 2026, 119, 115130. [Google Scholar] [CrossRef] [Scilit]
- Zhang, C.; Yan, R.; Liu, X.; Nie, X.; Peng, Y.; Peng, X. Empirical Simulation of Internal Validation Methods for Prediction Models: Comparing k-Fold Cross-Validation with Bootstrap-Based Optimism Correction. J. Clin. Epidemiol. 2026, 190, 112101. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wong, T.-T. Performance Evaluation of Classification Algorithms by k-Fold and Leave-One-out Cross Validation. Pattern Recognit. 2015, 48, 2839–2846. [Google Scholar] [CrossRef] [Scilit]
- Wu, J.; Chen, X.-Y.; Zhang, H.; Xiong, L.-D.; Lei, H.; Deng, S.-H. Hyperparameter Optimization for Machine Learning Models Based on Bayesian Optimization. J. Electron. Sci. Technol. 2019, 17, 26–40. [Google Scholar] [CrossRef]
- Lundberg, S.M.; Lee, S.-I. A Unified Approach to Interpreting Model Predictions. In Proceedings of the Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
- Chen, D.; Chen, S.; Xiong, X.; Li, X. Embodied Carbon Emissions from the Materialization of Office Buildings in China: A Systematic Review of Range, Characteristics and Impact Factors. J. Build. Eng. 2025, 113, 114065. [Google Scholar] [CrossRef] [Scilit]
- Fang, D.; Brown, N.; De Wolf, C.; Mueller, C. Reducing Embodied Carbon in Structural Systems: A Review of Early-Stage Design Strategies. J. Build. Eng. 2023, 76, 107054. [Google Scholar] [CrossRef] [Scilit]
- Zhao, Z.; Pan, C.; Yang, G.; Ji, C.; Massucato, C.J.; Attia, S. Establishing Industry-Based Benchmarks for Embodied Carbon in Buildings: A Statistical Analysis of 85 Case Studies in China. J. Build. Eng. 2025, 114, 114372. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.









