Next Article in Journal
Small Models, Big Cities: A Low-Cost AI Pipeline for Urban Regulatory Document Analysis in Metropolitan Planning
Previous Article in Journal
Energy- and Communication-Aware Federated Learning for Smart City Sensing and Urban Intelligence
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Environmental-Health Vulnerability and Respiratory Mortality in Europe: Evidence from Panel Econometrics, Clustering, and Machine Learning

1
Department of Clinical and Molecular Medicine, Sapienza University of Rome, 00185 Rome, Italy
2
Institute of Respiratory Disease, Department of Basic Medical Science, Neuroscience, and Sense Organs, University of Bari “Aldo Moro”, Piazza Giulio Cesare 11, 70125 Bari, Italy
3
IRCCS Don Gnocchi, 50143 Firenze, Italy
4
Dipartimento di Management, Finanza e Tecnologia, LUM University Giuseppe Degennaro, 70010 Casamassima, Italy
*
Author to whom correspondence should be addressed.
Urban Sci. 2026, 10(7), 351; https://doi.org/10.3390/urbansci10070351
Submission received: 2 April 2026 / Revised: 8 June 2026 / Accepted: 15 June 2026 / Published: 24 June 2026

Abstract

Respiratory mortality in Europe is associated with interacting environmental, infrastructural, climatic, and energy-related conditions. This study investigates country–year patterns of respiratory disease mortality by integrating panel-data econometrics, clustering analysis, and machine-learning prediction. The econometric results indicate that agricultural land use and coal-based electricity generation are positively associated with respiratory mortality, while access to electricity and freshwater withdrawals show negative associations. Cooling degree days capture a heat-related environmental-health dimension, although some coefficients become weaker under robust specifications. Sanitation and renewable energy display heterogeneous and specification-sensitive patterns, suggesting that they may partly reflect broader development gradients, infrastructure transitions, and regional heterogeneity rather than direct causal mechanisms. Hierarchical clustering identifies 10 country–year environmental-health profiles, highlighting differentiated combinations of energy systems, land use, infrastructure, climatic exposure, and respiratory mortality. This approach avoids treating countries as fixed homogeneous units and allows environmental-health profiles to vary over time. The selected hierarchical solution provides a balanced and interpretable structure relative to more polarized clustering alternatives. Machine-learning models are used as a complementary predictive exercise rather than as substitutes for econometric inference. Within the adopted validation framework, K-nearest neighbors achieves the strongest predictive performance. Additional stability checks and local additive explanations improve transparency regarding model tuning and prediction behavior, while confirming that machine-learning outputs should be interpreted as predictive rather than causal evidence. Overall, the findings support integrated and region-sensitive policy approaches combining air-quality management, infrastructure resilience, energy transition, climate adaptation, and public-health planning.

1. Introduction

One of the key features of the environmental health situation in Europe, in relation to current socioeconomic development trends, is respiratory mortality. Respiratory mortality in Europe arises from the interaction among infrastructural, energy-related, land-use, climate-related, and territorial factors [1,2,3]. While the current research is based on country–year-level data, its results are highly relevant to urban science researchers, as they pertain to environmental health in metropolitan, mid-sized cities, and peripheral regions [4,5]. Environmental health in metropolises depends not only on local circumstances, but also on regional and national-level characteristics, including infrastructure, energy, climate adaptation, and environmental management policies [1,2]. Thus, respiratory mortality should be considered a key marker of vulnerability to metropolitan and regional environmental health risks [5]. In this context, indicators such as electricity access, sanitation access, freshwater withdrawals, cooling degree days, electricity production through coal burning, agriculture land use, and share of renewable energy are important due to the fact that they are linked to the resilience of the infrastructural system, environmental stress, climatic stress, and the process of transition towards sustainability [1,4,6]. Cooling degree days can be seen as a measure of climatic stress specific to large cities, while a high reliance on coal-fired electricity production and agriculture indicates environmental stress [3,6,7]. Furthermore, electricity access, sanitation access, freshwater withdrawal, and renewable energy can be taken as proxies for development and regional inequality [1,4]. Analytical approaches used in order to discover determinants include regression modeling, sensitivity analyses, hierarchical clustering, and machine learning. Regression models will allow for identifying the most prominent predictors, while heterogeneity tests, variance inflation factor analyses, and the Driscoll–Kraay correction will increase the robustness of the results [5]. On the other hand, hierarchical clustering will allow for defining different profiles of environmental health risk and illustrating how environmental health risks transform due to infrastructural, energy-related, climatic, and environmental dynamics [5]. Machine learning will complement regression analysis in making predictions by accounting for training, testing, scaling, tuning, stability, and interpretability [5]. In summary, the present research contributes to the field of urban science by investigating the relationship between infrastructure transitions and regional environmental health risks. Policy recommendations will follow from the research findings and will relate to air pollution control, building resilient infrastructures, transitioning to renewable energy sources, climate change adaptation, and reducing environmental health risks overall [3,4,8,9].
The remainder of the paper is organized as follows. Section 2 presents the analytical framework and variable definitions. Section 3 describes the dataset, data harmonization procedures, and the multi-method research design. Section 4 reports the econometric results on environmental, infrastructural, and climatic associations with respiratory mortality. Section 5 identifies environmental-health vulnerability profiles through hierarchical clustering. Section 6 presents the machine-learning analysis and predictive modelling results. Section 7 integrates the econometric, clustering, and machine-learning evidence. Section 8 discusses the study’s limitations. Section 9 examines policy implications, while Section 10 concludes.

2. Integrated Analytical Framework and Variable Definitions

In this section, the relationships between environmental conditions, infrastructure, climate, energy consumption, and respiratory tract disease are analyzed using a quantitative approach. Respiratory tract disease is defined as the age-standardized mortality rate from respiratory diseases per 100,000 people. It is important to note that respiratory diseases have been linked to numerous environmental factors, including air pollution and industrialization [10,11,12]. ELEC shows the portion of electricity services among the studied population. AGRL shows the agricultural land portion in the country. The WTRW is the annual portion of water withdrawals compared to total renewable internal freshwater. CDDs are the annual cooling degree days, starting at 18° Celsius. To assess how coal pollution influences respiratory disease mortality, the proportion of coal used in electricity generation (COAL) was selected, as there is a correlation between the two variables. SANS may be understood as the proportion of sanitary services per capita. Lastly, RENE refers to the share of renewables in energy consumption. There are many cases in which respiratory diseases occur due to international differences in mortality rates. Also, air pollution in Europe has recently been identified as one of the main factors influencing pneumonia and should thus be accounted for in the analysis [11]. See Table 1.
The dataset was harmonized at the country–year level. All variables were retained in their original measurement units for panel regressions. No logarithmic transformation was applied. Feature scaling was applied only for clustering and machine-learning procedures.

3. Data Harmonization and Multi-Method Analytical Design

The empirical analysis is based on panel data consisting of 41 countries observed from 2010 to 2021. This choice of panel corresponds to the availability of comparable indicators for respiratory disease mortality rates, environmental factors, infrastructure, climate, energy use, and other related indicators. Indeed, it has already been shown that mortality from respiratory diseases can be considered a variable influenced by environmental factors. There are established links between industrial emissions, air quality, and the development of respiratory diseases [10,11,12]. The geographical region in this case covers all the countries included in the databases used as sources of data for the analysis, rather than the common definition of Europe. Thus, 41 countries were initially selected for the analysis, and after listwise deletion of incomplete observations, the panel consisted of 38 country–year observations. This means that the initial dataset included 41 countries, but the estimation sample included 38 cross-sections or country–year observations. The dataset used is an unbalanced panel because the observations within country–year pairs were not evenly distributed, depending on the data availability in international databases. The missing observations mostly referred to sanitation, renewable energy use, and cooling degree days. They are mostly concerned with Eastern and South-Eastern Europe. This means that one observation represents the condition of one country in a given year, not the country itself. The selection of countries was based on their TRD mortality data, and the indicators included in the regression model as regressors. The reasons for including Israel and Russia in the dataset are because of the inclusion of those countries in international comparative studies and the availability of data in the databases used. TRD mortality was used as the dependent variable, while the list of additional predictors included the share of population having access to electricity (ELEC), the share of agricultural land (AGRL), the amount of freshwater withdrawn per capita (WTRW), cooling degree days (CDDs), electricity generated by coal (COAL), the share of the population with safe sanitation access (SANS), and the share of renewable energy usage in the country (RENE). The inclusion of energy-related and environment-related variables can be justified based on prior evidence about the negative impact of air pollution, emissions caused by the usage of fossil fuels, and other factors [10,11,12,13,14]. Additionally, climatic stress is important, as there are established links between extreme temperatures and respiratory mortality in many countries [15]. Renewable energy-related indicators were used to describe particular characteristics and changes in the energy system. The missing-data handling method chosen for this study is listwise deletion. Every model was estimated with the maximum possible number of non-missing observations. Listwise deletion was preferred to imputation when working with environmental macro-level indicators, since artificial fluctuations in those indicators across the entire panel are unwanted. Fixed-effects and random-effects panel regressions were estimated. The variables in the dataset have been harmonized to the country–year level. Variables have not been transformed prior to estimation. See Figure 1.
Countries are represented in Figure 2.
The correlation matrix among the regressors used in the panel regressions on respiratory disease mortality (TRD) is shown in Figure 3 below. It provides an insight into the correlations among the environment, infrastructure, climate, and energy-related regressors across European countries. In general, the results show no sign of multicollinearity, supporting the credibility of the empirical results found in Table 2. Pairwise correlations have been used following prior literature on the interdependencies among climatic, water, agricultural, and electricity regressors in environmental and energy economics [16,17,18]. To begin with, the highest positive correlation coefficient has been observed between WTRW and CDDs (r = 0.61), suggesting that high temperatures are associated with higher water consumption. Research supports a positive relationship between cooling and water use depending on climate conditions [16,17]. Secondly, the moderately positive relationships have been found between agricultural land use and TRD (r = 0.32) and between agricultural land use and COAL (r = 0.31). This implies that structural vulnerability is associated with significant agricultural activity and coal use. Finally, RENE shows negative correlations with AGRL (r = −0.47), WTRW (r = −0.38), and CDDs (r = −0.25), indicating that lower resource consumption is associated with greater reliance on renewables. The results align with the existing literature on environmental and energy economics regarding energy consumption and demand in Europe and the world at large [16,18]. Lastly, the lack of strong correlations between ELEC and SANS and the other regressors in the regression models suggests a low level of shared explanatory power with the other regressors, indicating that positive relationships with health are unlikely to be explained by multicollinearity. See Figure 3.
The diagnostics for multicollinearity in Table 2 can be used to evaluate whether multicollinearity affects the regression models used in this research. Indeed, the application of VIF and tolerance as measures of multicollinearity, and their possible effects on regression coefficient distortion, is justified by their use as diagnostics [19,20,21]. Generally speaking, the results indicate that multicollinearity is not a concern in the analysis. The average VIF is relatively low and well below the threshold, at 1.57. Among the variables with the highest values, one may find WTRW (VIF = 2.01), followed by CDDs (VIF = 1.92) and RENE (VIF = 1.69), indicating that these regressors exhibit relatively low correlations with other variables. Moreover, only one additional VIF is slightly above 1.5, at 1.54, indicating some multicollinearity, but it is still tolerable for analysis. Variables with VIF under 1.5 include COAL (VIF = 1.38) and ELEC (VIF = 1.08). See Table 2.
Figure 4 is an infographic illustrating the most important findings of the present research and how various factors, such as the environment, infrastructure, climate, and energy production, impact the respiratory health of the population in Europe.
The key concept illustrated by the infographic is conveyed through the metaphor of a balance scale, which highlights the connection between risk factors on one hand and factors contributing to resilience on the other. Specifically, on the left-hand side, the risks associated with respiratory deaths are shown. In this respect, agricultural land use and electricity production from coal are depicted, as they have been identified as risk factors based on their positive coefficients in the econometric analysis. These two factors indicate increased levels of pollutants, PM, and environmental stressors, which may contribute to poor respiratory health [10,11]. Furthermore, heat stress is highlighted since CDDs have been found to significantly predict respiratory mortality. This risk increases as the weather becomes hotter and as heatwaves become more common in densely populated regions where heat accumulates [22,23]. On the right-hand side of the infographic, some examples of protective factors are shown. Electricity access and access to freshwater sources are presented as protective factors, as they indicate better living conditions and infrastructural resilience. In other words, they indicate better environmental and health security in Europe, which makes people less susceptible to respiratory conditions through infrastructure resilience and the proper functioning of the relevant infrastructure system [22,24]. At the bottom of the infographic, you can find information on the two most important additional findings. Indeed, as explained before, the results show the presence of 10 environmental-health profiles and high predictive performance using KNN. Together, these findings suggest that respiratory mortality does not derive from one factor only but results from the interaction between climate stress, infrastructural resilience, energy production systems, and environmental conditions [10,11,23,24]. Therefore, effective policies need to consider multiple determinants of such mortality.

4. Environmental, Infrastructural, and Climatic Associations with Respiratory Mortality

This subsection describes the panel analysis on respiratory disease mortality (TRD) based on the panel dataset with 38 effective cross-sectional observations following exclusion of data availability and missing values from the first group of 41 countries. This study uses both fixed effects and random effects model estimations to assess the relationship between TRD and its potential determinants, such as electricity access, agricultural land use, freshwater withdrawals, cooling degree days, coal-based electricity generation, sanitation access, and renewable energy consumption. Specifically, panel data estimations are used here to control for the presence of unobservable heterogeneities and time-specific effects, which is in line with established practice in country-wide analysis on the environment and energy impacts on health [11,24]. The model specification was tested through Breusch–Pagan, Hausman, and joint significance tests. Results showed negative relationship between TRD and electricity access and freshwater withdrawals, while a positive relationship exists between TRD and agricultural land use, coal-based electricity generation, and heat stress. Those results corresponded to recent studies suggesting negative impacts of air pollution and fossil fuel dependency on respiratory mortality along with other adverse public health outcomes due to climate change [11,22]. On the other hand, unexpected positive relationships between sanitation access and renewable energy consumption and TRD were reported; however, considering previous results on complex interplay between renewables and social factors affecting health, these positive relationships cannot be interpreted causally [24]. Correlation and VIF analyses indicated that severe multicollinearity was not a main cause of positive results. However, the role of possible omitted variable bias, reverse causality, development gradients, and regional variability cannot be neglected. Therefore, those results cannot be considered causally but only as a sign of the existence of relationships among the variables in heterogeneous European contexts [22,24].
T R D = α + β 1 E L E C i t + β 2 A G R L i t + β 3 W T R W i t + β 4 C D D i t + β 5 C O A L i t + β 6 S A N S i t + β 7 R E N E i t
where i denotes the cross-sectional country units included in the estimations (effective N = 38) and t denotes the annual time dimension over the period 2010–2021 (Table 3).
The TRD index combines information about energy share, agriculture, freshwater extraction, CDDs, sanitation, and renewables. In this case, it is appropriate to apply fixed- and random-effect regression models to address heterogeneity, as in other panel data research on health care and the environment [11,24]. Substantial heterogeneity was observed in the sample regarding respiratory mortality (mean = 38.9; SD = 16.8). A negative effect was observed for energy consumption (−0.8129); however, positive influences were observed for agricultural use (0.5151), energy from coal (0.1149), and heat (0.00449). The results align with other empirical studies, revealing a link between respiratory deaths and various environmental factors [11,22]. Moreover, the impact of freshwater (−0.1084) was negative, while sanitation (0.2605) and renewable energy (0.2497) had positive effects. Renewable energy deployment appears to affect public health differently across countries [24]. Fixed-effect regressions provided similar results. The diagnostics revealed an adequate functional form with serial correlation that requires the use of robust estimators. Thus, there are significant links between respiratory mortality rates and environmental, health, and climate factors.

4.1. Robustness Analysis Using Driscoll–Kraay Standard Errors: Addressing Cross-Sectional Dependence and Temporal Correlation

Because robust estimates are needed for both fixed and random effects, the Driscoll–Kraay robust standard errors were used. Heteroskedasticity, autocorrelation, and cross-sectional dependence are problems common to macro-panel data across several countries. Clustered standard errors can underestimate variance in such cases, leading to misleading inference. Driscoll–Kraay adjustments provide heteroskedasticity-consistent estimates that are valid even in the presence of both time and spatial correlations. Thus, this methodology allows for drawing more conservative and reliable conclusions from the statistical analysis. It seems quite appropriate for estimating relationships between variables that are components of an interlinked environment and the macroeconomy, such as those in Europe. Cross-sectional dependence in such panels arises from transboundary environmental processes, the integration of national energy markets, and similar responses to common policies across countries. Given the small number of periods in the dataset, the two-lag specification was chosen, as otherwise, there is a risk of overestimating temporal dependence and ending up with over-parameterization (Table 4).
The Driscoll–Kraay estimator applied to two lags in the fixed-effect regression accounts for both spatial and temporal dependence, yielding relatively conservative inference with N = 238 across 38 European countries. It is common practice to apply this methodology in panel regressions because of potential heteroscedasticity, autocorrelation, and cross-sectional dependence [25,26,27]. The model seems to fit well, as seen in the within R2 of 0.3909. The reported F-statistic is extremely high (F(17,10) = 62,298.16, p < 0.001); however, this should be interpreted carefully. In particular, it must be understood that it is the automated computation of the robust Wald-type test statistic, which uses the covariance matrix calculated using the Driscoll–Kraay estimator, rather than the F-statistic proper of the fixed-effect regression [25,28]. Large values are a consequence of robustness and appropriate degrees of freedom. According to the Driscoll–Kraay findings, only some of the baseline evidence is robust to the presence of cross-sectional and temporal dependencies [26,27]. For example, there is still a significant (p = 0.004) positive impact of agricultural land use (AGRL), thus confirming the association between environment-related factors in agriculture and respiratory health. Moreover, coal-based electricity generation (COAL) still has a significant effect (p = 0.033), thereby supporting the notion that energy production systems based on fossil fuels contribute to respiratory deaths. Finally, safe sanitation (SANS) remains significant too (p < 0.001); however, it is important to note that its influence cannot be considered directly negative due to other developmental and transitional determinants. Some [25,26,27] coefficients exhibit lower robustness when the Driscoll–Kraay estimator is used. Thus, electricity (ELEC) has become negatively significant (p = 0.056) in terms of its sign while remaining the same. Water withdrawals (WTRW) remain significant at the marginal level (p = 0.082). Cooling degree days (CDDs) become non-significant (p = 0.143), and the consumption of renewable energy (RENE) ceases to be statistically significant (p = 0.259). Hence, the Driscoll–Kraay estimator calls for greater caution regarding empirical evidence, and only the AGRL, COAL, and SANS variables remain highly significant. Thus, the discussion of the findings distinguishes between robust and unreliable estimates (i.e., those with non-significant p-values or lower p-values).

4.2. Sensitivity Analyses and Robustness Checks for Alternative Specifications

Table 5 presents sensitivity analyses to determine whether the positive coefficients for sanitation access (SANS) and renewable energy consumption (RENE) depend on the base-case specification with fixed effects.
The results show that SANS and RENE are positively related to respiratory deaths and statistically significant, with coefficients of 0.269 and 0.235, respectively. Removing SANS from the baseline model does not affect the results, as the RENE coefficient increases to 0.360 and remains positive and statistically significant. Hence, one can conclude that the positive effect of renewable energy consumption is not driven by its simultaneous presence with SANS in the model. The same applies when removing RENE from the baseline model, because the SANS coefficient becomes 0.330, indicating that the positive relation with SANS is not driven by the presence of RENE. Thus, after removing both variables from the model, the core environmental controls demonstrate relatively stable estimates, indicating the robustness of the base-case model setup. Such an analysis is crucial for determining the stability of estimated relations across various model specifications and preventing misinterpretation of observational panel estimates [29]. A lagged-regressor specification provides some insights into reverse causality and the time dependence of the coefficients. Here, the positive, statistically significant effect is associated with SANS, while RENE remains positive but statistically insignificant. It seems that sanitation access shows a more stable time effect than renewable energy. Nevertheless, the sensitivity analyses suggest that any conclusions should be drawn cautiously regarding the positive effects of SANS and RENE. In particular, the positive coefficient estimates should not be taken as evidence of a harmful impact of these variables. Instead, they are more likely related to development gradients, omitted-variable biases, transitional mechanisms, or geographic differences between European countries. Such an approach is consistent with the evidence on the complex relationships among clean energy, social indicators, infrastructure, and health outcomes [24,30].

4.3. Regional Heterogeneity Between Western and Eastern Europe

Table 6 displays regional fixed-effect estimates based on a split sample of Western and Eastern European countries. See Table 6.
There are clear differences in the explanatory power of the independent variables across the two European regions. Namely, in Western Europe, the AGRL, CDD, COAL, and RENE factors have positive coefficients that are also significant at the 1% level. At the same time, the SANS factor shows no statistically significant relationship. It is likely that RENE may capture certain structural and transitional processes related to energy transitions in highly developed countries with unique approaches towards the implementation of renewable energy sources and environmental restructuring [31]. On the other hand, Eastern Europe demonstrates that SANS plays an important positive role and is highly significant, whereas RENE does not show a statistically significant impact. Also, the EAC and FWTH variables become statistically significant, indicating that the impact of infrastructure and resource availability on respiratory mortality rates may be more pronounced in less economically developed regions. This statement is consistent with the idea that environmental conditions, infrastructure, public health factors, and economic development continue to play a key role in respiratory mortality in Eastern European countries [32]. Meanwhile, the energy transition in Central and Eastern Europe proceeds through different stages of institution building and economic development compared with Western Europe [33]. At the same time, the CDD variable continues to show a positive, significant coefficient across all European regions. In this respect, it can be stated that the effect of climate pressure on mortality rates is relatively consistent across all European countries. Overall, these findings imply that baseline coefficients may differ across European countries and may partially reflect differences in development levels and institutions [31,32,33].

4.4. Extended Fixed-Effects Specification Addressing Potential Omitted-Variable Bias

Table 7 contrasts the baseline fixed-effects regression model specification with one that incorporates additional demographic, healthcare, socioeconomic, and pollution-control variables.
These results provide a better understanding of robust and conservative predictors of respiratory disease mortality. The inclusion of additional demographic and environmental variables significantly increases the model’s explanatory power, as evidenced by the improvement in within R2 from 0.3909 to 0.5053, while the number of observations decreases from 238 to 237. Therefore, it can be assumed that demographic, healthcare, economic development, and pollutant controls provide relevant explanatory information but do not reduce the sample. The inclusion of demographic, environmental, and health control variables is consistent with recent research showing that aging, healthcare state, environmental pollution, and economic development are critical determinants of mortality and public health issues among EU countries [11,22,32]. As for environmental and infrastructural variables, the results remain highly robust, especially for the most significant environmental risk factors. Firstly, ELEC remains highly negative and significantly correlated with increased respiratory mortality, with a slight reduction in the coefficient from −1.276 to −1.196 (at 1%). Therefore, it can be concluded that electricity is a protective infrastructure variable, regardless of whether additional demographic, healthcare, and pollution factors were added to the baseline specification. Access to modern energy systems has previously been shown to be correlated with health outcomes in many cross-sectional and panel studies [30].
Additionally, AGRL shows a strong, positive correlation with respiratory mortality in two models, with a small increase in the coefficient from 0.636 to 0.648. Furthermore, WTRW shows negative, significant coefficients in both models, although the correlation strength increases slightly in the second case. It can be concluded that the protective effect is not due to any omitted variables in the second model. Recent studies on climate-related exposures and the environment have also demonstrated their crucial impact on health outcomes [22].
There was a slight change in CDDs and COAL. While both showed a highly positive relationship with respiratory mortality in the baseline regression model, in the second case, their significance dropped to 10%, down from 5% initially. This means that some effects from these variables are likely due to other, omitted factors in the baseline regression. However, the role of the environment in people’s health has recently been highlighted by many scholars working in the field of European public health [22].
It is noteworthy that the behavior of RENE and SANS changed. After adding the omitted controls to the baseline model, SANS and RENE were no longer positively or negatively correlated, respectively, and SANS lost significance, while RENE became almost zero. These findings directly relate to omitted-variable bias because they show that the initial positive coefficient for SANS was partially driven by the inclusion of variables such as aging, income, and healthcare.
POP65 shows a significant and positive relationship with respiratory mortality. GDPPC is negative, meaning that higher income levels are associated with lower respiratory mortality. The relationship between healthcare and respiratory mortality is very weak. As for PM25, it does not show any significance in this specification. Nevertheless, a large European study showed a high correlation between respiratory mortality and air pollution [11].
The graph in Figure 5 compares the coefficients obtained via estimation in the base fixed-effects model with those from a model that includes several additional control variables (population aging (POP65), health development (HEALTH), GDP per capita (GDPPC), and PM2.5 (pollution). See Figure 5.
Thus, the visualization helps assess the robustness of the findings and the potential influence of omitted-variable bias [29]. Several coefficients remain stable across the two models. In particular, access to electricity (ELEC) remains negatively associated with respiratory mortality, agricultural land use (AGRL) remains positively associated, and freshwater withdrawals (WTRW) retain a negative coefficient. These results suggest that the main findings for these variables are robust to the inclusion of additional controls. By contrast, the coefficients for sanitation access (SANS) and renewable energy consumption (RENE) move closer to zero in the extended model, suggesting that omitted demographic and socioeconomic factors may have influenced their baseline estimates. Among the added controls, POP65 shows the largest positive coefficient, while the health development indicator is negatively associated with mortality. GDP per capita also contributes to the model, whereas PM2.5 is not statistically significant in this specification. This insignificant PM2.5 result should be interpreted cautiously, given that previous European studies have documented adverse respiratory-health effects of particulate matter [11].

4.5. Interpreting Structural Environmental and Infrastructure Indicators

Interpretation of the country-level environmental and infrastructural indicators merits great attention due to the nature of these predictors—they are structural variables that can capture broader conditions than mere causal mechanisms. Thus, even if the regression models identify statistically significant links between some predictors and respiratory disease mortality, it is essential to remember that these relations are based on aggregate national data and, hence, coefficients are better understood in the context of the structural characteristics of environmental and infrastructural conditions in European countries [22,29]. Specifically, in terms of safe sanitation access (SANS), a positive coefficient estimated in the base model cannot be taken as evidence that improving sanitation will increase mortality rates from respiratory diseases. In contrast, the findings suggest some development-specific effects, such as infrastructure transitions or regional differences. The robustness of this interpretation is supported by other analyses carried out in this study. For example, when extending the fixed-effects model with control variables for demographics, healthcare provision, socioeconomic status, and air pollution, the effect of SANS becomes insignificant. These findings show that the initial positive coefficient might have been caused by an unmeasured development gradient. This suggestion is supported by the regional heterogeneity analysis, which shows significant variability in SANS coefficients across different European contexts [29,32]. Similarly to SANS, electricity access (ELEC) is an indicator that cannot be easily linked to the health effect at the country level. As mentioned above, while the coefficient is negative across all specifications and therefore implies a positive health effect, this indicator may be associated not only with access to electricity but also with more general development factors. In countries with highly developed electrical infrastructure, there tends to be strong general infrastructure, better access to healthcare, better indoor environments, and higher living standards. This means that the negative coefficient estimated for ELEC actually reflects this complex of factors rather than a health effect specifically associated with access to electricity [30]. This applies to freshwater withdrawals (WTRW) too.
As in the case of ELEC, the coefficient of WTRW is negative. This means that the presence of such a predictor implies reduced mortality from respiratory diseases. However, water withdrawal should not be associated only with this outcome. Instead, this country-level indicator may reflect broader differences in water resource management and overall economic, environmental, and institutional development across European countries. In particular, in countries where mobilizing water resources is easier, there is usually much better overall infrastructure and much higher institutional capacity for environmental protection. Consequently, WTRW is associated with these factors and not with water withdrawals alone. In addition, agricultural land use (AGRL) deserves a special discussion. Despite the positive coefficient obtained, the fact that high levels of agricultural land use are associated with higher mortality rates from respiratory diseases should not be explained by direct causation from farming activities. First, AGRL might reflect the extent of the impact of agricultural factors on people’s health, including agricultural emissions, biomass burning, dust particles, pesticides, and fertilizers. Second, agricultural land use may simply be associated with broader differences in environmental conditions, including the nature of the land use. Thus, it is appropriate to interpret the relation in question as a health condition associated with agricultural environments [11,32]. To sum up, interpreting country-level indicators requires great caution, as they are not individual predictors but rather reflect complex conditions associated with them. Therefore, rather than focusing on causal relationships, it is worth discussing the structural characteristics of environmental and infrastructural contexts. The interpretation presented here is consistent with the supplementary robustness checks conducted throughout the process, such as regional heterogeneity tests and extensions of the fixed-effects model. Thus, respiratory mortality rates in Europe depend on multiple factors that act jointly [22,31,33]. See Table 8.

4.6. Proxy Effects, Structural Confounding, and Endogeneity

While robustness checks address concerns about collinearity and omission bias, there may be other problems related to structural confounding and endogeneity. Several country-level indicators used in the study might act as proxies for underlying factors related to broader socioeconomic, institutional, and developmental contexts rather than causal mechanisms [29,30]. For example, measures such as electricity access, sanitation access, freshwater withdrawal, and renewable energy consumption may partly reflect variations in the level of country development and infrastructure, as well as institutional and governance characteristics. Collinearity was assessed using correlation and VIF values, which indicate that there is no multicollinearity among predictors in the regression model. However, despite low multicollinearity, the coefficients may still depend on the overall structure of each country–year pair. Therefore, it is necessary to consider these findings as a part of the general developmental context of each observation point [29]. To address possible structural confounding, a separate fixed-effects specification including population aging, healthcare development, per capita GDP, and PM2.5 exposure was estimated. As the change values for sanitation access and renewable energy consumption show, certain associations could be explained by underlying structural patterns and socioeconomic, demographic, and development-related gradients. The last concern regarding this analysis is endogeneity, which might limit causal interpretations of the results. There are several possible forms of endogeneity that might affect certain relationships. Specifically, reverse causality and simultaneity might apply to relationships such as investment in infrastructure causing respiratory mortality and the simultaneous presence of both strong infrastructure and respiratory mortality in more developed countries. In heterogeneous panel data analysis, endogeneity is a frequent problem, as noted by [27,29,30]. Therefore, one of the main goals of the current study is to identify reliable relationships between certain environmental and socioeconomic factors and respiratory mortality, rather than causal connections between predictors and the response variable [29,30].

5. Mapping Environmental-Health Vulnerability Through Hierarchical Clustering

The quality of the cluster analysis was evaluated based on several statistical measures. In JASP, feature scaling was applied during cluster analysis. In the final hierarchical clustering model, the dissimilarity measure was the Euclidean distance, and agglomeration was defined using the average linkage criterion. The number of clusters used in the final hierarchical model was determined using the Bayesian Information Criterion (BIC), testing alternative options from 1 to 10 clusters based on existing methodologies [34,35]. While density-based clustering fit well within internal validation criteria, it resulted in an extreme polarization of clusters (one major and three minor clusters) that affected the interpretation and representation of environmental-health heterogeneity. Thus, hierarchical clustering was chosen as the primary clustering method. It allowed obtaining a balanced composition of 10 clusters capturing 0.381 of the heterogeneity, with all clusters having positive silhouette coefficients (three clusters with coefficients > 0.6) [35,36] (Table 9).
Even though density-based clustering yielded better results on internal validity criteria, it was chosen as the final clustering procedure because it is more balanced and conceptually interpretable in its representation of heterogeneity. Hierarchical clustering was estimated using feature scaling, Euclidean distance, and the average linkage method in JASP. In addition, the BIC statistic was used to select the best model for hierarchical clustering, with the number of clusters ranging from 1 to 10 [35]. Density-based clustering resulted in a highly unbalanced structure with one main cluster and three small ones, which affected the interpretability of the results. Conversely, 10 country–year clusters consisting of 11, 69, 8, 60, 18, 12, 12, 24, 18, and 6 observations, respectively, were revealed by hierarchical clustering. Moreover, there was greater separation between the clusters, and silhouette values were above zero for hierarchical clustering, with several exceeding 0.6. Thus, hierarchical clustering performed better in the mentioned aspects [36,37]. See Table 10.
Comparison of hierarchical clustering solutions from k = 2 to k = 10 shows a progressive improvement in cluster balance, explained variance, between-cluster separation, and interpretability. As reported in Table 11, lower-k solutions remain strongly polarized and are therefore less suitable for representing the multidimensional heterogeneity of environmental-health conditions. Intermediate solutions improve differentiation but still retain some imbalance in cluster size. More substantial improvements emerge in the higher-k solutions, where cluster structures become more differentiated and less dominated by a single large group. The corrected k = 9 solution confirms that the nine-cluster partition is distinct from the eight-cluster solution. Nevertheless, the k = 10 solution provides the strongest overall compromise between heterogeneity representation, cohesion, separation, model fit, and interpretability. Therefore, k = 10 is retained as the preferred hierarchical clustering solution [35,38,39]. See Table 11.
Figure 6 below highlights the key features of the adopted hierarchical clustering solution (k = 10) as an expression of environmental and health diversity associated with respiratory mortality rates in Europe. See Figure 6.
The upper-left diagram illustrates the number of points in different clusters. Unlike other possible lower k solutions, this structure looks more balanced, not suffering from excessive polarization inherent to density-based clusters. While clusters 2 and 4 continue dominating other clusters, such distribution still indicates reasonable differentiation instead of the concentration of points in one major cluster. This feature positively affects the interpretive value of the classification [35]. The upper-right diagram displays silhouette scores of all clusters. All clusters possess positive silhouette scores and thus can be regarded as structurally consistent with low levels of cohesion and high distance from neighboring clusters. Specifically, clusters 1, 7, 9, and 10 have a score of 0.6 or above, which demonstrates significant structural distinction. Meanwhile, lower silhouette scores of clusters 2 and 4 indicate that these two clusters increase empirical heterogeneity in the dataset, making the hierarchy more complicated yet realistic [35,39]. The lower-left panel presents the within-cluster sum of squares (WSS), reflecting the dispersion within the clusters. As can be seen, clusters 2 and 4 demonstrate significantly higher values of the criterion, while other clusters look relatively compact due to lower WSSs. The key point to be emphasized is that dispersion is spread among several clusters, not concentrated in one oversized cluster as is characteristic of many hierarchical solutions [36]. In addition, the lower-right diagram represents a heatmap illustrating standardized means across all environmental-health variables in different clusters. Clearly, various distinct multidimensional clusters can be seen, especially for the variables of electrification level, dependence on coal, share of renewables, water withdrawals, and cooling degree days. Hence, the choice of the k = 10 hierarchical solution seems rational in terms of representing real differences among clusters in Europe [35,39].
Figure 7 shows the relationship between model fit and cluster compactness across hierarchical clustering results with k = 2–10 clusters. There are three lines illustrating the changes in AIC, BIC, and WSS scores as k increases. Low scores imply better clustering results, representing positive values of model fit, parsimony, and intra-cluster compactness. It is evident that as the number of clusters increases, each criterion decreases, indicating enhanced discrimination of environmental-health differences and lower within-group variability [35]. In general, the most dramatic declines in scores are observed at k = 5–k = 8. From this point on, there are fewer fluctuations and lower slopes, which indicates that any benefits from an increase in k would not be significant. The red dot corresponds to the chosen solution with k = 10, which implies the smallest BIC value. In this case, the red dot is especially meaningful, as the score considers both fit and parsimony, punishing the over-fragmentation. Therefore, the red indicator represents the optimal trade-off among fit, discrimination, and model simplicity, which justifies adopting the 10-cluster hierarchical clustering results [35,39]. See Figure 7.
Figure 8 depicts the final clustering structure obtained using hierarchical clustering (k = 10) based on standardized variables and Euclidean distance with average linkage. See Figure 8.
As shown, the high-dimensional clustering hierarchy is depicted in a two-dimensional space where different colors represent different environmental-health profiles of European countries that influence respiratory mortality patterns. In Figure 8, there is good visual cohesion and separation of clusters that reflects the heterogeneity discovered during hierarchical clustering [35]. A number of clusters appear compact and well-separated, indicating that the current clustering specification is stable. Specifically, clusters 9 and 10, located relatively far away from the main bulk of points, have unique structures. From this, it can be concluded that some countries may exhibit extreme environmental-health profiles characterized by peculiar combinations of electrification rates, climate-related stressors, infrastructure development, and respiratory disease mortality rates. This conclusion is corroborated by silhouette scores above the accepted threshold of 0.6, indicating well-separated clusters [35,36]. As regards the middle zone of the graph, one can note clusters with overlapping points, as they are partially identical and correspond to intermediate environmental-health characteristics. However, the general structure is much better than the density-based alternative because there are not many observations concentrated in a single cluster, and points are evenly distributed across all clusters. Finally, the structure with branches and spatial dispersion is indicative of several interconnected structural determinants of respiratory mortality. Specifically, clusters differ by electrification level, agricultural intensity, coal dependence, renewable energy share, sanitation status, water stress, and climate risk. Thus, Figure 8 shows that mortality rates from respiratory disease depend on multiple interconnected factors rather than a single isolated variable [40]. Taking everything into account, Figure 8 is useful because it shows that the k = 10 hierarchical solution was correct, providing satisfactory results in terms of cohesion, separation, and interpretability.
Hierarchical clustering highlights the connections among various factors and respiratory disease-related deaths in Europe. Electrification is consistently associated with lower mortality rates in areas with high electrification rates. Agricultural density, coal use, and heat stress are all found to increase RD risk, as confirmed by studies on particulates and climatic factors. Water availability and renewables appear to be inversely correlated, with a common tendency for them to be positively correlated with TRD—a consequence of transition phenomena associated with urbanization. This shows how climate can affect health. The most extreme cluster highlights the different national trends in time series: while some countries show very high rates of renewable energy use, others show strong persistence in their coal use. See Table 12.
The dendrogram presented displays hierarchical clustering of environmental health profiles related to respiratory mortality across Europe. This diagram illustrates how data points are hierarchically clustered depending on the similarity between them, considering standardized measures: respiratory mortality (TRD), electricity access (ELEC), agricultural land use (AGRL), freshwater withdrawals (WTRW), cooling degree days (CDDs), coal-based electricity production (COAL), sanitation access (SANS), and renewable energy consumption (RENE). The hierarchical clustering method used in this analysis uses Euclidean distance and average linkage. It allows revealing nested similarity patterns in observations [35,39]. Among the essential features of this dendrogram is the existence of 10 distinguishable clusters, each of which is painted in a particular color. As for the structure of clusters, their separation is illustrated with branch height; the closer the observations are clustered together on the graph, the less dissimilar they are. Thus, several compact clusters can be seen in the graph, having relatively small internal branches and, therefore, high cohesion. Additionally, there are clusters located far apart from one another; their branches connect relatively high up, suggesting significant heterogeneity in the environmental-health profiles of countries in Europe [35,36]. Another advantage of the hierarchical approach is the lack of excess polarization in the graph. Compared with the density-based solution, which clusters all observations into a single dominant group, the presented solution allows observing a relatively uniform distribution of observations across clusters. In doing so, clusters reveal distinct combinations of environmental, climatic, infrastructural, and energy factors associated with respiratory mortality across European regions [35]. Furthermore, certain clusters appear relatively isolated from the others in the graph, suggesting that some environmental-health profiles have unique features associated with extreme values of variables (e.g., renewable energy penetration, coal dependence, or climatic stress). Thus, it shows that the determinants of respiratory mortality are truly multidimensional [37,39]. Consequently, the dendrogram shows that hierarchical clustering is an efficient and easily interpretable technique for revealing the pattern of heterogeneity in environmental health profiles across Europe. See Figure 9.
As shown in the figure, there are 10 clusters, each representing the standardized mean of the environmental-health variables used in this analysis. Each of the 10 clusters has its own unique multidimensional structure. Positive numbers represent above-average cases, whereas negative numbers show below-average cases. A number of clusters exhibit clear-cut differences in characteristics. Clusters 10 and 8 represent cases of high use of renewable energy resources and climatic stresses on one hand, and dependency on coal and higher sanitation scores on the other hand. In turn, Cluster 3 demonstrates low electricity use and higher levels of infrastructural risk factors. Different agricultural, climatic, and energy structures are evident across clusters; this clearly indicates the multidimensionality of the problem and the involvement of environmental and climate-related factors, along with infrastructure, in determining respiratory mortality. In contrast to the polarization of density-based clusters, hierarchical clustering produces a set of balanced, well-defined clusters that avoid collapse into the dominant cluster. Thus, as the figure shows, the hierarchical clustering technique is more appropriate for analyzing respiratory mortality rates across Europe. See Figure 10.

6. Non-Linear Environmental-Health Patterns and Predictive Modelling

This part describes the machine learning analysis, which is intended to complement previous results from the econometric model and clustering techniques by examining the predictive ability of competing models of respiratory disease mortality. It should be noted that, unlike econometric analysis, machine learning does not involve the estimation of causal or inferential relationships between variables and serves only to identify whether indicators associated with environment, infrastructure, climate, and energy consumption provide useful information about TRD.
Table 13 outlines the general scheme used in machine learning analysis. See Table 13.
The sample includes 238 observations with one outcome variable and seven predictors: TRD, ELEC, AGRL, WTRW, CDDs, COAL, SANS, and RENE. These observations were divided into three sets: the first (training) set comprised 64%, while the validation and holdout sets comprised 16% and 20%, respectively. This allows training and tuning on training observations, then evaluating the model on validation and holdout sets that were not used during the estimation process. Prior to estimation, feature scaling was performed because KNN is based on computing distances between objects. In particular, this model uses the Euclidean distance measure with a rectangular kernel, and the number of neighbors is automatically optimized in the range k from 1 to 10. To evaluate selection consistency, an additional 10-fold cross-validation was performed across different k values [41]. Performance measures, including MSE, RMSE, MAE/MAD, MAPE, and R2, were used to quantify model performance and predictive ability. All these metrics allow making inferences about prediction errors and the predictive ability of the models in both relative and absolute terms. Finally, local additive explanations of each model’s predictions in the KNN procedure were used to interpret the contributions of predictors to predictions. Such an explanation can only be regarded as descriptive and is valid only at the local level; it should not be interpreted causally [42,43].
Table 14 presents the predictive performance of the machine learning models used for mortality estimation. See Table 14.
Instead of using solely normalized estimates, we provide the actual figures for MSE, RMSE, MAE/MAD, MAPE, and R2 to facilitate interpretation. Using several evaluation metrics is a standard practice in machine learning, as they complement each other in capturing predictive performance. Under the validation paradigm of our study, the K-Nearest Neighbors (KNN) regression yields the lowest prediction errors and the highest R2 among all models considered. In particular, the KNN model achieves the lowest MSE (10.031), RMSE (3.167), MAE/MAD (1.790), and MAPE (4.22%), along with the highest R2 (0.976). In practice, RMSE indicates that the model produces average prediction errors of around 3.17 deaths per 100,000 inhabitants, while MAPE indicates an average percentage error of 4.22%. Thus, the results suggest good prediction quality under the given validation paradigm. The normalized values in parentheses are equal to 1.000 for KNN, since it is the best-performing model across all indicators in the present comparison. The next-best prediction results are from the Random Forest model, with RMSE = 6.450, MAPE = 14.32%, and R2 = 0.894. These figures imply that tree-based methods capture a substantial portion of the variation in respiratory mortality. Yet, their errors are greater than those of KNN in the present dataset. It is known that Random Forest models often demonstrate high predictive ability when there are nonlinear interactions among predictors. The Decision Tree and Boosting Regression models register mediocre performance. The Decision Tree regression shows RMSE = 10.799 and R2 = 0.635, while the Boosting Regression regression demonstrates RMSE = 12.181 and R2 = 0.509. The figures imply that the mentioned models capture some structure but produce higher prediction errors than KNN and Random Forest models. The predictive power of linear models proves to be weaker than that of other algorithms in the presented comparison. Regularized Linear Regression, Linear Regression, and SVM produce larger RMSEs, smaller R2s, and greater MAPEs. Linear Regression and SVMs, in particular, fail to predict outcomes, as they account for only a small portion of the variance (R2 = 0.158 and 0.081, respectively). Therefore, there may be nonlinear or locally heterogeneous relationships between the predictors and the dependent variable. In conclusion, the KNN method performs best within the proposed validation scheme. Still, one needs to be cautious when interpreting such a result, since it does not imply causality. The KNN results merely indicate better predictions but do not allow us to claim about structural generalizability. Its superiority might be explained by the peculiarities of the selected sample, the chosen variables, and the neighborhood in KNN itself. Thus, the prediction results are rather additional evidence rather than any causal finding. Machine learning must be treated as a tool for picking predictive patterns and supporting panel econometrics and clustering.

6.1. Explaining Environmental-Health Predictions Through Additive Contributions

Indeed, K-Nearest Neighbors (KNN) is essentially a forecasting algorithm and thus cannot be considered inherently interpretable, unlike linear or panel data regressions. In other words, KNN does not return the coefficients that could be directly used for interpretation purposes in terms of marginal effects, causality, and so forth. As such, the interpretability of the applied model in this study relies on after-the-fact explanation procedures, which can be applied to any black-box algorithm [44,45]. To ensure the results were understandable, local additive explanations were generated using the explanation methodology developed by JASP. In essence, this procedure decomposes each forecast into a base prediction and a contribution for each predictor, showing how each predictor increases or decreases the predicted level of respiratory mortality for a given observation. Thus, the explanations reveal how the model works locally and cannot be used to draw conclusions about the global data-generating process [46]. Moreover, the additive contributions obtained as part of the explained analysis should not be confused with coefficients, causalities, elasticities, and other similar measures that allow researchers to understand what happens when predictors change their values. All these metrics are returned as a result of econometric analysis, whereas in the case of KNN, additive contributions merely show the contribution that each predictor makes towards the prediction in question. Thus, the explanations should be seen not as a way of interpreting the algorithm, but rather as a means to increase transparency [45,46]. Overall, the machine learning method can be seen as a complement to the panel data and cluster analysis carried out in the study. While econometrics provides statistical evidence of association, machine learning helps predict possible nonlinear dynamics [46]. See Table 15.
Table 15 presents examples of local additive explanations for KNN predictions of mortality from respiratory diseases (TRD). Local interpretation is based on the “Additive Explanations for Predictions” method provided in JASP, which decomposes each individual prediction into a baseline value (“Base”) and separate additive contributions of predictors. Additive contributions represent the effect of a particular variable on the predicted TRD at a specific test observation point. This method can be used as a part of post hoc interpretability approaches for individual predictions made by machine learning algorithms [43,47]. Five observations were selected because the same number of predictions is shown when the “Explain predictions” option in JASP is automatically applied. That is why it seems reasonable to treat them as a demonstration of local characteristics rather than a representative sample. KNN parameters include Euclidean distance, rectangle weighting, feature scaling, and optimal number of nearest neighbors. Mean drop-out loss is the metric that helps measure the global predictive relevance of predictors in the KNN model. Among the listed predictors, AGRL, RENE, CDDs, and WTRW have the highest mean drop-out loss values and are thus considered the most relevant predictors. As far as local interpretation is concerned, both the sign and the absolute value of additive contributions differ across observations, reflecting diverse environmental-health relations. Positive values are associated with higher predicted TRD, while negative values are associated with lower values of the dependent variable. Local additive contributions are descriptive. In other words, it is necessary to keep in mind that there are no statistical tests, confidence intervals, or bootstrapped estimates for the local interpretations presented above [43,48].
Figure 11 shows the predictive performance and interpretability structure of the K-Nearest Neighbors (KNN) regression model used to predict TRD outcomes in European countries per year.
In the top-left plot, we visualize the agreement between predicted and observed TRD values. As seen, most observations have been predicted with very high accuracy, with almost no prediction error. The observed scatter pattern in the upper-left plot corroborates our previous quantitative results (R2, RMSE, and MAPE). The ability to predict outcomes with high precision indicates that the KNN regression model can capture complex nonlinear relationships between TRD outcomes and explanatory variables. In the top-right plot, we visualize the parameter tuning process as a function of MSE vs. the number of nearest neighbors (k). In this figure, the red point indicates the optimal k value obtained during parameter tuning using cross-validation. According to the plot, the minimum validation error is observed for k = 1. Thus, more precise estimates can be obtained only when the similarity structure is maintained in all possible resolutions. Increasing k leads to progressive oversmoothing and, thus, to growing errors. This effect is caused by an increase in heterogeneity in environmental-health profiles [49]. Mean drop-out loss values calculated using permutations are visualized in the lower-left plot. Agricultural land use (AGRL), renewable energy consumption (RENE), cooling degree days (CDD), and freshwater withdrawals (WTRW) are the most important predictors, contributing to prediction accuracy and representing important structural determinants of TRD outcomes. Finally, in the lower-right plot, we visualize the local additive explanations for selected observations in the test set. Positive values increase the predicted TRD, while negative values decrease it. We observe considerable heterogeneity across observations in their contributions. In particular, sanitation access (SANS), renewable energy consumption (RENE), and agricultural land use (AGRL) exhibit substantial heterogeneity in their local contributions. Thus, the use of machine learning models, along with interpretability methods, may improve the transparency of predictive models [50,51]. It should be mentioned that the results of such analysis should be considered only as predictive, not causal [52].
Figure 12 displays the residual diagnostic plot of the KNN regression used to forecast respiratory disease deaths (TRD). See Figure 12.
The graph shows the model residuals versus predicted TRD. The horizontal red dashed line represents the zero-prediction-error line; hence, all points on this line correspond to cases with perfect predictions of the outcome. It is noticeable that the residuals have been clustered around zero for the most part, indicating that the KNN model performs fairly accurately in predicting country–year observations. The lack of a clear systematic pattern in the residuals indicates that the proposed KNN regression model successfully accounts for non-linear associations between environmental, infrastructural, climate, and energy-based factors and respiratory deaths. Moreover, residuals have not shown high variability when predicting low and moderate respiratory mortality rates using the KNN model. On the contrary, larger deviations in prediction errors are observed among high mortality rates, suggesting inaccuracies in the predictions for a few cases, as expected in such a complex environmental-health system with highly heterogeneous structural features. Importantly, a funnel-shaped residual pattern was not observed in the plot, indicating no strong evidence of heteroscedasticity. In summary, the analyzed residuals provide evidence of the robustness of the KNN model for forecasting respiratory mortality.

6.2. Hyperparameter Optimization and Model Configuration

Table 16 below presents the training, validation, and hyperparameter tuning for the K-Nearest Neighbors regression algorithm used to forecast mortality from respiratory diseases. Table 15 contributes to greater methodological clarity in the application of the machine learning approach, highlighting the processes of model estimation, tuning, and assessment. See Table 16.
The present study uses 238 country–year observations, split into three subsets: training (152 observations), validation (39 observations), and holdout test (47 observations). As such, the use of three subsets allows for model fitting using one subset, model tuning using another subset, and performance assessment on observations excluded from the model-fitting process [53,54]. The choice of the internal validation subset is highly relevant for KNN, as the algorithm is sensitive to the number of neighbors used to make predictions. In this regard, the optimal number of neighbors k = 1 was determined by optimizing k, varying it until the validation mean squared error was minimized among candidates up to k = 10. Hence, the result implies that the best validation performance was obtained with k = 1, indicating that prediction based on the nearest point in the scaled feature space was the best [49]. Feature scaling was applied prior to model estimation, which is correct for the KNN. Specifically, because the KNN algorithm requires measuring distances between observations, variables with large numerical values can easily distort these comparisons. Therefore, scaling ensures that all features (electricity access, agricultural land use, freshwater withdrawals, cooling degree days, coal reliance, sanitation access, and renewable energy use) equally influence the similarity measure. Since Euclidean distance and rectangular weights were used, predictions were created based on the closest observations without weighting distances. The predictive performance on the test subset is excellent, with a test MSE of 10.031, test RMSE of 3.167, test MAE/MAD of 1.790, test MAPE of 4.22%, and an R-squared of 0.976. The results imply that the model produces highly accurate mortality predictions for respiratory diseases when validated using the specified 10-fold cross-validation framework. In practical terms, the RMSE indicates that the average prediction error is approximately 3.17 deaths per 100,000 people. At the same time, MAPE indicates a low average relative error in predictions. However, the optimal hyperparameter k = 1 must be interpreted carefully. Specifically, models with small values of k produce highly accurate predictions but become increasingly sensitive to sample characteristics. Therefore, the optimal hyperparameter must be considered alongside the results of cross-validation stability analysis. Taken together, the provided information suggests that the optimal KNN model specification emerged after exhaustive hyperparameter tuning [50,55]. Despite the positive findings, the results must be considered only as predictive, not causal or inferential, evidence. More precisely, KNN identifies local similarities among observations but does not analyze structural effects or causal relationships.

6.3. KNN Stability Across Alternative Neighborhood Specifications

Table 17 presents an analysis of the sensitivity of the K-Nearest Neighbors (KNN) regression model to alternative neighborhood size parameters.
As the use of the algorithmically derived optimal neighborhood size (k = 2) may lead to overfitting and high sensitivity of predictions to random fluctuations in the data and outliers, the aim of this additional analysis is to examine the model’s stability. To do so, a comparison of performance measures (RMSE, MAE, R2, and their standard deviations calculated via 10-fold cross-validation) was undertaken for various neighborhood sizes (k = 1–25) [56,57]. Accordingly, the results reveal that the KNN model achieves its best predictive performance with very small neighborhood sizes. Specifically, the optimal neighborhood size is k = 2, as confirmed by the corresponding measures of predictive performance: the lowest RMSE (2.067), the lowest MAE (1.342), and the highest R2 (0.986). Notably, similar outcomes are obtained for other small values of k, such as k = 1 and k = 3, indicating that the model works well when its predictions are based on similarity relations between highly localized observations. At the same time, the differences between the predictive performances of KNN for k = 1, 2, and 3 are relatively minor, implying that the strength of predictions is not contingent upon a single parameter value choice. This outcome shows that the KNN model’s high predictive accuracy is stable across the low-neighborhood-size range [58]. Furthermore, it becomes evident that predictive performance progressively deteriorates as the number of neighbors increases. At k = 5, the RMSE measure increases to 3.341, and R2 drops to 0.963. The deterioration continues as the neighborhood size increases, with RMSE reaching 8.782 and R2 dropping to 0.756 for k = 10. Even worse results are observed at k = 15, where RMSE is 11.268, and R2 is 0.571. The largest k sizes examined—k = 20 and k = 25—produce even worse outcomes, with RMSE measures above 12 and R2 below 0.50. Similar trends are observed for MAE, which rises from 1.342 at k = 2 to over 10 at k = 25. This outcome implies that respiratory mortality rates in the dataset have considerable variability across neighboring country–year pairs. On the one hand, small neighborhood sizes allow the model to exploit subtle similarity structures, whereas large values of k lead to overgeneralization and prevent the identification of such patterns. This means that averaging among too many neighboring cases obfuscates true data dynamics. The conclusion drawn is consistent with the overall conclusions regarding the dataset, emphasizing its heterogeneity in environmental and health variables [59,60]. Additionally, the stability diagnostics in Table 17 support the conclusion. Namely, the standard deviations of the RMSE, R2, and MAE values increase as the number of neighboring cases increases. While a large k size is generally recommended in certain studies (for instance, k ≈ √n), it is shown in the current analysis that this guideline might not always guarantee the highest predictive performance for heterogeneous datasets related to environmental and health outcomes [61]. The goal of the analysis described here is to demonstrate that the selection of the KNN model was not accidental but based on a cross-validation procedure that reduced prediction errors. Thus, the results show that low-neighborhood-size values tend to perform better than large ones within the analyzed dataset. However, these findings cannot be considered inferential. Other validation techniques, such as repeated cross-validation or bootstrap resampling, may provide additional information on model robustness and serve as a basis for future research. Still, the results confirm the superiority of the KNN model over other models, while also identifying the importance of local similarity structures for respiratory mortality outcomes [56,57,61].
Figure 13 provides the findings of the KNN stability analysis, conducted with different neighborhood sizes.
The goal is to evaluate the robustness of the KNN algorithm’s predictive accuracy to variations in the neighborhood size parameter, addressing the potential problem of overfitting caused by extremely small neighborhood sizes. To achieve this goal, a set of KNN models was estimated with k = 1 to k = 25 via 10-fold cross-validation as recommended in machine learning studies on hyperparameter tuning [57,59,60]. The dash-dot line marks the optimal value chosen during cross-validation (k = 2), and the dotted vertical line shows the rule-of-thumb neighborhood size equal to approximately √n (approximately 15 neighbors). Panel A plots the development of RMSEs for different numbers of neighbors. The lowest prediction error is achieved at k = 2, indicating the highest accuracy of predictions from the KNN model. Slightly smaller errors are found for k = 1 and k = 3 and can be considered indicative of model stability when a small number of neighbors is used. At the same time, the prediction errors increase systematically as the neighborhood size grows. For all k > 5, the deterioration is considerable, and the errors are high at k = 10, 15, 20, and 25. Panel B illustrates the cross-validated R2s obtained for KNN models. The highest explanatory accuracy is achieved with k = 2, yielding a cross-validated R2 of nearly 0.99. With increasing numbers of neighbors, the explanatory accuracy decreases monotonically until falling to less than 0.60 for k = 15 and less than 0.40 for k = 25. Overall, the two graphs indicate strong localized similarity structures in the dataset used for respiratory mortality. With small neighborhood sizes, these structures remain preserved and yield accurate predictions, whereas larger neighborhood sizes smooth them out excessively, resulting in lower predictive power. In this way, the choice of k = 2 cannot be considered arbitrary, as it is based on evidence from the cross-validation analysis. On the other hand, the analysis shows that KNN predictions are strongly dependent on neighborhood size and highlights the need for careful interpretation of the results [56].

6.4. Limitations and Interpretation of the Machine-Learning Framework

Despite its usefulness as a supplementary source of evidence for predicting respiratory disease mortality, machine learning analysis poses some issues that must be addressed in advance. First, the purpose of the machine-learning analysis differs significantly from that of the econometric analysis in this paper. Specifically, while econometrics focuses on identifying significant relationships between environmental, infrastructure, climate, and energy parameters on the one hand, and respiratory mortality on the other, machine learning aims to achieve maximum predictive power. The difference between the approaches is clearly defined in the modern literature [62,63]. Hence, machine-learning results cannot be treated as inferential but only as predictive evidence. Therefore, such evidence does not imply causality between the individual predictors and the health indicators. Second, the number of observations in the dataset remains relatively limited for machine learning. While there are 238 observations in the country–year panel, this is too few to conduct machine learning according to best-practice guidelines. In particular, it can negatively affect model validity when making out-of-sample predictions. Thus, performance depends on data features, and generalization may be problematic due to limited sample size [64]. Third, K-Nearest Neighbors (KNN) is known to be sensitive to local structure and the choice of the neighborhood parameter k. That is why the revised version of the manuscript adds a separate section on stability analysis. Specifically, the latter is conducted via 10-fold cross-validation on a variety of neighborhood sizes. The importance of validation techniques in addressing overfitting is obvious [64]. The result demonstrates that better performance occurs with small-neighborhood specifications. Predictive power deteriorates rapidly as neighborhood size increases. Hence, the selected specification is stable but dependent on the underlying data structure. Next, despite the interpretability benefits of additive explanations and dropout-loss values, they cannot be equated with econometric parameters or marginal effects. As a matter of fact, the question of whether the difference exists in the context of contemporary debates is crucial to modern methodology [63,65]. Lastly, although the paper uses machine learning to predict mortality from respiratory diseases, this type of analysis complements econometric modeling and clustering. Therefore, the use of KNN results is justified only in the sense of adding information on nonlinear relationships. Hence, following best practices in the predictive analytics literature, the machine-learning evidence needs to supplement the overall framework of environmental determinants of health proposed in this paper.

7. Integration of Econometric, Clustering, and Machine-Learning Evidence

It is important to note that these approaches answer different, yet complementary questions. On the one hand, the panel regression analysis provides inferential information on average within- and between-country associations between environmental determinants and respiratory mortality, after accounting for omitted-variable bias and time effects. Conversely, hierarchical clustering provides no estimates of marginal effects but helps us identify unsupervised profiles of environmental-health relations across different country–year periods. Last, but not least, k-nearest neighbors regression is a predictive modeling exercise that allows us to evaluate how precisely respiratory mortality can be predicted using the set of environmental health indicators and to identify the most relevant predictors among them. These three methods of analysis have their distinct purposes but are also related in that respect: we can consider hierarchical clustering and k-nearest neighbors as methods for exploratory data analysis that complement our inferential modeling. It is important to distinguish between these two types of analysis, and the difference between them is widely discussed in the literature [63,66]. In particular, clustering and k-nearest neighbors regression show a similar pattern with regard to agricultural land use, coal reliance, and climatic stress: all three are risk factors in the panel regression estimates and are part of distinctive country–year clusters as well as key predictors in the k-nearest neighbors analysis. On the other hand, electricity access also shows consistent evidence as a protective structural factor in the panel regression analysis and as an important part of several low-respiratory-mortality health-environment profiles, whereas the contributions of sanitation facilities, renewable energy, and freshwater withdrawals depend on the method used to analyze these variables. Consequently, we cannot interpret clustering and machine learning results as confirmations of our econometric estimates—on the contrary, heterogeneous patterns in these results suggest that these variables might represent contextual factors. See Table 18.

8. Limitations in Data, Inference, and Predictive Modelling

The paper provides a comprehensive examination of environmental, infrastructural, climatic, and energy-related factors associated with respiratory disease mortality in Europe. However, several limitations should be acknowledged. The first limitation concerns the analytical scale. The study is based on a panel of European countries and therefore relies on country-level indicators. While this approach allows for consistent cross-national comparison, it cannot fully capture intra-national variation in environmental exposures, health conditions, urbanization patterns, and pollution intensity. This issue is particularly relevant for PM2.5. In Table 7, the estimated coefficient for PM2.5 is negative and statistically insignificant. This result should not be interpreted as evidence that PM2.5 has no effect on respiratory mortality. Rather, it may reflect the use of annual country-level average PM2.5 values, which are too aggregated to capture the substantial spatial and temporal heterogeneity of particulate matter exposure. Atmospheric pollutants are unevenly distributed across territories and over time, with concentrations varying across cities, industrial areas, rural regions, seasons, and short-term pollution episodes. Consequently, the use of one annual average PM2.5 value per country may attenuate the estimated relationship between particulate matter and mortality. This limitation is consistent with recent evidence showing that anthropogenic impurities in the atmosphere are highly heterogeneous in both space and time [67,68]. Therefore, the PM2.5 result should be interpreted cautiously and should not be considered as contradicting the broader epidemiological literature on the adverse health effects of particulate matter. Future research should use higher-resolution air-quality data, preferably at regional, city, or grid-cell level and with monthly or daily frequency, to better capture exposure variability and its relationship with respiratory mortality.
A related limitation concerns the representation of climate variables. In this study, climate exposure is proxied only by cooling degree days (CDDs), which capture thermal stress and heat-related energy demand. However, CDDs do not describe the atmospheric processes that influence the accumulation, dilution, and dispersion of pollutants near the surface. Relevant mechanisms include boundary-layer depth, temperature inversions, atmospheric stability, wind fields, and air-mass transport. These processes may strengthen or weaken the environmental burden experienced by populations by determining whether pollutants remain trapped near the ground or are dispersed. Therefore, the climatic component of the model should be interpreted as a heat-stress indicator rather than as a complete representation of atmospheric pollution-dispersion mechanisms. Future research should incorporate meteorological and atmospheric-dynamics variables, such as boundary-layer height, wind speed and direction, inversion frequency, and air-mass circulation, preferably at higher spatial and temporal resolution.
The second major limitation concerns the type of dataset analyzed. Although the paper employs fixed-effects and random-effects models and uses Driscoll–Kraay standard errors, it cannot establish causation with a high degree of certainty. The coefficients obtained in the regressions should therefore be interpreted as statistical associations rather than causal effects, as suggested by contemporary studies in causal inference [69,70]. Endogeneity may remain even when advanced specification techniques are used. Future studies could address this issue through difference-in-differences with staggered adoption, instrumental-variable approaches, or double/debiased machine-learning methods. The third limitation concerns the use of variables that may serve as proxies for broader structural conditions rather than isolated mechanisms. For example, access to electricity, sanitation, freshwater withdrawals, and renewable energy consumption may also reflect the level of development, institutional capacity, infrastructure quality, and regional inequality of each country. For this reason, the interpretation of these coefficients requires caution and should be placed within current debates on environmental health, development gradients, and infrastructure transitions. The fourth limitation concerns the imbalance of the panel and the use of listwise deletion. Listwise deletion avoids introducing artificial variation through imputation and is therefore suitable for macro-level environmental indicators. At the same time, it reduces the number of observations, lowers statistical power, and may limit the generalizability of the results. This issue has been highlighted in recent evaluation studies. The fifth limitation refers to the clustering analysis. Although hierarchical clustering reveals meaningful patterns among countries with similar environmental-health profiles, the method is sensitive to scaling procedures, distance measures, and linkage criteria. Therefore, the clusters should be interpreted as exploratory profiles rather than fixed or definitive country classifications. The sixth limitation concerns the machine-learning analysis. K-nearest neighbors showed strong predictive performance within the adopted validation framework, but this does not necessarily imply generalizability or causality. Predictive accuracy does not establish causal interpretation [71]. Moreover, KNN is sensitive to neighborhood composition and to the selected number of neighbors. The relatively small sample size is also a constraint for machine-learning applications. Cross-validation supports the robustness of the low-neighborhood specifications, but further validation with external datasets would strengthen the reliability of the predictions. Similarly, SHAP values and dropout-loss measures are used only to improve interpretability. They should not be interpreted as regression coefficients or causal effects. Finally, the policy recommendations should be interpreted as broad strategic indications rather than location-specific prescriptions. Since the analysis is conducted mainly at the country level, the findings support integrated policy approaches combining air-quality management, infrastructure resilience, energy transition, and climate-change adaptation. However, these conclusions should be further verified using more detailed spatial information and higher-resolution environmental and health data.

9. Urban Vulnerability, Infrastructure Transitions, and Policy Implications

The results could be interpreted within a broader research context focused on the relationships among urban sustainability, environmental health vulnerabilities, infrastructure resilience, and urban sustainability transitions. Even though the empirical findings pertain to country-level data, the identified determinants are directly linked to urban processes that lead to health vulnerabilities in metropolitan areas, intermediate cities, and peri-urban zones. Such interpretation is consistent with recent discussions in urban science on urban sustainability transitions and resilience planning [72]. First of all, it is worth noting that the positive correlation between cooling degree days and respiratory mortality reflects the growing role of heat as a key environmental challenge in contemporary urban development. As a result of warming, European cities will face increasingly hot temperatures and heatwaves. As a result, densely populated environments will become more vulnerable to health issues [72]. Furthermore, the positive correlation between coal dependency and respiratory mortality could be explained by the fact that urbanization in these cities will generate additional environmental health risks related to increased emissions and higher pollution levels. Therefore, the results support the idea of an accelerated urban energy transition and the transformation of urban environments into greener ones [73]. Additionally, the negative correlation between health outcomes and access to electricity could be explained using the theory of urban infrastructural resilience. The availability of basic services contributes to improved living conditions and reduces the risk of respiratory problems caused by adverse environmental factors [72]. In conclusion, the results indicate significant heterogeneity of environmental health patterns across European countries. In such situations, it is necessary to adopt regionalized policies that consider a wide range of factors, including climate conditions, urban infrastructures, energy systems, and environmental exposures. Instead of introducing unified urban policies, public health authorities and urban planners should focus on regionalized solutions tailored to specific territories. Overall, it is possible to conclude that urban health indicators are strongly associated with environmental changes, infrastructure development, and climate adaptations. Effective public health interventions require comprehensive approaches that draw on a wide range of environmental, energy, infrastructure, and climate factors [72,73]. See Table 19.

10. Conclusions

This paper examined the relationship between respiratory disease mortality and environmental, infrastructural, climatic, and energy-related conditions in Europe through an integrated methodological framework combining panel-data econometrics, robustness checks, hierarchical clustering, and machine-learning prediction. The study was based on an initial dataset of 41 countries over the period 2010–2021, with an effective estimation sample of 38 cross-sectional units and 238 country–year observations after accounting for data availability. Overall, the findings show that respiratory mortality is associated with broader structural environmental-health conditions rather than isolated causal mechanisms. The econometric results indicate that land use, fossil-fuel dependence, infrastructure access, water-resource conditions, and climatic stress are relevant dimensions of respiratory vulnerability. However, the sensitivity of some coefficients across robust specifications confirms that these variables should be interpreted as country-level structural indicators rather than direct causal effects. In this sense, the results highlight the importance of development gradients, infrastructure quality, energy systems, environmental exposure, and regional heterogeneity in shaping respiratory-health outcomes. The clustering and machine-learning analyses complement the econometric evidence by showing that respiratory mortality patterns are heterogeneous across country–year profiles and cannot be reduced to a single European trajectory. Hierarchical clustering identifies differentiated environmental-health configurations, while the predictive models confirm that nonlinear similarities among observations can provide additional insight into vulnerability patterns. These methods are not used to infer causality, but to strengthen the interpretation of heterogeneity, prediction, and profile-based environmental-health risk. From an urban science perspective, the study suggests that respiratory mortality is embedded in broader systems connecting air quality, energy transition, infrastructure resilience, climatic exposure, and water-resource management. Although the analysis is conducted at the country level, the findings are relevant for metropolitan, intermediate, and peri-urban areas because these territories are shaped by the same structural conditions. Future research should employ higher-resolution spatial and temporal data to better capture local exposure, intra-national inequalities, atmospheric dynamics, and city-level vulnerability. Policy responses should therefore be integrated and region-sensitive, combining air-quality management, resilient infrastructure, clean-energy transition, climate adaptation, and public-health planning.

Author Contributions

Conceptualization, E.R., O.R., P.L., A.C., A.L.; methodology, E.R., O.R., P.L., A.C., A.L.; validation, E.R., O.R., P.L., A.C., A.L.; formal analysis, E.R., O.R., P.L., A.C., A.L.; investigation, E.R., O.R., P.L., A.C., A.L.; resources, E.R., O.R., P.L., A.C., A.L.; data curation, E.R., O.R., P.L., A.C., A.L.; writing—original draft preparation, E.R., O.R., P.L., A.C., A.L.; writing—review and editing, E.R., O.R., P.L., A.C., A.L.; supervision, E.R., O.R., P.L., A.C., A.L.; project administration E.R., O.R., P.L., A.C., A.L. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data were acquired from the following websites: Institute for Health Metrics and Evaluation, Link: https://www.healthdata.org/ (accessed on 10 January 2025), World Bank, Sovereign ESG Data Portal, https://databank.worldbank.org/source/world-development-indicators (accessed on 12 January 2025).

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

AcronymDefinition
TRDMortality Rate by Respiratory Disease
ELECAccess to Electricity
AGRLAgricultural Land
WTRWFreshwater Withdrawals
CDDsCooling Degree Days
COALCoal Electricity
SANSSafe Sanitation
RENERenewable Energy
KNNK-Nearest Neighbors
MSEMean Squared Error
RMSERoot Mean Squared Error
MAEMean Absolute Error
MAPEMean Absolute Percentage Error
SVMSupport Vector Machine
ANNArtificial Neural Network

References

  1. Kriit, H.K.; Chen-Xu, J.; Semenza, J.C.; Heiliger, H.; Markandya, A.; Dasandi, N.; Jankin, S.; van Daalen, K.R.; Achebak, H.; Alari, A.; et al. The 2026 Europe report of the Lancet Countdown on health and climate change: Narrowing window for decisive health action. Lancet Public Health 2026, 11, e386–e407. [Google Scholar] [CrossRef] [PubMed]
  2. Vicedo-Cabrera, A.M.; Melén, E.; Forastiere, F.; Gehring, U.; Katsouyanni, K.; Yorgancioglu, A.; Ulrik, C.S.; Hansen, K.; Powell, P.; Ward, B.; et al. Climate change and respiratory health: A European Respiratory Society position statement. Eur. Respir. J. 2023, 62, 2201960. [Google Scholar] [CrossRef] [PubMed]
  3. Tran, H.M.; Tsai, F.-J.; Lee, Y.-L.; Chang, J.-H.; Chang, L.-T.; Chang, T.-Y.; Chung, K.F.; Kuo, H.-P.; Lee, K.-Y.; Chuang, K.-J.; et al. The impact of air pollution on respiratory diseases in an era of climate change: A review of the current evidence. Sci. Total Environ. 2023, 898, 166340. [Google Scholar] [CrossRef] [PubMed]
  4. Im, U.; Ye, Z.; Schuhen, N.; Chowdhury, S.; Christensen, J.H.; Geels, C.; Hänninen, R.; Hodnebrog, Ø.; Marelle, L.; Sofiev, M.; et al. Europe will struggle to meet the new WHO Air Quality Guidelines. npj Clean Air 2025, 1, 13. [Google Scholar] [CrossRef]
  5. Robben, J.; Antonio, K.; Kleinow, T. The short-term association between environmental variables and mortality: Evidence from Europe. J. R. Stat. Soc. Ser. A Stat. Soc. 2026, 189, 1131–1153. [Google Scholar]
  6. Liebensteiner, M.; Kimani, A.M. Environmental and health costs of Europe’s shift from gas to coal amidst the energy crisis. J. Econ. Behav. Organ. 2026, 241, 107397. [Google Scholar] [CrossRef]
  7. Tasev, G.; Makreski, P.; Jovanovski, G.; Životić, D.; Boev, I.; Jelenkovic, R. The environmental and health damage caused by the use of coal. ChemTexts 2025, 11, 3. [Google Scholar] [CrossRef]
  8. Adamopoulos, I.; Syrou, N.; Vito, D. Climate Change Risks and Impacts on Public Health Correlated with Air Pollution—African Dust in South Europe. Med. Sci. Forum 2025, 33, 1. [Google Scholar] [CrossRef]
  9. Pashley, C.; Hemming, D.; Adams-Groom, B.; Borman, A.; Johnson, E.; Anees-Hill, S.; Skjøth, C. Outdoor airborne allergenic pollen & fungal spores. In Health Effects of Climate Change (HECC) Report; UK Health Security Agency (UKHSA): London, UK, 2023. [Google Scholar]
  10. Genowska, A.; Strukcinskiene, B.; Jamiołkowski, J.; Abramowicz, P.; Konstantynowicz, J. Emission of in-dustrial air pollution and mortality due to respiratory diseases: A birth cohort study in Poland. Int. J. Environ. Res. Public Health 2023, 20, 1309. [Google Scholar] [CrossRef] [PubMed]
  11. Liu, S.; Lim, Y.-H.; Chen, J.; Strak, M.; Wolf, K.; Weinmayr, G.; Rodopolou, S.; de Hoogh, K.; Bellander, T.; Brandt, J.; et al. Long-term air pollution exposure and pneumonia-related mortality in a large pooled European cohort. Am. J. Respir. Crit. Care Med. 2022, 205, 1429–1439. [Google Scholar] [CrossRef] [PubMed]
  12. Nazar, W.; Niedoszytko, M. Air pollution in Poland: A 2022 narrative review with focus on respiratory dis-eases. Int. J. Environ. Res. Public Health 2022, 19, 895. [Google Scholar] [CrossRef] [PubMed]
  13. Liu, Y.; Liu, S.; Huang, J.; Liu, Y.; Wang, Q.; Chen, J.; Sun, L.; Tu, W. Mitochondrial dysfunction in metabolic disorders induced by per-and polyfluoroalkyl substance mixtures in zebrafish larvae. Environ. Int. 2023, 176, 107977. [Google Scholar] [CrossRef] [PubMed]
  14. Anenberg, S.C.; Henze, D.K.; Tinney, V.; Kinney, P.L.; Raich, W.; Fann, N.; Malley, C.S.; Roman, H.; Lamsal, L.; Duncan, B.; et al. Estimates of the global burden of ambient PM2. 5, ozone, and NO2 on asthma incidence and emergency room visits. Environ. Health Perspect. 2018, 126, 107004. [Google Scholar] [CrossRef] [PubMed]
  15. Alahmad, B.; Khraishah, H.; Royé, D.; Vicedo-Cabrera, A.M.; Guo, Y.; Papatheodorou, S.I.; Achilleos, S.; Acquaotta, F.; Armstrong, B.; Bell, M.L.; et al. Associations between extreme temperatures and cardiovascular cause-specific mortality: Results from 27 countries. Circulation 2023, 147, 35–46. [Google Scholar] [CrossRef] [PubMed]
  16. Castaño-Rosa, R.; Barrella, R.; Sánchez-Guevara, C.; Barbosa, R.; Kyprianou, I.; Paschalidou, E.; Thomaidis, N.S.; Dokupilova, D.; Gouveia, J.P.; Kádár, J.; et al. Cooling degree models and future energy demand in the residential sector. A seven-country case study. Sustainability 2021, 13, 2987. [Google Scholar] [CrossRef]
  17. Egbo, O.P.; Ezeaku, H.C.; Okolo, V.O.; Anisiuba, C.A.; Ibe, G.I.; Okeke, O.M.; Igwe, P.A. Enhancing agricultural and industrial productivity through freshwater withdrawals and management: Implications for the BRICS countries. Environ. Dev. Sustain. 2023, 25, 3771–3799. [Google Scholar]
  18. Nishimwe, A.M.R.; Reiter, S. Estimation, analysis and mapping of electricity consumption of a regional building stock in a temperate climate in Europe. Energy Build. 2021, 253, 111535. [Google Scholar] [CrossRef]
  19. Kim, J.H. Multicollinearity and misleading statistical results. Korean J. Anesthesiol. 2019, 72, 558–569. [Google Scholar] [CrossRef] [PubMed]
  20. Daoud, J.I. Multicollinearity and regression analysis. J. Phys. Conf. Ser. 2017, 949, 012009. [Google Scholar] [CrossRef]
  21. Senaviratna, N.A.M.R.; Cooray, T.M.J.A. Diagnosing multicollinearity of logistic regression model. Asian J. Probab. Stat. 2019, 5, 1–9. [Google Scholar] [CrossRef]
  22. van Daalen, K.R.; Tonne, C.; Semenza, J.C.; Rocklöv, J.; Markandya, A.; Dasandi, N.; Jankin, S.; Achebak, H.; Ballester, J.; Bechara, H.; et al. The 2024 Europe report of the Lancet Countdown on health and climate change: Unprecedented warming demands unprecedented action. Lancet Public Health 2024, 9, e495–e522. [Google Scholar] [CrossRef] [PubMed]
  23. Huang, W.T.K.; Masselot, P.; Bou-Zeid, E.; Fatichi, S.; Paschalis, A.; Sun, T.; Gasparrini, A.; Manoli, G. Economic valuation of temperature-related mortality attributed to urban heat islands in European cities. Nat. Commun. 2023, 14, 7438. [Google Scholar] [CrossRef] [PubMed]
  24. Caruso, G.; Colantonio, E.; Gattone, S.A. Relationships between renewable energy consumption, social factors, and health: A panel vector auto regression analysis of a cluster of 12 EU countries. Sustainability 2020, 12, 2915. [Google Scholar] [CrossRef]
  25. Hoechle, D. Robust standard errors for panel regressions with cross-sectional dependence. Stata J. 2007, 7, 281–312. [Google Scholar] [CrossRef]
  26. Chudik, A.; Mohaddes, K.; Pesaran, M.H.; Raissi, M. Long-Run Effects in Large Heterogeneous Panel Data Models with Cross-Sectionally Correlated Errors. In Essays in Honor of Man Ullah; Emerald Group Publishing Limited: Leeds, UK, 2016; Volume 36, pp. 85–135. [Google Scholar]
  27. Juodis, A.; Karavias, Y.; Sarafidis, V. A homogeneous approach to testing for Granger non-causality in heterogeneous panels. Empir. Econ. 2021, 60, 93–112. [Google Scholar]
  28. Cox, N.J. Speaking Stata: Loops, again and again. Stata J. 2021, 20, 999–1015. [Google Scholar]
  29. Henningsen, A.; Low, G.; Wuepper, D.; Dalhaus, T.; Storm, H.; Belay, D.; Hirsch, S. Estimating causal effects with observational data: Guidelines for agricultural and applied economists. J. Agric. Econ. 2025, 77, 356–382. [Google Scholar] [CrossRef]
  30. Roy, A. A panel data study on the role of clean energy in promoting life expectancy. Dialogues Health 2025, 6, 100201. [Google Scholar] [CrossRef] [PubMed]
  31. Lupu, I.; Hurduzeu, G.; Lupu, R.; Popescu, M.-F.; Gavrilescu, C. Drivers for Renewable Energy Consumption in European Union Countries. A Panel Data Insight. Amfiteatru Econ. 2023, 25, 380–396. [Google Scholar] [CrossRef]
  32. Popescu, G.H.; Nica, E.; Kliestik, T.; Alpopi, C.; Bîgu, A.-M.P.; Niță, S.-C. The impact of ecological footprint, urbanization, education, health expenditure, and industrialization on child mortality: Insights for environment and public health in Eastern Europe. Int. J. Environ. Res. Public Health 2024, 21, 1379. [Google Scholar] [CrossRef] [PubMed]
  33. Apostu, S.A.; Panait, M.; Balsalobre-Lorente, D.; Ferraz, D.; Rădulescu, I.G. Energy transition in non-euro countries from central and eastern Europe: Evidence from panel vector error correction model. Energies 2022, 15, 9118. [Google Scholar] [CrossRef]
  34. Charrad, M.; Ghazzali, N.; Boiteau, V.; Niknafs, A. NbClust: An R package for determining the relevant number of clusters in a data set. J. Stat. Softw. 2014, 61, 1–36. [Google Scholar] [CrossRef]
  35. Akhanli, S.E.; Hennig, C. Comparing clusterings and numbers of clusters by aggregation of calibrated clustering validity indexes. Stat. Comput. 2020, 30, 1523–1544. [Google Scholar] [CrossRef]
  36. Srinivasan, A.; Xue, L.; Zhan, X. Identification of microbial features in multivariate regression under false discovery rate control. Comput. Stat. Data Anal. 2023, 181, 107621. [Google Scholar] [CrossRef]
  37. Plante, R.L.; Becker, C.A.; Medina-Smith, A.; Brady, K.; Dima, A.; Long, B.; Bartolo, L.M.; Warren, J.A.; Hanisch, R.J. Implementing a registry federation for materials science data discovery. Data Sci. J. 2021, 20, 10–5334. [Google Scholar] [CrossRef]
  38. Vendramin, L.; Campello, R.J.G.B.; Hruschka, E.R. Relative clustering validity criteria: A comparative overview. Stat. Anal. Data Min. ASA Data Sci. J. 2010, 3, 209–235. [Google Scholar] [CrossRef]
  39. Zhou, S.; Xu, H.; Zheng, Z.; Chen, J.; Li, Z.; Bu, J.; Wu, J.; Wang, X.; Zhu, W.; Ester, M. A comprehensive survey on deep clustering: Taxonomy, challenges, and future directions. ACM Comput. Surv. 2024, 57, 1–38. [Google Scholar] [CrossRef]
  40. Bi, X.; Nie, H.; Zhang, X.; Zhao, X.; Yuan, Y.; Wang, G. Unrestricted multi-hop reasoning network for inter-pretable question answering over knowledge graph. Knowl.-Based Syst. 2022, 243, 108515. [Google Scholar] [CrossRef]
  41. Chaudhary, A.; Chouhan, K.S.; Gajrani, J.; Sharma, B. Deep learning with PyTorch. Machine Learning and Deep Learning in Real-Time Applications; Engineering Science Reference: Hershey, PA, USA, 2020; pp. 61–95. [Google Scholar]
  42. Mi, J.-X.; Li, A.-D.; Zhou, L.-F. Review study of interpretation methods for future interpretable machine learning. IEEE Access 2020, 8, 191969–191985. [Google Scholar] [CrossRef]
  43. Lundberg, S.M.; Erion, G.; Chen, H.; DeGrave, A.; Prutkin, J.M.; Nair, B.; Katz, R.; Himmelfarb, J.; Bansal, N.; Lee, S.-I. From local explanations to global understanding with explainable AI for trees. Nat. Mach. Intell. 2020, 2, 56–67. [Google Scholar] [CrossRef] [PubMed]
  44. Belle, V.; Papantonis, I. Principles and practice of explainable machine learning. Front. Big Data 2021, 4, 688969. [Google Scholar] [CrossRef] [PubMed]
  45. Arrieta, A.B.; Díaz-Rodríguez, N.; Del Ser, J.; Bennetot, A.; Tabik, S.; Barbado, A.; Garcia, S.; Gil-Lopez, S.; Molina, D.; Benjamins, R.; et al. Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Inf. Fusion 2020, 58, 82–115. [Google Scholar] [CrossRef]
  46. Vilone, G.; Longo, L. Notions of explainability and evaluation approaches for explainable artificial intelligence. Inf. Fusion 2021, 76, 89–106. [Google Scholar] [CrossRef]
  47. Madsen, A.; Reddy, S.; Chandar, S. Post-hoc interpretability for neural nlp: A survey. ACM Comput. Surv. 2022, 55, 1–42. [Google Scholar] [CrossRef]
  48. Murdoch, W.J.; Singh, C.; Kumbier, K.; Abbasi-Asl, R.; Yu, B. Definitions, methods, and applications in interpretable machine learning. Proc. Natl. Acad. Sci. USA 2019, 116, 22071–22080. [Google Scholar] [CrossRef] [PubMed]
  49. Bischl, B.; Binder, M.; Lang, M.; Pielok, T.; Richter, J.; Coors, S.; Thomas, J.; Ullmann, T.; Becker, M.; Boulesteix, A.; et al. Hyperparameter optimization: Foundations, algorithms, best practices, and open challenges. Wiley Interdiscip. Rev. Data Min. Knowl. Discov. 2023, 13, e1484. [Google Scholar] [CrossRef]
  50. Biecek, P.; Burzykowski, T. Explanatory Model Analysis: Explore, Explain, and Examine Predictive Models; Chapman & Hall/CRC: Boca Raton, FL, USA, 2021. [Google Scholar]
  51. Minakova, S.; Stefanov, T. Memory-throughput trade-off for CNN-based applications at the edge. ACM Trans. Des. Autom. Electron. Syst. 2022, 28, 1–26. [Google Scholar] [CrossRef]
  52. Bzdok, D.; Krzywinski, M.; Altman, N. Machine learning: Supervised methods. Nat. Methods 2018, 15, 5–6. [Google Scholar] [CrossRef] [PubMed]
  53. Balise, R.R.; Grealis, K.; Brandt, L.; Luo, S.X.; Odom, G.; Jean-Baptiste, G.; Feaster, D.J. Applying machine learning in predicting medication treatment outcomes for opioid use disorder. J. Subst. Use Addict. Treat. 2025, 182, 209847. [Google Scholar] [CrossRef] [PubMed]
  54. Karashchuk, P.; Tuthill, J.C.; Brunton, B.W. The DANNCE of the rats: A new toolkit for 3D tracking of animal behavior. Nat. Methods 2021, 18, 460–462. [Google Scholar] [CrossRef] [PubMed]
  55. Cai, J.; Mandelbaum, A.; Nagaraja, C.H.; Shen, H.; Zhao, L. Statistical theory powering data science. Stat. Sci. 2019, 34, 669–691. [Google Scholar] [CrossRef]
  56. Nader, Y.; Sixt, L.; Landgraf, T. DNNR: Differential nearest neighbors regression. Int. Conf. Mach. Learn. 2022, 162, 16296–16317. [Google Scholar]
  57. Arif, M.S.; Mukheimer, A.; Asif, D. Enhancing the early detection of chronic kidney disease: A robust machine learning model. Big Data Cogn. Comput. 2023, 7, 144. [Google Scholar] [CrossRef]
  58. Gajan, S. Modeling of seismic energy dissipation of rocking foundations using nonparametric machine learning algorithms. Geotechnics 2021, 1, 534–557. [Google Scholar] [CrossRef]
  59. Bouqentar, M.A.; Terrada, O.; Hamida, S.; Saleh, S.; Lamrani, D.; Cherradi, B.; Raihani, A. Early heart disease prediction using feature engineering and machine learning algorithms. Heliyon 2024, 10, e38731. [Google Scholar] [CrossRef] [PubMed]
  60. Li, Q.; Zhang, Z.; Ma, Z. Raman spectral pattern recognition of breast cancer: A machine learning strategy based on feature fusion and adaptive hyperparameter optimization. Heliyon 2023, 9, e18148. [Google Scholar] [CrossRef] [PubMed]
  61. Marco, R.; Ahmad, S.S.S.; Ahmad, S. Bayesian hyperparameter optimization and ensemble learning for machine learning models on software effort estimation. Int. J. Adv. Comput. Sci. Appl. 2022, 13, 3. [Google Scholar] [CrossRef]
  62. Shmueli, G.; Koppius, O.R. Predictive analytics in information systems research. MIS Q. 2011, 35, 553–572. [Google Scholar] [CrossRef]
  63. Athey, S.; Imbens, G.W. Machine learning methods that economists should know about. Annu. Rev. Econ. 2019, 11, 685–725. [Google Scholar] [CrossRef]
  64. Poldrack, R.A.; Huckins, G.; Varoquaux, G. Establishment of best practices for evidence for prediction: A review. JAMA Psychiatry 2020, 77, 534–540. [Google Scholar] [CrossRef] [PubMed]
  65. Haring, R.; Ghannad, M.; Bertizzolo, L.; Page, M.J. No evidence found for an association between trial characteristics and treatment effects in randomized trials of testosterone therapy in men: A meta-epidemiological study. J. Clin. Epidemiol. 2020, 122, 12–19. [Google Scholar] [CrossRef] [PubMed]
  66. Mullainathan, S.; Spiess, J. Machine learning: An applied econometric approach. J. Econ. Perspect. 2017, 31, 87–106. [Google Scholar] [CrossRef]
  67. Shikhovtsev, M.Y.; Makarov, M.M.; Aslamov, I.A.; Tyurnev, I.N.; Molozhnikova, Y.V. Application of Modern Low-Cost Sensors for Monitoring of Particle Matter in Temperate Latitudes: An Example from the Southern Baikal Region. Sustainability 2025, 17, 3585. [Google Scholar] [CrossRef]
  68. Shikhovtsev, M.Y.; Shikhovtsev, A.Y.; Kovadlo, P.G.; Obolkin, V.A.; Molozhnikova, Y.V. Influence of Air Movement Structure on Microphysical Properties of the Atmosphere over Listvyanka. Atmos. Ocean Opt. 2025, 38, 180–187. [Google Scholar] [CrossRef]
  69. Hernán, M.A.; Robins, J.M. Causal Inference: What If; Chapman & Hall/CRC: Boca Raton, FL, USA, 2020. [Google Scholar]
  70. Angrist, J.D.; Pischke, J.S. Mastering ’Metrics: The Path from Cause to Effect; Princeton University Press: Princeton, NJ, USA, 2014. [Google Scholar]
  71. Varoquaux, G.; Cheplygina, V. Machine learning for medical imaging: Methodological failures and recommendations for the future. npj Digit. Med. 2022, 5, 48. [Google Scholar] [CrossRef] [PubMed]
  72. Sharifi, A. Urban sustainability assessment: An overview and bibliometric analysis. Ecol. Indic. 2021, 121, 107102. [Google Scholar] [CrossRef]
  73. Mihajlović, B. The role of consumers in the achievement of corporate sustainability through the reduction of unfair commercial practices. Sustainability 2020, 12, 1009. [Google Scholar] [CrossRef]
Figure 1. Integrated Methodological Framework for Environmental Determinants of Respiratory Health. Note: The figure illustrates the seven-stage analytical pipeline, from data collection to policy implications, combining panel econometrics, hierarchical clustering, and machine learning to examine environmental, infrastructural, climatic, and energy-related determinants of respiratory mortality across 38 European countries (2010–2021).
Figure 1. Integrated Methodological Framework for Environmental Determinants of Respiratory Health. Note: The figure illustrates the seven-stage analytical pipeline, from data collection to policy implications, combining panel econometrics, hierarchical clustering, and machine learning to examine environmental, infrastructural, climatic, and energy-related determinants of respiratory mortality across 38 European countries (2010–2021).
Urbansci 10 00351 g001
Figure 2. Geographical distribution of the 41 countries initially included in the empirical analysis. Countries belonging to the European continental context are highlighted in blue. The effective econometric panel was reduced to 38 cross-sectional units over 2010–2021 due to data availability constraints.
Figure 2. Geographical distribution of the 41 countries initially included in the empirical analysis. Countries belonging to the European continental context are highlighted in blue. The effective econometric panel was reduced to 38 cross-sectional units over 2010–2021 due to data availability constraints.
Urbansci 10 00351 g002
Figure 3. Correlation Structure of Environmental and Infrastructural Determinants of Respiratory Mortality in Europe. Note. Correlation heatmap of baseline regressors included in the panel estimations for respiratory mortality across European countries. The figure highlights moderate associations among environmental and infrastructural variables, while excluding severe multicollinearity patterns that could compromise econometric inference.
Figure 3. Correlation Structure of Environmental and Infrastructural Determinants of Respiratory Mortality in Europe. Note. Correlation heatmap of baseline regressors included in the panel estimations for respiratory mortality across European countries. The figure highlights moderate associations among environmental and infrastructural variables, while excluding severe multicollinearity patterns that could compromise econometric inference.
Urbansci 10 00351 g003
Figure 4. Environmental-Health Vulnerability and Respiratory Mortality in Europe: Integrated Evidence from Econometrics, Clustering, and Machine Learning. Note: The infographic summarizes the main empirical findings, highlighting environmental risk factors, infrastructural protection mechanisms, 10 country–year vulnerability profiles, and KNN predictive performance. This figure is an authors’ elaboration created with NotebookLM.
Figure 4. Environmental-Health Vulnerability and Respiratory Mortality in Europe: Integrated Evidence from Econometrics, Clustering, and Machine Learning. Note: The infographic summarizes the main empirical findings, highlighting environmental risk factors, infrastructural protection mechanisms, 10 country–year vulnerability profiles, and KNN predictive performance. This figure is an authors’ elaboration created with NotebookLM.
Urbansci 10 00351 g004
Figure 5. Baseline and Extended Fixed-Effects Estimates for Respiratory Disease Mortality in Europe. Note: The figure compares coefficient estimates from baseline and extended fixed-effects models, illustrating the robustness of key predictors and the influence of additional demographic, socioeconomic, healthcare, and pollution controls on estimated associations.
Figure 5. Baseline and Extended Fixed-Effects Estimates for Respiratory Disease Mortality in Europe. Note: The figure compares coefficient estimates from baseline and extended fixed-effects models, illustrating the robustness of key predictors and the influence of additional demographic, socioeconomic, healthcare, and pollution controls on estimated associations.
Urbansci 10 00351 g005
Figure 6. Diagnostic Evaluation of the Ten-Cluster Hierarchical Solution for Environmental-Health Profiles. Note: The figure summarizes the adopted hierarchical clustering solution. The upper-left panel reports cluster-size distribution, showing how observations are allocated across the 10 clusters. The upper-right panel presents silhouette scores, indicating cluster cohesion and separation. The lower-left panel shows within-cluster sum of squares, measuring internal dispersion. The lower-right heatmap displays standardized cluster means across environmental-health variables, highlighting distinct multidimensional vulnerability profiles.
Figure 6. Diagnostic Evaluation of the Ten-Cluster Hierarchical Solution for Environmental-Health Profiles. Note: The figure summarizes the adopted hierarchical clustering solution. The upper-left panel reports cluster-size distribution, showing how observations are allocated across the 10 clusters. The upper-right panel presents silhouette scores, indicating cluster cohesion and separation. The lower-left panel shows within-cluster sum of squares, measuring internal dispersion. The lower-right heatmap displays standardized cluster means across environmental-health variables, highlighting distinct multidimensional vulnerability profiles.
Urbansci 10 00351 g006
Figure 7. Model-Fit and Compactness Criteria for Hierarchical Clustering Solutions. Note: The figure reports AIC, BIC, and WSS values across alternative hierarchical clustering solutions from k = 2 to k = 10. The red marker identifies the selected ten-cluster solution.
Figure 7. Model-Fit and Compactness Criteria for Hierarchical Clustering Solutions. Note: The figure reports AIC, BIC, and WSS values across alternative hierarchical clustering solutions from k = 2 to k = 10. The red marker identifies the selected ten-cluster solution.
Urbansci 10 00351 g007
Figure 8. Two-Dimensional Projection of the Ten-Cluster Hierarchical Environmental-Health Structure. Note: The figure visualizes the ten-cluster hierarchical solution in two-dimensional space, showing cluster cohesion, separation, and overlap among country–year observations, while highlighting heterogeneous environmental-health profiles associated with respiratory mortality patterns.
Figure 8. Two-Dimensional Projection of the Ten-Cluster Hierarchical Environmental-Health Structure. Note: The figure visualizes the ten-cluster hierarchical solution in two-dimensional space, showing cluster cohesion, separation, and overlap among country–year observations, while highlighting heterogeneous environmental-health profiles associated with respiratory mortality patterns.
Urbansci 10 00351 g008
Figure 9. Hierarchical Dendrogram of Environmental-Health Profiles Associated with Respiratory Mortality in Europe. Note: The dendrogram illustrates the hierarchical structure of 10 environmental-health clusters based on standardized variables. Branch heights represent dissimilarity levels, while colored segments identify cluster membership, highlighting multidimensional heterogeneity across European country–year observations.
Figure 9. Hierarchical Dendrogram of Environmental-Health Profiles Associated with Respiratory Mortality in Europe. Note: The dendrogram illustrates the hierarchical structure of 10 environmental-health clusters based on standardized variables. Branch heights represent dissimilarity levels, while colored segments identify cluster membership, highlighting multidimensional heterogeneity across European country–year observations.
Urbansci 10 00351 g009
Figure 10. Hierarchical clustering reveals structural divergence in environmental-health profiles and respiratory mortality across Europe. Note: Ten clusters capture contrasting trajectories. Electrification is protective, coal and agriculture increase risk, while water scarcity and heat stress intensify fragility. Renewables and sanitation reflect transitional dynamics.
Figure 10. Hierarchical clustering reveals structural divergence in environmental-health profiles and respiratory mortality across Europe. Note: Ten clusters capture contrasting trajectories. Electrification is protective, coal and agriculture increase risk, while water scarcity and heat stress intensify fragility. Renewables and sanitation reflect transitional dynamics.
Urbansci 10 00351 g010
Figure 11. KNN Predictive Performance, Parameter Optimization, Feature Importance, and Local Additive Explanations. Note: The figure summarizes the KNN modelling framework, showing prediction accuracy, cross-validated parameter optimization, permutation-based feature importance, and local additive explanations. Together, these diagnostics improve transparency regarding predictive performance, variable relevance, and observation-specific prediction behavior.
Figure 11. KNN Predictive Performance, Parameter Optimization, Feature Importance, and Local Additive Explanations. Note: The figure summarizes the KNN modelling framework, showing prediction accuracy, cross-validated parameter optimization, permutation-based feature importance, and local additive explanations. Together, these diagnostics improve transparency regarding predictive performance, variable relevance, and observation-specific prediction behavior.
Urbansci 10 00351 g011
Figure 12. Residual Diagnostics of the K-Nearest Neighbors Regression Model for Respiratory Mortality Prediction. Note: The figure presents residuals against predicted TRD values for the KNN model. Residuals are generally centered around zero without systematic patterns, supporting predictive adequacy while highlighting larger forecasting errors among high-mortality observations.
Figure 12. Residual Diagnostics of the K-Nearest Neighbors Regression Model for Respiratory Mortality Prediction. Note: The figure presents residuals against predicted TRD values for the KNN model. Residuals are generally centered around zero without systematic patterns, supporting predictive adequacy while highlighting larger forecasting errors among high-mortality observations.
Urbansci 10 00351 g012
Figure 13. Cross-Validation Stability of KNN Regression Across Alternative Neighborhood Sizes. Note: (A) shows the root-mean-square errors (RMSEs) from the 10-fold cross-validation analysis for different neighborhood sizes. Low RMSE values indicate high predictive power. Predictive accuracy is maximal when k = 2 because highly localized neighborhood structures give the best predictions of respiratory mortality. The corresponding R2s from the cross-validation are shown in (B), where a higher value of R2 indicates more predictive power and model explanatory power. Predictions of respiratory mortality are most accurate and best explain environmental risk factors when k = 2; thereafter, the value of R2 decreases as the number of neighbors increases. Dashed vertical lines show the optimal neighborhood size obtained from the cross-validation, and the dotted vertical lines show k ≈ √n (around 15 in this case). From these figures, one can see a clear deterioration in prediction as the number of neighbors increases, indicating that respiratory mortality exhibits a strong locally similar structure.
Figure 13. Cross-Validation Stability of KNN Regression Across Alternative Neighborhood Sizes. Note: (A) shows the root-mean-square errors (RMSEs) from the 10-fold cross-validation analysis for different neighborhood sizes. Low RMSE values indicate high predictive power. Predictive accuracy is maximal when k = 2 because highly localized neighborhood structures give the best predictions of respiratory mortality. The corresponding R2s from the cross-validation are shown in (B), where a higher value of R2 indicates more predictive power and model explanatory power. Predictions of respiratory mortality are most accurate and best explain environmental risk factors when k = 2; thereafter, the value of R2 decreases as the number of neighbors increases. Dashed vertical lines show the optimal neighborhood size obtained from the cross-validation, and the dotted vertical lines show k ≈ √n (around 15 in this case). From these figures, one can see a clear deterioration in prediction as the number of neighbors increases, indicating that respiratory mortality exhibits a strong locally similar structure.
Urbansci 10 00351 g013
Table 1. Key Environmental and Infrastructural Predictors of Respiratory Disease Mortality (TRD).
Table 1. Key Environmental and Infrastructural Predictors of Respiratory Disease Mortality (TRD).
AcronymVariableRevised Definition
TRDMortality Rate by Respiratory DiseaseAge-standardized mortality rate from respiratory diseases, expressed as deaths per 100,000 population.
ELECAccess to ElectricityPercentage of the population with access to electricity services; % of population.
AGRLAgricultural LandAgricultural land as a percentage of total land area; % of land area.
WTRWFreshwater WithdrawalsAnnual freshwater withdrawals as a percentage of total renewable internal freshwater resources; %.
CDDsCooling Degree DaysAnnual cooling degree days, measured as degree-days above a standard baseline temperature of 18 °C.
COALCoal ElectricityShare of electricity generated from coal sources; % of total electricity generation.
SANSSafe SanitationPercentage of the population using safely managed sanitation services; % of population.
RENERenewable EnergyRenewable energy consumption as a percentage of total final energy consumption; %.
Note. Source: Institute for Health Metrics and Evaluation, Link: https://www.healthdata.org/ (accessed on 10 January 2025), World Bank, Sovereign ESG Data Portal, https://databank.worldbank.org/source/world-development-indicators, accessed on 12 January 2025.
Table 2. Variance Inflation Factor Diagnostics for Baseline Regressors.
Table 2. Variance Inflation Factor Diagnostics for Baseline Regressors.
VariableVIFTolerance (1/VIF)
WTRW2.010.498265
CDDs1.920.521613
RENE1.690.591785
AGRL1.540.649988
COAL1.380.722825
SANS1.380.724796
ELEC1.080.928494
Mean VIF1.57
Note: VIF diagnostics were computed after the baseline OLS specification. All values remain below conventional multicollinearity thresholds, indicating that collinearity does not substantially bias the interpretation of coefficients in the panel regression framework.
Table 3. Comparison of Random-Effects GLS and Fixed-Effects Panel Estimates for TRD.
Table 3. Comparison of Random-Effects GLS and Fixed-Effects Panel Estimates for TRD.
Random-Effects (GLS), Using 238 Observations
Included 38 Cross-Sectional Units
Time-Series Length: Minimum 6, Maximum 11
Dependent Variable: TRD
Fixed-Effects, Using 238 Observations
Included 38 Cross-Sectional Units
Time-Series Length: Minimum 6, Maximum 11
Dependent Variable: TRD
CoefficientStd. ErrorzCoefficientStd. Errort-ratio
Constant71.84 ***26.522.7068.52 **27.472.49
ELEC−0.81 ***0.25−3.23−0.81 ***0.25−3.21
AGRL0.51 ***0.105.060.59 ***0.134.48
WTRW−0.10 ***0.039−2.74−0.13 ***0.04−3.00
CDDs0.004 ***0.0013.690.004 ***0.0013.40
COAL0.114 ***0.042.750.12 ***0.042.70
SANS0.26 ***0.054.920.26 ***0.054.50
RENE0.24 ***0.073.390.23 ***0.072.95
StatisticsMean dependent var38.9038.9
Sum squared resid60652.10766.63
Log-likelihood−997.04−476.90
Schwarz criterion2037.861200.06
rho0.500.500
S.D. dependent var16.7616.76
S.E. of regression16.201.99
Akaike criterion2010.081043.81
Hannan–Quinn2021.281106.78
Durbin–Watson0.70.74
Tests‘Between’ variance = 261.99
‘Within’ variance = 3.97
Mean theta = 0.95
Joint test on named regressors -
Asymptotic test statistic: Chi-square(7) = 109.07
with p-value = 1.42877 × 10−20
Joint test on named regressors -
Test statistic: F(7, 193) = 15.32
with p-value = P(F(7, 193) > 15.3211) = 6.93623 × 10−16
Breusch–Pagan test –
Null hypothesis: Variance of the unit-specific error = 0
Asymptotic test statistic: Chi-square(1) = 580.148
with p-value = 3.48302 × 10−128
Test for differing group intercepts -
Null hypothesis: The groups have a common intercept
Test statistic: F(37, 193) = 341.847
with p-value = P(F(37, 193) > 341.847) = 1.59215 × 10−156
Hausman test -
Null hypothesis: GLS estimates are consistent
Asymptotic test statistic: Chi-square(7) = 7.92934
with p-value = 0.338866
Note: Asterisks indicate statistical significance levels: *** p < 0.01; ** p < 0.05.
Table 4. Fixed-Effects Regression with Driscoll–Kraay Standard Errors for TRD (2010–2021).
Table 4. Fixed-Effects Regression with Driscoll–Kraay Standard Errors for TRD (2010–2021).
ItemValueItemValue
Regression typeFixed-effects regressionStandard errorsDriscoll–Kraay
Number of observations238Number of groups38
Group variable (i)nMaximum lag2
F(17,10)62,298.16Prob > F0
Within R-squared0.39
VariableCoefficientStd. Err.tp > |t|95% Conf. Interval (Lower)95% Conf. Interval (Upper)
ELEC−1.270.591−2.160.056−2.590.042
AGRL0.630.163.750.0040.251.01
WTRW−0.110.05−1.930.082−0.230.01
CDD0.0030.0021.590.143−0.0010.008
COAL0.110.042.470.0330.010.21
SANS0.170.035.1300.090.25
RENE0.120.101.20.259−0.100.35
20100
20110.450.085.0600.250.65
20120.420.301.390.193−0.251.09
20130.940.185.030.0010.521.36
20140.750.302.450.0350.061.43
20151.470.562.610.0260.212.72
20161.970.375.3201.142.80
20172.720.456.0101.713.73
20183.490.536.5602.304.68
20194.110.695.9402.565.65
20203.840.745.1702.185.5
20210
_cons121.0862.921.920.083−19.12261.29
Table 5. Sensitivity Analyses for Baseline Fixed-Effects Specification.
Table 5. Sensitivity Analyses for Baseline Fixed-Effects Specification.
ModelSANSRENEObservationsGroupsInterpretation
Baseline FE0.269 ***0.235 ***23838Both variables remain positive and statistically significant.
Without SANS-0.360 ***23838RENE remains positive and robust after excluding sanitation.
Without RENE0.330 ***-23838SANS remains positive and statistically significant.
Without SANS and RENE--23838Core environmental controls remain generally stable.
Lagged regressors0.254 ***0.15620038SANS remains significant; RENE loses significance under temporal lag structure.
Note: Fixed-effects sensitivity analyses use the baseline sample of 238 observations and 38 groups, except for the lagged-regressor model, which includes 200 observations because first-year panel observations are lost after lagging. Asterisks indicate statistical significance: *** p < 0.01.
Table 6. Fixed-Effects Estimates by European Regional Context.
Table 6. Fixed-Effects Estimates by European Regional Context.
VariableWestern EuropeEastern Europe
ELEC−8.816−1.006 ***
AGRL1.151 ***0.052
WTRW−0.116 **−0.233 ***
CDDs0.004 **0.004 **
COAL0.250 ***−0.013
SANS0.1340.324 ***
RENE0.519 ***0.043
Constant859.511105.964 ***
Observations120118
Groups2018
Within R-squared0.4780.447
F-statistic12.18 ***10.76 ***
Note: Fixed-effects regressions were estimated separately for Western and Eastern Europe using the baseline sample. Asterisks indicate statistical significance levels: *** p < 0.01, ** p < 0.05. Within R-squared values refer to within-panel explanatory power.
Table 7. Fixed-Effects Panel Estimates for Respiratory Disease Mortality: Baseline and Extended Specifications.
Table 7. Fixed-Effects Panel Estimates for Respiratory Disease Mortality: Baseline and Extended Specifications.
VariableBaseline FE CoefficientBaseline FE Std. ErrorExtended FE CoefficientExtended FE Std. Error
ELEC−1.276 ***0.314−1.196 ***0.295
AGRL0.636 ***0.1350.648 ***0.124
WTRW−0.110 **0.049−0.128 ***0.046
CDDs0.004 **0.0010.003 *0.001
COAL0.112 **0.0460.085 *0.044
SANS0.175 **0.0680.0740.065
RENE0.1250.104−0.0030.098
POP652.617 ***0.562
HEALTH−0.001 *0.001
GDPPC−0.00014 **0.00007
PM25−0.0020.147
Year fixed effectsYesYes
Country fixed effectsYesYes
Observations238237
Countries3838
Within R20.39090.5053
Between R20.09780.1959
Overall R20.09460.1897
F-statistic6.91 ***8.66 ***
Prob > F00
Sigma_u17.04116.897
Sigma_e1.9921.807
Rho0.9870.989
Note: The dependent variable is TRD, respiratory disease mortality. Standard errors are reported in parentheses. Both specifications include country and year fixed effects. The extended specification adds controls for population aging (POP65), healthcare-system development or health expenditure (HEALTH), GDP per capita (GDPPC), and particulate exposure (PM25). Statistical significance: *** p < 0.01, ** p < 0.05, * p < 0.10.
Table 8. Conceptual Interpretation of Country-Level Environmental and Infrastructure Indicators.
Table 8. Conceptual Interpretation of Country-Level Environmental and Infrastructure Indicators.
VariableDirect InterpretationAlternative Country-Level InterpretationImplications for Interpretation
ELEC (Access to Electricity)Population access to electricity servicesProxy for economic development, infrastructure quality, housing conditions, institutional capacity, and access to health-related servicesThe negative coefficient should not be interpreted exclusively as a direct effect of electrification. It may reflect broader socioeconomic and infrastructural advantages associated with more developed countries.
WTRW (Freshwater Withdrawals)Intensity of freshwater withdrawals relative to renewable resourcesProxy for water-resource availability, infrastructure capacity, environmental governance, and resource-management capabilityThe negative coefficient does not imply that higher water withdrawals directly reduce mortality. It may capture broader differences in resource availability and infrastructure resilience.
AGRL (Agricultural Land)Share of land devoted to agricultural activitiesProxy for agricultural intensity, land-use patterns, rural environmental exposure, pesticide use, dust emissions, and biomass burningThe positive coefficient should not be interpreted as evidence that agriculture directly causes respiratory mortality. Rather, it may reflect environmental exposures associated with intensive agricultural systems.
SANS (Safe Sanitation Access)Population access to safely managed sanitation servicesProxy for development gradients, infrastructure-transition processes, regional heterogeneity, and broader socioeconomic conditionsThe positive coefficient observed in the baseline model does not imply that sanitation worsens health outcomes. The loss of significance in the extended specification suggests that the relationship is likely driven by broader contextual factors.
RENE (Renewable Energy Consumption)Share of renewable energy in final energy consumptionProxy for energy-transition processes, institutional change, development pathways, and regional heterogeneityThe coefficient may reflect transitional dynamics rather than the direct health effects of renewable energy itself.
CDDs (Cooling Degree Days)Climatic demand for cooling above 18 °CIndicator of climatic stress, heat exposure, and environmental vulnerabilityThe positive coefficient is consistent with increased respiratory-health vulnerability associated with rising temperatures and heat-related environmental pressures.
COAL (Coal-Based Electricity Generation)Share of electricity generated from coalIndicator of fossil-fuel dependence, pollution exposure, and carbon-intensive energy systemsThe positive coefficient is consistent with higher exposure to pollutants and environmental stressors associated with coal-intensive energy structures.
Note: The variables included in the analysis are measured at the country–year level and should not be interpreted as isolated causal mechanisms. Several indicators may act as proxies for broader environmental, infrastructural, socioeconomic, institutional, and development-related conditions. Consequently, the estimated coefficients should be interpreted as associations reflecting complex structural processes rather than direct causal effects.
Table 9. Cluster Validity Indices for Competing Algorithms in Environmental–Health Data.
Table 9. Cluster Validity Indices for Competing Algorithms in Environmental–Health Data.
RankModelMaximum DiameterMinimum SeparationPearson’sPearson’s γDunn IndexEntropy
1Density Based0.211.00.621.01.00.0
2Hierarchical0.870.261.00.440.120.95
3Neighborhood Based1.00.140.860.290.051.0
4Random Forest0.610.370.290.510.00.64
5Model Based0.390.050.450.060.060.52
6Fuzzy C-Means0.00.00.00.00.140.18
Table 10. Comparative Evaluation of Hierarchical and Density-Based Clustering Solutions for Environmental-Health Profiles.
Table 10. Comparative Evaluation of Hierarchical and Density-Based Clustering Solutions for Environmental-Health Profiles.
MetricHierarchical ClusteringDensity-Based ClusteringComparative Interpretation
Number of clusters104 (+1 noisepoint)Hierarchical clustering provides greater segmentation and finer differentiation across environmental-health profiles.
Cluster-size distribution, country–year observations11, 69, 8, 60, 18, 12, 12, 24, 18, 6219, 6, 6, 6Density-based clustering is highly polarized, with one dominant cluster containing almost all observations.
Minimum cluster size, country–year observations66Similar minimum size, but hierarchical clustering distributes observations more evenly.
Maximum cluster size, country–year observations69219Density-based clustering concentrates observations in a single oversized cluster.
Average cluster size, country–year observations23.859.3Hierarchical clustering achieves a more balanced partition structure.
Explained proportion within-cluster heterogeneityMaximum = 0.381Maximum = 0.998Density-based clustering captures variance mainly through one dominant cluster, reducing representativeness.
Within sum of squares (largest cluster)180.2831474.5The dominant density-based cluster exhibits extremely high internal concentration.
Silhouette score range0.203–0.8080.165–0.853Both methods show positive cohesion, but hierarchical clustering maintains cohesion across a larger number of balanced groups.
Clusters with silhouette > 0.6Clusters 1, 9, and 10Clusters 2, 3, and 4Hierarchical clustering achieves strong cohesion without excessive concentration.
Between sum of squares1422.22397.8Hierarchical clustering explains substantially more between-group variation.
Total sum of squares18961875.56Both methods operate on comparable total variance structures.
InterpretabilityHighLimitedHierarchical clustering provides more interpretable and policy-relevant environmental-health profiles.
Final selectionSelectedNot selectedHierarchical clustering was preferred because it balances cohesion, heterogeneity, and cluster distribution.
Note: The table compares clustering performance across country–year observations, showing that hierarchical clustering provides more balanced, interpretable, and policy-relevant environmental-health profiles than density-based clustering, which produces one dominant cluster.
Table 11. Comparative Assessment of Hierarchical Clustering Solutions from k = 2 to k = 10.
Table 11. Comparative Assessment of Hierarchical Clustering Solutions from k = 2 to k = 10.
kR2BICSilhouetteBetween SSCluster Size DistributionMain Interpretation
20.161679.50.361304.058220/18Extremely polarized structure dominated by one cluster
30.2731509.540.293517.799200/20/18Slight improvement, but still highly unbalanced
40.3121479.350.285591.761200/20/12/6Persistent concentration in one dominant cluster
50.4181322.650.261792.245182/20/18/12/6Improved segmentation, but imbalance remains substantial
60.531153.540.3111005.127153/29/20/18/12/6Better differentiation, though one oversized cluster persists
70.5711119.950.3111082.5153/20/18/18/12/11/6Increased separation, but still structurally polarized
80.68957.60.341288.62593/60/20/18/18/12/11/6Markedly improved balance and heterogeneity representation
90.696970.880.3421319.12693/60/18/18/12/12/11/8/6Further differentiation, with nine distinct clusters and improved representation of environmental-health heterogeneity
100.75911.560.3751422.21969/60/24/18/18/12/12/11/8/6Best compromise between cohesion, separation, and interpretability
Note: The table compares alternative hierarchical clustering solutions, showing that k = 10 provides the best balance between explained variance, BIC, silhouette score, between-cluster separation, cluster-size distribution, and interpretability.
Table 12. Structural relationships between environmental determinants and respiratory disease mortality across European clusters.
Table 12. Structural relationships between environmental determinants and respiratory disease mortality across European clusters.
ClusterTRDELECAGRLWTRWCDDCOALSANSRENE
10.1240.404−0.9640.2261.514−1.713−0.673−0.562
2−0.330−0.456−0.6270.2610.0930.331−0.686−0.510
3−0.109−0.401−0.886−3.204−0.945−0.107−1.183−0.557
40.8460.0300.0780.316−0.3860.2291.3390.370
50.1390.1701.8750.0710.091−2.512−0.2270.017
60.7310.2350.739−2.8600.034−0.628−0.197−0.026
7−1.1082.572−0.9640.320−0.9880.327−0.1231.042
80.702−0.5851.2970.320−0.6350.821−0.334−0.177
9−2.075−0.835−0.7650.3202.1650.4850.348−0.624
10−0.9963.4591.3160.320−1.0470.589−1.0744.436
Table 13. Machine-Learning Training, Validation, and Tuning Framework.
Table 13. Machine-Learning Training, Validation, and Tuning Framework.
ComponentSpecification
Dataset size238 country–year observations
PredictorsELEC, AGRL, WTRW, CDD, COAL, SANS, RENE
Target variableTRD
Training sample64%
Validation sample16%
Holdout test sample20%
Feature scalingEnabled
Distance metric (KNN)Euclidean
Weighting schemeRectangular
Hyperparameter tuningAutomatic neighbor selection
Candidate k values1–10 (JASP optimization)
Additional robustness analysis10-fold cross-validation
Performance metricsMSE, RMSE, MAE/MAD, MAPE, R2
Interpretation toolLocal additive explanations
Note: Table 13 provides a comprehensive overview of the machine-learning training, validation, tuning, and interpretation procedures adopted in this study. The workflow was designed to ensure transparency and comparability across predictive models. Feature scaling was applied before estimation, particularly because distance-based methods such as KNN are sensitive to the measurement scale of predictors. The holdout test set was used to assess final predictive performance, while the validation sample and additional 10-fold cross-validation were used to evaluate model tuning and stability. The reported configuration should therefore be interpreted as a transparent predictive framework rather than as a causal or inferential modelling strategy.
Table 14. Absolute and Normalized Predictive Performance Metrics for Machine Learning Models.
Table 14. Absolute and Normalized Predictive Performance Metrics for Machine Learning Models.
ModelMSERMSEMAE/MADMAPE (%)R2
K-Nearest Neighbors (KNN)10.031 (1.000)3.167 (1.000)1.790 (1.000)4.22 (1.000)0.976 (1.000)
Random Forest41.602 (0.894)6.450 (0.820)4.771 (0.750)14.32 (0.720)0.894 (0.916)
Decision Tree116.619 (0.730)10.799 (0.530)7.506 (0.570)18.59 (0.500)0.635 (0.651)
Boosting Regression148.382 (0.500)12.181 (0.330)9.700 (0.360)30.49 (0.200)0.509 (0.521)
Regularized Linear Regression (Lasso)231.449 (0.300)15.213 (0.180)11.839 (0.180)32.70 (0.070)0.167 (0.171)
Linear Regression299.124 (0.040)17.295 (0.020)14.422 (0.030)47.51 (0.160)0.158 (0.162)
Support Vector Machine (SVM)302.732 (0.300)17.399 (0.180)14.222 (0.200)48.45 (0.000)0.081 (0.083)
Note: Values outside parentheses report the absolute performance metrics obtained on the test data, whereas values in parentheses report the normalized scores used for relative comparison across models. Lower values of MSE, RMSE, MAE/MAD, and MAPE indicate better predictive performance, while higher R2 values indicate greater predictive accuracy. The normalized scores are retained solely to facilitate model comparison and should not be interpreted as measures of substantive predictive quality. Therefore, the absolute metrics provide the primary basis for evaluating model performance and interpreting prediction errors in respiratory mortality.
Table 15. Additive Feature Attributions and Mean Dropout Loss for KNN Predictions of Respiratory Disease Mortality (TRD) Across Case Profiles.
Table 15. Additive Feature Attributions and Mean Dropout Loss for KNN Predictions of Respiratory Disease Mortality (TRD) Across Case Profiles.
VariableMean Dropout LossCase 1Case 2Case 3Case 4Case 5
Predicted TRD25.69535.45037.13516.99516.380
Base TRD38.92138.92138.92138.92138.921
ELEC5.84−1.1820.8900.907−8.8000.039
AGRL13.41−3.982−5.386−5.848−0.940−0.522
WTRW11.63−3.652−4.062−3.999−0.934−4.229
CDD11.555.869−3.529−3.571−1.121−4.087
COAL7.968−3.190−3.645−3.333−6.175−1.192
SANS9.176−6.7397.2859.5590.528−7.829
RENE13.111−0.3514.9774.499−4.485−4.721
Table 16. KNN Training, Validation, and Hyperparameter Configuration.
Table 16. KNN Training, Validation, and Hyperparameter Configuration.
ComponentSpecification
Total observations238 country–year observations
Training sample152 observations (64%)
Validation sample39 observations (16%)
Holdout test sample47 observations (20%)
Validation strategyInternal validation set
Optimization criterionValidation Mean Squared Error (MSE)
Feature scalingEnabled (standardization applied before estimation)
Distance metricEuclidean distance
Weighting schemeRectangular weights
Neighbor selectionAutomatic optimization
Maximum number of neighbors considered10
Selected number of neighbors (k)1
Validation MSE110.192
Test MSE10.031
RMSE3.167
MAE/MAD1.79
MAPE4.22%
R20.976
Note: The KNN model was estimated using scaled predictors, Euclidean distance, and rectangular weights. Twenty percent of the observations were reserved as holdout test data, while 20% of the remaining training observations were used for validation. The number of neighbors was selected automatically by minimizing validation-set MSE among candidate neighborhood sizes up to k = 10.
Table 17. KNN Stability Analysis Across Alternative Neighborhood Sizes.
Table 17. KNN Stability Analysis Across Alternative Neighborhood Sizes.
Number of Neighbors (k)RMSER2MAERMSE SDR2 SDMAE SD
12.2180.9831.4830.390.010.283
22.0670.9861.3420.3770.0060.258
32.1830.9841.4040.530.010.328
53.3410.9632.3830.7970.0190.562
75.9070.8964.7151.3250.051.084
108.7820.7567.271.6120.0861.261
1511.2680.5719.1342.2290.1571.678
2012.5370.46610.0612.170.1581.76
2513.2940.39410.6192.1590.1631.842
Note: Performance metrics were obtained using 10-fold cross-validation. Predictors were centered and scaled before KNN estimation. Lower RMSE and MAE values indicate better predictive performance, while higher R2 values indicate greater explanatory accuracy.
Table 18. Integrated Interpretation of Panel Regression, Hierarchical Clustering, and KNN Evidence.
Table 18. Integrated Interpretation of Panel Regression, Hierarchical Clustering, and KNN Evidence.
Evidence DimensionPanel RegressionHierarchical ClusteringKNN RegressionIntegrated Interpretation
Main analytical roleInferential estimation of average associationsUnsupervised identification of country–year profilesPredictive modelling of TRDThe three methods provide complementary evidence, not identical confirmation
Question addressedWhich variables are associated with TRD?Which environmental-health profiles emerge from the data?Which variables improve prediction of TRD?Each method answers a different empirical question
AGRLPositive association with TRDContributes to differentiated high-risk profilesHigh predictive relevanceStrong convergence: agricultural intensity is risk-enhancing
COALPositive association with TRDCharacterizes structurally vulnerable profilesRelevant predictorDirectional convergence: coal dependence contributes to respiratory risk
CDDsPositive association with TRDDifferentiate climate-stressed profilesHigh predictive relevanceStrong convergence: heat stress is an important risk factor
ELECNegative association with TRDAssociated with lower-mortality profilesContext-dependent local contributionConsistent evidence of protective infrastructural relevance
WTRWNegative association with TRDVaries across cluster profilesHigh predictive relevance, locally variableContext-dependent effect, likely linked to regional and infrastructural heterogeneity
SANSPositive but cautiously interpretedMixed across profilesLocally heterogeneous contributionContext-dependent, possibly reflecting transitional infrastructure
RENEPositive or unstable across specificationsMixed across profilesHigh predictive relevance but locally variableContext-dependent, likely reflecting energy-transition heterogeneity
Interpretation ruleInferential evidenceHeterogeneity evidencePredictive evidenceStrong conclusions are drawn only where methods are directionally consistent
Note: The table summarizes how econometric, clustering, and predictive evidence complement one another, distinguishing inferential associations, country–year heterogeneity profiles, and machine-learning prediction while identifying convergent and context-dependent environmental-health patterns.
Table 19. Urban-Science Interpretation of Environmental-Health Findings and Policy Implications.
Table 19. Urban-Science Interpretation of Environmental-Health Findings and Policy Implications.
Empirical FindingUrban-Science InterpretationPlanning and Public-Health Implication
Positive association between CDDs and respiratory mortalityHeat stress and urban climatic exposureStrengthen heat-adaptation plans, urban cooling strategies, green infrastructure, and protection of vulnerable groups
Positive association between COAL and respiratory mortalityCarbon-intensive energy systems and urban air-quality vulnerabilityAccelerate clean-energy transitions, reduce fossil-fuel dependence, and integrate air-quality goals into urban energy planning
Negative association between ELEC and respiratory mortalityInfrastructure access and service reliabilityImprove resilient access to essential urban services, especially in vulnerable peri-urban and less connected areas
Negative association between WTRW and respiratory mortalityWater-resource availability and infrastructure resilienceIntegrate water security, climate adaptation, and public-health planning
Heterogeneous effects of SANS and RENETransitional infrastructure and uneven development pathwaysInterpret sanitation and renewable-energy indicators in relation to regional context, infrastructure quality, and implementation stage
Hierarchical clustering identifies differentiated country–year profilesTerritorial heterogeneity in environmental-health vulnerabilityAvoid one-size-fits-all policies; design region-sensitive urban and regional planning strategies
KNN identifies strong predictive relevance of environmental-health variablesNon-linear and local vulnerability structuresUse predictive tools as complementary support for targeting environmental-health interventions
Note: The table translates the empirical findings into an urban-science framework. It does not imply that the analysis directly measures intra-urban morphology or city-level inequalities; rather, it identifies structural determinants that condition urban and regional environmental-health vulnerability.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Resta, E.; Resta, O.; Liuzzi, P.; Costantiello, A.; Leogrande, A. Environmental-Health Vulnerability and Respiratory Mortality in Europe: Evidence from Panel Econometrics, Clustering, and Machine Learning. Urban Sci. 2026, 10, 351. https://doi.org/10.3390/urbansci10070351

AMA Style

Resta E, Resta O, Liuzzi P, Costantiello A, Leogrande A. Environmental-Health Vulnerability and Respiratory Mortality in Europe: Evidence from Panel Econometrics, Clustering, and Machine Learning. Urban Science. 2026; 10(7):351. https://doi.org/10.3390/urbansci10070351

Chicago/Turabian Style

Resta, Emanuela, Onofrio Resta, Piergiuseppe Liuzzi, Alberto Costantiello, and Angelo Leogrande. 2026. "Environmental-Health Vulnerability and Respiratory Mortality in Europe: Evidence from Panel Econometrics, Clustering, and Machine Learning" Urban Science 10, no. 7: 351. https://doi.org/10.3390/urbansci10070351

APA Style

Resta, E., Resta, O., Liuzzi, P., Costantiello, A., & Leogrande, A. (2026). Environmental-Health Vulnerability and Respiratory Mortality in Europe: Evidence from Panel Econometrics, Clustering, and Machine Learning. Urban Science, 10(7), 351. https://doi.org/10.3390/urbansci10070351

Article Metrics

Back to TopTop