4.1. Results of the Empirical Analysis
Correlation analysis was conducted to assess the strength and direction of pairwise linear relationships between variables and to identify potential multicollinearity before constructing the panel model. The results of the analysis are presented as a heat map in
Figure 2.
Analysis of the matrix suggests that the Pearson correlation coefficients indicate a strong positive relationship between variable y (RIC, investment per capita) and factor X7 (0.63) (Interaction Industry & Budget). This suggests that, at this stage, of all the regressors, X7 appears to be the most promising factor for explaining Y. X1 (Regional budget per capita) and X3 (Industry’s share of GVA) also have a positive relationship with Y, but a weaker one: 0.32 and 0.29, respectively. This means they may be useful in the model, but their influence initially appears less pronounced than that of X7. X4 (Share of agricultural production in GVA) shows a moderate negative relationship with Y (−0.34). For X2, X5, X6, X8, and X9, the relationship with Y is weak or virtually nonexistent.
At the same time, relatively high absolute correlation coefficients were recorded between several independent variables, particularly for the pairs X5–X9 (r = 0.78), X2–X8 (r = 0.75), X3–X6 (r = 0.66), and X3–X9 (r = −0.63). This indicates possible multicollinearity, so the next step requires calculating the variance inflation factors (VIF). Their calculation revealed that VIF values range from 2.18 to 7.31. The highest values were recorded for variables X3 (7.31), X9 (6.96), and X4 (5.63), indicating moderate multicollinearity. However, none of the variables exceeded the critical threshold of 10, so there is no basis for automatically eliminating factors at this stage. Consequently, the set of regressors can be retained for subsequent estimation of the panel model.
Next, we estimated pooled OLS, fixed-effects, and random-effects models. Pooled OLS results demonstrated high explanatory power, but the pooling test rejected the hypothesis of homogeneity of panel units, ruling out the pooled regression as the final specification (
Table 3).
In the fixed effects and random effects models, variables X2 (Regional subventions per capita), X3, and X4 (with negative signs), as well as X7 (Interaction Industry & Budget) and X8 (Interaction Agro & Budget) (with positive signs) were consistently significant. Variable X5 (Business density) was statistically significant only in the fixed effects model, while X1 and X6, significant in the pooled OLS, lost significance after accounting for individual regional effects.
To determine the final choice between the fixed and random effects models, the Hausman test was used. The resulting statistical value was 84.808 with 9 degrees of freedom, p-value = 1.78 × 10−14. Since the significance level was significantly lower than 0.05, the null hypothesis of the random effects model is rejected. This indicates a correlation between the individual regional effects and the explanatory variables. Therefore, a fixed effects model is preferable for further analysis.
According to the fixed effects model, factors X2, X3, X4, X5, X7, and X8 have a statistically significant effect on the dependent variable. The corresponding estimated equation is:
To verify the correctness of the error structure in the selected fixed-effects model, the Breusch–Pagan test for heteroscedasticity and the Wooldridge test for first-order autocorrelation in panel data were conducted. The results of the Breusch–Pagan test revealed a statistically significant deviation from the homoscedasticity assumption (LM = 71.812, p < 0.001; F = 9.846, p < 0.001), indicating heterogeneity in the residual variance. Concurrently, the Wooldridge test revealed first-order autocorrelation (t = −4.949, p < 0.001), indicating a time-dependent dependence of the model errors within panel units.
Therefore, using standard errors for the final interpretation of the coefficients is inappropriate. Given the identified heteroscedasticity and autocorrelation, further testing for possible endogeneity of the key regressors was performed using the control function procedure within a panel model with fixed effects and clustered standard errors (
Table 4). X2, X3, X5, X7, and X8 were considered as potentially endogenous variables. In the first stage, auxiliary equations were estimated for each of them using first-order lags as endpoints. In the second stage, the residuals from the first-stage equations were included in the structural equation. The results revealed no statistically significant endogeneity for X2, X3, X5, and X8. For X7, only a weak signal was detected at the 10% significance level. The joint test failed to reject the null hypothesis of exogeneity for this block of variables, suggesting that the final model is generally robust to the endogeneity problem.
The results of the first stage showed that the lags of potentially endogenous variables possess high explanatory power and statistical significance. For all regressors considered, the F-statistic values significantly exceed the thresholds used to diagnose weak instruments, and the p-values are close to zero. This suggests the high relevance of the selected lagged instruments. Therefore, further endogeneity testing using the control function is methodologically justified.
The data in
Table 5 show that no significant endogeneity issue was identified in the model. The only variable that raises caution is X7, as it shows a weak signal at the 10% level. However, overall, the set of key regressors can be considered empirically acceptable for this test.
According to the results of the final fixed-effects model with clustered standard errors, factors X2, X3, X5, X7, and X8 have a statistically significant effect on the dependent variable. Variables X2 and X3 are negative, indicating an inverse relationship with the dependent variable, while X5, X7, and X8 have a positive effect. Variable X4 demonstrates borderline significance at the 10% level and can therefore only be considered a weak additional factor. Variables X1, X6, and X9 did not show a statistically significant effect.
The estimated model equation is as follows:
The resulting model indicates that the level of development of the business environment in the regions, specifically the number of active companies per 1000 people (X5), has a positive impact on regional investment growth. The positive and significant impact of both coefficients for variables X7 (Interaction Industry & Budget) and X8 (Interaction Agro & Budget) indicates that budget funds are more effective in industrially and agriculturally developed regions. In this case, this refers to government spending that is targeted to support specific programs and projects for industrial and agricultural development. The negative effect of subventions indicates that their growth reduces per capita investment, which weakens investment incentives for regional business development. The negative effect of the share of industry most likely indicates saturation and concentration in already developed industrial regions.
Thus, the results demonstrate that regional investment allocation is shaped by opposing structural and fiscal factors. The model shows that subsidies and the share of industry are negatively correlated with investment, reinforcing divergence. This also points to fiscal dependence and structural concentration. The positive moderating effect of the budget-industry interaction indicators suggests that government spending is only effective in regions with a developed industrial and agricultural base. Consequently, investment dynamics follow a pattern of regional differentiation, where growth accelerates in structurally developed regions, leading to spatial concentration rather than uniform regional development.
4.2. A Comparative Analysis of Investment Concentration and Regional Inequality in Kazakhstan
The analysis reveals persistent differences between Kazakhstan’s regions due to structural characteristics. Indicators based on development level (investment and GRP per capita, average wage) show discrepancies across the full sample, largely due to the influence of the leading region, Atyrau Oblast, which significantly increases the overall gap (
Table 6).
Log difference analysis reveals persistent divergence in per capita investment driven by the dominance of resource-intensive regions. Atyrau Region consistently acts as the leading benchmark, creating a consistent gap compared to other regions. Although some industrial regions (Pavlodar, Karaganda) demonstrate partial convergence, overall inequality remains structurally determined. The Mangystau Region (with its developed oil and gas industry) and Astana (the capital and investment center) are closest to the leader. However, in Astana, the main influx of investment goes to the service sector; industry represents only about 10% of the region’s GVA. Furthermore, investment activity in Astana is largely driven by the capital’s administrative status and the concentration of state and quasi-state projects. Unlike Atyrau, investment in Astana is less associated with export-oriented sectors and does not generate a similar effect of spatial concentration. This limits its comparability with resource-intensive regions.
The coefficient of variation confirms that the variance is largely determined by outliers (
Table 7). Excluding Atyrau and Mangystau from the calculations significantly reduces inequality, indicating that most regions are following a more homogeneous trajectory. Management decisions that can be drawn from this analysis indicate that fiscal redistribution alone (as demonstrated by the model and the role of interaction indicators) is insufficient. Instead, emphasis should be placed on industrial diversification, increasing business density, and supporting the structural transformation of the regional economy. Analysis of similar comparisons by GRP per capita revealed the following conclusions (
Table 8).
Empirical results reveal persistent and structurally driven regional disparities in Kazakhstan. Log difference analysis reveals that Atyrau Oblast consistently acts as the dominant benchmark, while southern regions (Turkestan and Zhambyl Oblasts) remain structurally lagging, with log deviations exceeding −2. Meanwhile, a number of industrial regions (e.g., Pavlodar Oblast) demonstrate partial convergence, suggesting the role of industrial potential in narrowing disparities. The data for Almaty (low log difference) are explained by the growth of regional GRP due to the expansion of the service sector in the financial and investment hub.
The coefficient of variation confirms the dual structure of inequality (
Table 9). Including the leaders, the dispersion remains high (CV approximately 0.65–0.75), reflecting resource-driven polarization. Without the leaders, the dispersion decreases significantly (CV approximately 0.27), indicating relative convergence within the main group of regions. This suggests that inequality in Kazakhstan is not uniform but is driven by extreme deviations, primarily in oil-producing regions. The results support the hypothesis that regional inequality is shaped by the concentration of investment in extractive industries, uneven structural transformation, and limited spillover effects between regions. An analysis of similar comparisons for average wages (
Appendix A) revealed that comparing the log-difference and the coefficient of variation for all three indicators yields consistently similar results. Specifically, regional inequality in Kazakhstan is structural in nature and is reproduced through the concentration of resources in a limited number of regions. This confirms the hypotheses that investment allocation acts as a factor of divergence and is spatially concentrated in resource-oriented economies, while similar dynamics are observed in both household income and output.
4.3. Cross-Country Data on Investment Distribution and Regional Differences
The logarithmic difference results indicate a persistent but gradually narrowing gap between the selected countries and the leader. All countries show partial convergence toward Norway, particularly Lithuania and Estonia. Kazakhstan demonstrates a slower convergence and even a temporary widening of the gap. This points to structural constraints and lower investment transformation efficiency compared to European countries. The logarithmic difference results confirm that the relative distance from the leader remains significant, both across countries and across Kazakhstan’s regions, especially for income and investment indicators, indicating limited convergence in the medium term.
The coefficient of variation confirms that cross-country inequality depends significantly on the leading economy (
Table 11). Including Norway leads to consistently higher variance, although the downward trend through 2024 suggests partial convergence. Excluding Norway leads to significantly lower and more stable variance values, indicating relative homogeneity in development among the remaining countries. This supports the hypothesis that inequality has structural causes and is exacerbated by highly developed countries (regions), while countries (regions) with more similar economic conditions develop along comparable trajectories.
The structural transformation assessment focuses on analyzing the differences and structural patterns in the development of the selected countries. Structural indicators show that economies with a stronger industrial base and a more balanced structure exhibit higher investment intensity. Conversely, economies dominated by the service sector or structurally weaker economies exhibit lower investment indicators. In this context, Kazakhstan exhibits characteristics of a resource-oriented and structurally unbalanced economy, where investments are unevenly distributed and closely linked to industry specialization (
Figure 3).
The study showed that Norway, a resource-rich country, has an industrially oriented economy and a minimal share of agriculture in GDP, while the service sector remains high at approximately 52%. Latvia, with a low share of industry (until 2022), led in services, later passing to Estonia. Kazakhstan is characterized by higher shares of industry and agriculture than the group average, but a lower share of services (54.3%). Poland and Lithuania demonstrate a balanced structure. Poland leads in government expenditure, but Norway has a higher average (45.4%), while Kazakhstan lags at 27.1%. The results reflect different models of adaptation to external shocks. Norway demonstrates an effective transformation of its resource-based economy through industrialization and a strong role for the state, which ensures resilience in geopolitical instability. The Baltic countries and, to some extent, Poland rely on a service model that is vulnerable to external demand and digital risks, yet flexible. Kazakhstan maintains a raw materials-based and, to some extent, agro-industrial structure, with a low role for the state compared to other countries.