The empirical design is a descriptive distribution-diagnostics framework. It compares scale, weighting, concentration, persistence, and horizon across the WEO historical series and the harmonized WEO–WDI–PWT panel. The design does not estimate a treatment effect or a structural break; each statistic is interpreted according to the empirical object it summarizes.
3.1. Data Sources and Samples
The main source is the International Monetary Fund’s April 2026 World Economic Outlook database (
International Monetary Fund, 2026), supplied as WEOApr2026all.xlsx. The main historical baseline contains 138 countries observed continuously from 1980 through 2024. A separate balanced 137-country sample extends through the provisional 2025 WEO estimates; it is reported as an extension rather than as the historical baseline.
The WEO income variable is NGDPRPPPPC, GDP per capita at constant prices and purchasing power parity in 2021 international dollars. This PPP-adjusted constant-price measure is used to compare real income levels across countries and over time, rather than nominal values directly affected by exchange-rate movements. The population variable is LP, reported in millions. LP supplies country-level population shares for the population-weighted Gini and Theil indices, whereas unweighted calculations treat each country as a single observational unit.
External comparisons use the World Development Indicators (WDI) (
World Bank, 2026) and Penn World Table (PWT) 11.0 (
Feenstra et al., 2025), with the PWT methodology attributed separately to
Feenstra et al. (
2015). WDI uses NY.GDP.PCAP.PP.KD and SP.POP.TOTL; PWT GDP per capita is rgdpe/pop. Their feasible terminal years differ, so harmonized comparisons end in 2023. The source windows are consequently distinct: WEO extends provisionally through 2025, WDI through 2024 where available, and PWT through 2023. WDI and PWT do not replace WEO as the principal long-run source. They provide external checks on whether the direction and metric sensitivity of the WEO findings persist under alternative data-construction systems and feasible terminal years. Source-specific magnitudes are not expected to coincide because the series differ in national-account revisions, PPP implementation, real-income concepts, coverage, and extrapolation practices.
To separate source construction from country composition, the analysis retains 167 countries observed in all three sources for every year from 2000 through 2023, yielding 12,024 source-country-year observations. ISO3 codes define matches. Detailed source-specific coverage and exclusions are reported in
Appendix A Table A2 and
Table A3 and the replication materials.
Balanced and maximum-coverage samples answer different questions. The fixed WEO panels support historical comparison, while maximum samples preserve broader country coverage. The harmonized panel controls cross-source country and year composition but does not make the three measurement systems conceptually identical.
Table 2 summarizes the variables, periods, and sample roles. Further implementation details and complete country lists are retained in
Appendix A and replication outputs rather than repeated in the main text.
3.2. Distributional Measures
Let
denote real GDP per capita at purchasing power parity for country
i in year
t, and let
. WEO and WDI supply PPP-adjusted constant-price GDP per capita directly; PWT GDP per capita is constructed as described in
Section 3.1. Observations with non-positive income are excluded from calculations requiring logarithms. Population weights are each country’s share of sample-year population:
The weights sum to one within each sample-year. Equal-country measures give every economy the same influence; population-weighted Gini and Theil give larger economies greater influence while retaining one country-average income value. They therefore remain between-country measures rather than estimates of global interpersonal inequality.
Income-group definitions depend on the diagnostic. Long-run mobility analysis uses fixed initial-year income quartiles: countries are classified once from GDP per capita in the initial year of the relevant sample and are followed under that classification. Benchmark-year transition matrices compare the quartile associated with the initial benchmark to the country’s terminal-year quartile. By contrast, the 2020–2023 variance decomposition in
Section 4.4 uses fixed 2020 income quartiles applied to both years. These conventions prevent annual reclassification from obscuring persistence and keep the decomposition groups constant over the post-2020 comparison window.
Relative income is defined, where required, as a country’s GDP per capita divided by the contemporaneous cross-country mean or median: , or . It is used only as a descriptive distributional diagnostic. Rank measures the order of countries by GDP per capita at selected benchmark years and supports Spearman correlations, transition matrices, same-quartile persistence, and upward or downward quartile movement.
Sigma dispersion is the cross-sectional standard deviation of log GDP per capita. With
, cross-country mean
, and
countries in year
t, it is defined in Equation (1):
A decline in denotes sigma convergence and an increase denotes sigma widening. The sample-standard-deviation convention uses denominator , rather than , and the convention is applied consistently to the historical series, harmonized comparisons, decomposition, and paired resampling. Equations (1)–(9) and all empirical specifications are otherwise unchanged.
Sigma measures proportional dispersion around the cross-country log-income mean and limits the dominance of extreme level-income observations relative to an untransformed standard deviation. Post-2020 sigma changes are examined in the main historical WEO baseline, the provisional 2025 WEO extension, source-specific WDI and PWT samples over their feasible windows, and the harmonized 167-country panel. This sequence evaluates whether the measured direction is sensitive to a provisional endpoint, source choice, or country coverage without treating the sources as identical.
Gini and Theil indices are computed from country-level GDP per capita because level-income inequality need not move with log-income dispersion. Let
denote the mean GDP per capita across the
countries in year t. The unweighted Gini is defined in Equation (2):
For population weights
satisfying
and weighted mean
, the population-weighted Gini is defined in Equation (3):
The unweighted Theil index is defined in Equation (4):
The population-weighted Theil index is defined in Equation (5):
The unweighted Gini and Theil measures of unweighted between-country GDP per capita inequality assign equal weight to countries. Their population-weighted counterparts measure population-weighted between-country GDP per capita inequality by weighting the same country-level income values by population shares. Neither specification observes income dispersion within countries. Comparisons among sigma, Gini, and Theil are therefore interpreted as evidence on metric sensitivity: sigma describes log-income spread, while Gini and Theil summarize the level-income distribution under different weighting schemes.
All annual metrics are computed within the country set defined for the relevant specification. Consequently, differences between balanced, maximum-unbalanced, and harmonized results can reflect both the income paths represented and the composition rule. The manuscript reports these designs separately rather than pooling them into one series or treating their levels as directly interchangeable.
Tail diagnostics comprise P90/P10, P75/P25, and the ratio of the mean GDP per capita of the ten highest-income countries to that of the ten lowest-income countries. The first two are
and
. The top10/bottom10 ratio is defined in Equation (6):
P90/P10 emphasizes the outer deciles, P75/P25 captures a broader separation around the middle half of the distribution, and top10/bottom10 compares the means of the extreme groups. These measures help distinguish extreme-tail concentration from wider interquartile movement. They remain supplementary because ratios based on distribution tails can be sensitive to small-country observations, source-specific measurement, and exceptionally high incomes; the related sample restrictions are described in
Section 3.4.
3.3. Beta-Convergence and Rank-Mobility Diagnostics
Beta convergence is evaluated using descriptive cross-country growth regressions. For country i, annualized log GDP per capita growth between the initial year
and terminal year
is computed in Equation (7):
The descriptive beta-convergence regression is specified in Equation (8):
A negative descriptive beta-convergence coefficient indicates that initially poorer countries grew faster on average during the specified period. A positive coefficient indicates a positive initial-income–growth gradient, while a coefficient closer to zero than in an earlier window indicates attenuation of catching-up. The regressions are estimated for 2000–2023, for the pre-2020 period 2000–2019, and for the short post-2020 window 2020–2023. These windows permit comparison of the full common-sample relationship with the pre- and post-2020 gradients without treating the latter as a structural break estimate.
The specification is unconditional. The coefficients are used as descriptive growth-gradient summaries, with the short 2020–2023 window interpreted cautiously because it may reflect temporary cross-country macroeconomic variation. Heteroskedasticity-robust White/HC1 standard errors are reported. Each regression is a single cross-section of 167 countries rather than repeated observations within clusters, so no natural clustering dimension is available, and no clustering adjustment is applied. R2 values of approximately 0.01–0.15 are consistent with a summary growth gradient rather than a fully specified growth model; no conditional-convergence claim is made.
Rank mobility provides a separate test of whether average growth differences correspond to changes in relative position. Spearman’s rank correlation between GDP per capita ranks in initial and terminal benchmark years is defined in Equation (9):
Higher indicates greater persistence in the cross-country income ranking, while lower values indicate more extensive reordering. Benchmark-year quartile transition matrices classify countries by initial benchmark-year GDP per capita and record their terminal-year quartile. The matrices yield same-quartile persistence and the shares moving upward or downward by at least one quartile. These diagnostics complement the descriptive beta-convergence regressions because an initially poorer country may grow faster yet remain in the same rank range or quartile when the initial income gap is large; conversely, modest changes near a quartile boundary may produce mobility without a large change in aggregate dispersion.
3.4. Bootstrap, Influence Diagnostics, and Interpretation Boundaries
Sample and source checks combine the balanced historical WEO panel, the separate provisional 2025 extension, the maximum-unbalanced WEO sample, WDI and PWT external comparisons, and the harmonized 167-country panel. Balanced samples maintain fixed composition, the maximum-unbalanced sample preserves wider WEO coverage, and excluding the provisional endpoint tests whether the historical conclusion depends on 2025. The harmonized panel holds countries and years constant across all three sources, while source-specific samples retain their feasible coverage.
Appendix A Table A4 compares maximum- and common-sample results. Together, these designs assess sensitivity to terminal-year status, source construction, and country composition without assigning identical interpretation to all samples. Cross-source agreement is evaluated directionally rather than by requiring equal magnitudes. WEO, WDI, and PWT retain differences in PPP benchmarks, national-account revisions, extrapolation, and real GDP concepts even after harmonization. The common sample reduces the possibility that apparent source disagreement is produced only by different countries, whereas remaining discrepancies may reflect measurement-system differences. This role is distinct from the influence and resampling diagnostics applied within a specified country set.
For selected changes, countries are sampled with replacement, with the same resampled country indices applied to both endpoints. The paired statistic summarizes aggregate change under repeated country draws; its interval and positive-resample share describe sign stability, not country breadth. Leave-one-country-out and predefined exclusion checks assess sensitivity to single observations and sample definitions. Paired resampling, additive contributions, and joint exclusion therefore retain distinct estimands.
The interpretation boundary is fixed throughout. The post-2020 period is treated as a descriptive event window rather than an exogenous treatment; decomposition, contribution, exclusion, beta-convergence, rank, and cross-source diagnostics characterize observed distributional patterns rather than identify specific shocks or mechanisms. Unweighted and population-weighted results remain measures of between-country GDP per capita inequality rather than estimates of global interpersonal inequality.
3.5. Historical Magnitude and Concentration Diagnostics
For country
in year t, the additive contribution to the cross-sectional variance of log GDP per capita is defined as
where
xit is log GDP per capita,
is its cross-country mean in year
t, and
N is the number of countries. The country-specific contribution to the 2020–2023 variance change is
By construction, these country-specific changes sum to the full-sample variance change, subject to machine precision. Top-five and top-ten shares are computed within the original full sample and identify concentration; they do not by themselves describe the recomputed distribution after exclusions.
The joint-exclusion exercise ranks countries by positive and jointly removes the first k contributors from both endpoints for recomputing sigma and variance on each remaining sample. Because exclusion changes the mean, sample size, and all deviations, the residual path need not equal one minus the original cumulative contribution share. The exercise is reported as an ex-post concentration stress test. Historical magnitude is assessed with on the locked 138-country WEO 1980–2024 series. The primary benchmark compares 2020–2024 with the 36 preceding four-year windows ending by 2019; the complete 41-window distribution is supplementary. Because the windows overlap and are serially dependent, ranks and empirical cumulative shares are descriptive and are not interpreted as p-values, confidence intervals, filters, or structural-break tests. Given the short four-to-five-year terminal window, the rolling-window benchmark is used instead of a filtered trend-cycle decomposition, which would introduce additional smoothing, tuning, and endpoint assumptions that the present descriptive design cannot discipline reliably.