1. Introduction
Sustainable development has become a central guiding principle for urban policy and planning in the 21st century and is commonly conceptualized around three interdependent pillars: economic, social, and environmental sustainability [
1,
2]. While these dimensions are mutually reinforcing, economic sustainability occupies a foundational position, as the ability to generate stable output, productive employment, and fiscal resources over time conditions the feasibility of both social welfare provision and environmental protection [
3,
4,
5]. In urban contexts, economic sustainability is not an abstract or autonomous outcome; rather, it is fundamentally shaped by cities’ economic performance, which determines their capacity to sustain growth, absorb economic shocks, and support long-term development trajectories [
6,
7,
8,
9].
Accordingly, economic performance constitutes an indispensable prerequisite for urban sustainability [
10]. Cities with strong economic performance are better positioned to generate employment, mobilize public revenues, and finance investments in infrastructure, social services, and environmental improvement [
11,
12,
13]. By contrast, persistent weaknesses in urban economic performance constrain fiscal capacity, exacerbate social inequalities, and undermine long-term sustainability objectives. In an increasingly globalized economic environment, the most successful cities are those capable of formulating and implementing development strategies that extend beyond national economic dynamics and position them as competitive actors in international markets [
14]. This research therefore focuses on comparative assessment of urban economic performance, rather than sustainability in a broad sense.
Despite the acknowledged importance of benchmarking urban economic performance, systematic evaluation poses substantial analytical challenges because urban economies are inherently multi-dimensional and heterogeneous. In response, multi-criteria decision-making (MCDM) approaches enable the simultaneous consideration of multiple, often conflicting, performance dimensions within a unified structure and make their relative importance explicit through structured weighting and aggregation procedures, thereby facilitating robust and comparable assessments of complex alternatives [
15]. Despite the widespread application of MCDM techniques in economic, environmental, and sustainability-related studies, existing applications have predominantly focused on countries or regions [
16,
17,
18,
19]. City-level analyses—especially those explicitly centred on economic performance—remain comparatively limited, despite the pivotal role of cities in contemporary economic systems.
Within this context, the Global Power City Index (GPCI) provides an internationally recognized empirical foundation for evaluating cities from a multi-dimensional perspective [
20]. Its Economy function comprises 13 indicators organized into six groups: market size, market attractiveness, economic vitality, human capital, business environment, and ease of doing business, which capture the core structural dimensions shaping urban economic performance. The index nevertheless remains descriptive in construction. It aggregates indicator scores without differentiating the relative importance of these six dimensions, so their contribution is fixed by construction rather than derived from a decision-analytic procedure; it also reports a single score per city and year, leaving the imprecision of the underlying measurement unexpressed. Closing both gaps requires an MCDM structure in which criterion importance is derived explicitly and imprecision is carried through the evaluation.
To address these gaps, the current study proposes an integrated grey-based MCDM approach to assess and compare cities’ economic performance. Building on the economic dimension of the GPCI, the proposed approach uses a grey-based RANCOM tool to determine criterion importance and a grey-based MUNRA tool to generate robust city rankings.
The usage of interval grey numbers in this research follows from the kind of information available at each stage of the analysis. A fuzzy model would require a membership function, and such functions are generally specified from expert intuition, which varies across decision makers and weakens methodological consistency [
21,
22]. A probabilistic model would require treating a city’s five annual observations as draws from a stable distribution, whereas they are successive measurements of a changing quantity. A grey number requires only a lower and an upper bound [
23,
24], and bounds are exactly what the two inputs of this study supply. The published GPCI scores carry no dispersion estimate, so the range observed over the 2021–2025 period is the only imprecision in the source documents. The experts, in turn, are asked for rank positions rather than degrees of preference (
Section 4.2), so each pairwise relation is known only as one of three ordinal states. Such states can be bounded, but they cannot identify a membership function. Grey numbers are therefore adopted because they represent what is known about the data and the judgments without adding what is not.
Section 5.5 examines whether this representation also yields better rankings.
Within this interval grey environment, the two methods serve complementary functions at the two stages of the decision process. G-RANCOM requires only a ranking of the criteria, which reduces the cognitive effort demanded of the experts and permits ties between criteria of equal importance. It is executed for each expert individually, so divergence of opinion is preserved in the width of the weight intervals rather than eliminated by a consensus ranking. G-MUNRA evaluates the alternatives under linear, vector, and non-linear normalization simultaneously and aggregates the three outcomes, which limits scale sensitivity and the risk of rank reversal. Grey models operating at both stages already exist (
Section 2), but they obtain the weights either from the data alone, without consulting experts, or from experts who must state how much more important one criterion is than another, and they rank the alternatives under a single normalization scheme. RANCOM and MUNRA have not previously been formulated with interval grey numbers, and their combination carries uncertainty from the first expert judgement to the final ranking without asking the experts for more than an order.
To demonstrate the framework’s effectiveness and practical applicability, we conduct a comprehensive case study. The case study compares the economic performance of sixteen cities across G7 economies, representing some of the most influential urban centres in the global economic system.
Accordingly, the study addresses three research questions. RQ1: Which dimensions of the GPCI economic framework carry the greatest importance when criterion weights are elicited from expert rankings under interval uncertainty?
RQ2: How do G7 cities rank when imprecision is propagated through both the weighting and the ranking stage rather than removed at the outset?
RQ3: How robust is the resulting ranking to the elicitation protocol, to the parameters of the grey model, and to the choice of decision-making paradigm?
This paper contributes to the literature regarding urban economics, city benchmarking, and multi-criteria decision-making in several important ways. First, it repositions city-level economic performance as the primary object of evaluation, whereas existing urban MCDM applications typically embed economic indicators within broader composite constructs of smartness, livability, or sustainability. Second, it introduces a hybrid grey-based MCDM methodology that combines RANCOM and MUNRA and, to the authors’ knowledge, provides the first formulation of either technique in an interval grey environment, so uncertainty carries through both the weighting and ranking stages. Third, by grounding the assessment in internationally recognized indicators and by documenting the weight structure, the uncertainty treatment, and the stability of the resulting ranking, it offers a benchmarking tool that policymakers, investors, and other stakeholders can apply to cities in advanced economies.
The remainder of the paper is organized as follows:
Section 2 reviews the relevant literature on urban economic performance and MCDM-based evaluation approaches.
Section 3 presents the proposed grey-based RANCOM–MUNRA methodology, while
Section 4 reports the empirical case study and application results.
Section 5 reports the sensitivity and comparative robustness analyses,
Section 6 discusses the findings together with the theoretical contributions and policy implications,
Section 7 states the limitations and directions for future research, and
Section 8 concludes the paper.
2. Literature Review
This section reviews the evolving literature on applying MCDM approaches to assess urban and economic performance. It first synthesizes earlier urban MCDM studies that evaluate cities through multi-dimensional performance frameworks, followed by an examination of MCDM-based economic performance assessments conducted at the national and regional levels. By juxtaposing these two strands of research, this section highlights prevailing methodological tendencies, identifies limitations in how economic performance and uncertainty are treated at the city level, and delineates the research gap that motivates the current manuscript.
2.1. MCDM-Based Assessment of Urban Performance
Recent scholarship increasingly employs multi-criteria decision-making (MCDM) methods to examine urban performance by integrating diverse indicators reflecting dimensions such as smartness, sustainability, livability, and competitiveness. In these analyses, cities are typically framed as decision alternatives and evaluated across systematically defined criteria, often informed by international benchmarking indices or composite indicator systems. For instance, Gelmez and Özceylan [
19] ranked 118 cities in the Smart City Index by distinguishing between technological and structural dimensions, using entropy-based weighting and COPRAS and ARAS ranking methods. Similarly, Ye et al. [
25] provided a methodologically relevant contribution by developing a hybrid MCDM framework that explicitly incorporated digital-economy-related indicators as a core assessment dimension. Focusing on nine cities in China’s Pearl River Delta region, the study assessed urban performance across digital infrastructure, smart living, and digital economy dimensions using entropy-based weighting in combination with WSM, TOPSIS, and VIKOR. Moreover, Sotirelis et al. [
9] assessed the smart performance of global cities using a PROMETHEE II outranking model based on six widely recognized smart city pillars.
Building on this general framework, several studies have concentrated on refining indicator structures and ranking mechanisms. For instance, Yürüyen and Ulutaş [
16] assessed 17 European cities employing LOPCOW for weighting and RAWEC for ranking based on GPCI data, while Işık et al. [
15] introduced a hybrid LOPCOW–CRITIC–ALWAS framework to assess the competitive performance of 12 European cities. Sustainability-oriented applications have followed similar methodological patterns; Özekenci [
26] analyzed 10 European capital cities using MEREC for weighting and compared RAWEC, Extended AROMAN, and MARA ranking results based on the Sustainable Cities Index.
Urban MCDM studies have also expanded to developing-country contexts and environmental dimensions. Taş and Alptekin [
17] assessed the smart city performance of 30 Turkish metropolitan municipalities using MEREC and MARCOS, while Komasi et al. [
27] evaluated the environmental competitiveness of 14 Iranian cities using ITARA–FUCOM weighting and multiple ranking algorithms (MARCOS, CODAS, EDAS, and TOPSIS) to ensure robustness. In addition, efficiency-based approaches have complemented MCDM rankings; Kutty et al. [
28] applied a double-frontier DEA framework to evaluate the long-term sustainability of 35 European smart cities, incorporating undesirable outputs and productivity dynamics.
Few studies have explicitly addressed uncertainty in urban performance evaluation. Aytaç Adali et al. [
29] introduced a grey-based MCDM framework to assess the smart performance of 17 European cities by modelling temporal variability and data imprecision through grey numbers, G-LBWA weighting, and G-EDAS ranking. Moreover, Ogrodnik [
30] highlighted methodological sensitivity, demonstrating that rankings of Polish cities varied substantially across TOPSIS and PROMETHEE, underscoring the importance of method selection in urban assessment.
2.2. MCDM-Based Assessment of Economic Performance
Parallel to the urban literature, MCDM tools have been extensively employed to assess economic performance at the country and regional levels. This body of work consistently underscores that macroeconomic performance is inherently multi-dimensional and cannot be adequately captured through single-indicator approaches. Ulutaş et al. [
31] proposed a grey-based MCDM method to analyze the macroeconomic performance of G7 countries, using LOPCOW-G weighting and RAWEC-G ranking, explicitly addressing uncertainty in macroeconomic data. Similarly, Üre et al. [
32] assessed the macroeconomic performance and innovation capacity of the G7 countries and Türkiye employing objective weighting (LODECI–PSI) and a distance-based ranking approach (WEDBA).
Broader comparative studies have examined larger country groups and methodological robustness. Baydaş et al. [
33] assessed the economic performance of G20 countries by applying multiple MCDM algorithms and normalization techniques, and validated the rankings against GDP per capita and the Environmental Performance Index. Oussama et al. [
34] constructed a TOPSIS-based macroeconomic performance index for 16 MENA countries over a long time horizon, while Komasi et al. [
35] analyzed economic development levels in South America using modified ITARA weighting and AROMAN–RAM ranking. At the subnational scale, Tekman and Ordu [
36] compared 26 Turkish development regions using SWARA-based weights and the CoCoSo method, and Kuncova and Seknickova [
37] introduced a two-stage PROMETHEE II framework to capture temporal importance in regional economic performance.
Table 1 summarizes the reviewed urban and economic performance studies by analysis unit, performance focus, weighting and ranking tools, uncertainty handling, and external validation, positioning the present research within this body of literature.
2.3. Research Gap
The existing literature indicates a clear separation between urban MCDM applications and MCDM-based economic performance assessments. Urban studies have predominantly focused on smartness, sustainability, or general competitiveness, with economic indicators typically treated as secondary components rather than as the primary evaluation objective. Conversely, MCDM-based economic performance analyses—particularly those adopting uncertainty-aware or grey-based approaches—have largely concentrated on countries or regions, overlooking the city level as the main locus of economic activity.
This methodological divide has limited existing frameworks’ ability to capture the spatial concentration and heterogeneity of economic performance within advanced economies. In globally comparable contexts such as the G7, economic dynamics are shaped largely by major metropolitan areas rather than by national aggregates. However, current MCDM studies do not sufficiently address this city-level perspective, nor do they explicitly integrate uncertainty into urban economic benchmarking.
Accordingly, a clear research gap remains in developing an advanced MCDM framework that integrates uncertainty and incomplete information, explicitly tailored to assess and rank the economic performance of cities within advanced economies. Bridging this gap necessitates positioning economic performance as the primary assessment objective and adapting grey-based decision-making methodologies from national-level applications to the urban context.
2.4. Methodological Motivation and Contribution
Motivated by this gap, this study proposes a grey-based multi-criteria decision-making framework to comparatively evaluate economic performance in G7 cities. Unlike existing urban MCDM applications that typically embed economic indicators within broader composite constructs, the proposed framework treats city-level economic performance as the primary decision objective. To operationalize this focus, we develop and integrate two novel grey-based methods, G-RANCOM for determining criterion weights and G-MUNRA for ranking alternatives, for the first time within a single analytical framework.
Adopting grey system theory allows the model to explicitly address uncertainty, imprecision, and variability inherent in urban economic data, thereby extending grey-based macroeconomic MCDM approaches to the city scale in a systematic manner. Within this structured decision architecture, economic performance is assessed using six core criteria: Market Size, Market Attractiveness, Economic Vitality, Human Capital, Business Environment, and “Ease of Doing Business” derived from the Global Power City Index (GPCI) framework of the Mori Memorial Foundation. These criteria capture the fundamental dimensions of urban economic competitiveness while ensuring consistency and comparability across G7 cities.
By adopting grey-number-based weighting and ranking mechanisms to accommodate the uncertainty inherent in both data and expert assessments, the recommended decision-making framework enables a systematic derivation of criterion importance and a reliable prioritization of cities as decision alternatives. Within this context, our manuscript offers three main contributions. First, it extends urban MCDM research by explicitly focusing on assessing economic performance at the city level. Second, it contributes to the grey decision-making literature by developing and applying, for the first time, the G-RANCOM weighting approach and the G-MUNRA ranking procedure. Third, it presents a theoretically well-founded, policy-oriented benchmarking framework that facilitates comparative analysis of economic performance across major global cities.
3. Research Methodology
This research employs a hybrid grey group decision-making approach to comparatively assess the economic performance of G7 cities using multi-dimensional urban economic indicators. The methodological framework, illustrated in
Figure 1, is designed to handle uncertainty, data heterogeneity, and incomplete information commonly encountered in city-level economic analysis by combining expert judgments with objective secondary data within a grey-system environment. At the outset, the core elements of the decision problem, including the economic performance criteria, city alternatives, and expert committee, are clearly defined. To provide the theoretical basis of the proposed approach, the essential concepts of grey numbers are briefly introduced. The analytical procedure then proceeds in two sequential stages. In the first stage, the G-RANCOM tool identifies the weights of the economic criteria, yielding a set of grey criterion weights that reflect the expert panel’s collective assessments. In the second stage, the G-MUNRA technique derives the final ranking of the G7 cities.
3.1. Grey Numbers
Deng Julong introduced Grey System Theory in the early 1980s to analyze systems characterized by incomplete, limited, or partially observable information. The theory was developed as an alternative to classical deterministic (crisp) and fuzzy-based modelling frameworks, which represent uncertainty either through precise point values or through membership functions. In many decision-making environments, neither exact numerical values nor well-defined membership functions can be reliably specified due to limited data availability or ambiguous expert judgments. Grey system theory addresses this gap by modelling uncertainty through bounded intervals, thereby avoiding the need for probabilistic assumptions or predefined membership structures.
Within this framework, information is classified into three categories by its degree of observability. White information refers to fully known and precisely determined values, whereas black information represents quantities for which no meaningful information is available. Grey information lies between these two extremes and characterizes situations in which exact numerical values are unknown but can be reasonably confined within known lower and upper bounds. Since most decision-making problems involve such partially known information, grey system theory offers a mathematically tractable and conceptually transparent approach to uncertainty modelling. In this context, grey numbers represent uncertainty, preserving interval-based information while avoiding the need for probability density functions or fuzzy membership functions.
3.1.1. Definitions of Grey Numbers [38]
Definition 1. Grey number
A grey number, denoted by , is a numerical quantity whose exact value is unknown, while its range of possible values is known.
Definition 2. Interval grey numbers
A grey number is defined as in Equation (1) Here, and denote the lower and upper bounds of the grey number, respectively. This interval-based representation explicitly incorporates uncertainty into mathematical modelling while preserving computational simplicity.
3.1.2. Basic Arithmetic Operations of Grey Numbers [23,39]
Let and be two interval grey numbers. The fundamental arithmetic operations are defined as follows:
Definition 3. Addition of grey numbers
The addition of two grey numbers is given by Equation (2) Definition 4. Subtraction of grey numbers
The subtraction of two grey numbers is defined in Equation (3). Definition 5. Multiplication of grey numbers
The multiplication of two grey numbers is provided in Equation (4) Definition 6. Division of grey numbers
Provided that , the division of two grey numbers is formulated in Equation (5). 3.1.3. Scalar and Functional Operations of Grey Numbers [40]
Beyond basic arithmetic, several scalar and functional operations are defined on grey numbers to support normalization, transformation, and aggregation procedures commonly used in multi-criteria decision-making.
Definition 7. Scalar multiplication of a grey number
Let be a scalar and be a grey number. Scalar multiplication is expressed in the following equation. Definition 8. Power of a grey number
For a non-negative real exponent , the power of a grey number is defined in Equation (7). Definition 9. Inverse of a grey number
If , the inverse of a grey number is given by Equation (8). 3.1.4. Whitenization of Grey Numbers [41]
To enable comparison, aggregation, and ranking, grey numbers are often transformed into crisp values employing a whitenization procedure.
Definition 10. Whitenization function
For a grey number , the whitened value is defined as: The whitenization coefficient reflects the decision maker’s attitude toward uncertainty. A neutral assessment is obtained when = 0.5, which corresponds to the midpoint of the grey interval.
3.2. Grey RANCOM-MUNRA Methodology
3.2.1. Stage 1: Grey Ranking Comparison Method (G-RANCOM) Algorithm for Criterion Weighting
RANCOM is a subjective weighting procedure that identifies the relative importance of criteria solely from ordinal expert judgments [
42]. Unlike traditional weighting procedures that rely on extensive pairwise comparisons, the original RANCOM methodology utilizes only ranking information, thereby reducing the cognitive burden on experts and naturally accommodating equal-importance relationships among criteria. In this study, we extend the original RANCOM methodology in two ways. First, we integrate interval grey numbers to explicitly capture the ambiguity and imprecision inherent in expert evaluation. Hence, the importance of a criterion is expressed through lower and upper bounds rather than crisp values. Second, the procedure is implemented individually for each member of the expert panel, and the individual grey weight vectors are combined in the final step using grey expert-importance coefficients. Consequently, the members of the panel do not need to agree on a consensus ranking before the model is run, and the differences in opinion among the experts are not eliminated, but rather preserved as information. A step-by-step explanation of the G-RANCOM technique is provided below. Within the multi-criteria procedure,
denote the criteria, whereas
indicate the experts participating in the assessment process.
Step 1. Ascertain the ranking order of the criteria.
In this step, the relevant criteria for the decision problem are determined. Each expert then separately ranks the assessment criteria by their relative importance.
Each expert’s judgement is formalized as a ranking function . This ranking function assigns a rank position to each criterion, where a lower rank reflects higher significance. In situations where criteria are deemed equally important, ties are admissible, indicating that .
Step 2. Form the grey criterion comparison matrices.
Based on the importance ranking obtained in Step 1, a grey criterion comparison matrix is constructed for every expert
as:
Here, stands for the grey relative importance of the criterion with respect to as judged by the expert and and denote the lower and upper bounds of the corresponding grey interval, respectively.
The elements of each grey criterion comparison matrix are identified through the use of the grey importance scale detailed in
Table 2. Expert evaluations use a three-level grey linguistic framework based on this scale, which translates ordinal preference information into interval-valued representations. First, the ranking positions
and
are compared to identify the relative significance, which is then assigned to the corresponding grey interval according to the scale in
Table 2. In this manner, the comparison matrix systematically encodes whether criterion
is assessed as less important than, equally important as, or more important than criterion
.
Step 3. Calculate the summed grey criteria values.
The elements of each matrix
derived in Step 2 are then aggregated in this step to calculate the total values for each criterion. The aggregation is performed independently for each expert employing the matrices associated with them. These values represent intermediate dominance measures derived from the comparison structure and serve as input for the normalization step used to obtain the final grey criterion weights. The summed grey values for each criterion and expert are computed as:
In the equation above, elements constitute the j-th row of , where the summation index extends across the columns.
Step 4. Compute the individual grey weight coefficients.
The grey weight coefficient of criterion
under the judgement of expert
, denoted
, is computed by normalizing the summed grey criterion value obtained in Step 3 with respect to the total of the summed grey values of that expert.
In Equation (12), the summation index works over the set of criteria, so that the normalization procedure is performed within the opinion of the single expert .
Step 5. Assign the grey expert-importance coefficients
Following [
43], the importance of expert
is expressed as an interval grey number
chosen from the five-level linguistic scale given in
Table 3.
Step 6. Calculate the final grey weight coefficients.
In the last step, each individual weight vector obtained in Step 4 is multiplied by the importance coefficient of the corresponding expert and averaged over the
experts, as formulated in Equation (13). The aggregated values are then normalized via Equation (14) to produce the final grey criterion weights.
In Equation (13), stands for the grey multiplication operator. Finally, the grey weight vector calculated based on Equation (14) is used as input in Stage 2.
3.2.2. Stage 2: Grey Multiple Normalization Ranking Analysis (MUNRA) Algorithm for City Ranking
The MUNRA technique is a recent MCDM method introduced by [
44] for ranking alternatives by simultaneously using linear, vector, and non-linear normalization procedures. By integrating multiple normalization perspectives within a unified aggregation structure, MUNRA enhances robustness and reduces the sensitivity of ranking results to the choice of a single normalization scheme.
In this research, the classical MUNRA framework is extended into a grey system environment by incorporating interval-valued grey numbers to explicitly model uncertainty and incomplete information inherent in urban economic performance data. The resulting Grey MUNRA (G-MUNRA) tool operates entirely within a grey arithmetic structure, allowing uncertainty to be propagated throughout the normalization, aggregation, and ranking stages. The computational steps of the developed G-MUNRA approach are presented below.
Step 1. Form the initial grey performance matrix
.
Here, stands for the grey performance value of the city alternative under the economic-related criterion. Here, denotes the number of city alternatives and the number of criteria.
Step 2: Construct the normalized grey performance matrix .
In this step, the grey performance matrix is normalized using three complementary normalization techniques, namely linear, vector, and non-linear normalization, in order to ensure comparability across criteria with different measurement scales. The normalization procedures are applied separately for benefit-type and cost-type criteria, yielding three normalized grey performance matrices.
In the linear normalization procedure, the normalized grey performance values are computed via Equations (16) and (17). For the vector normalization procedure, Equations (18) and (19) are applied. Finally, non-linear normalization is employed to enhance discrimination among alternatives, with the normalized grey values computed using Equations (20) and (21).
In Equations (16)–(21), and denote the sets of benefit-type and non-benefit-type criteria, respectively.
Step 3. Aggregation of weighted grey scores under each normalization.
After applying the linear, vector, and non-linear normalization procedures, the normalized grey performance values are aggregated using the grey criterion weights computed via Equation (14) of Stage 1. This step performs a weighted aggregation in a grey environment, combining criterion-level contributions to obtain an overall grey performance score for each alternative under each normalization procedure. Thus, the aggregated weighted grey scores for the alternative
are computed according to Equations (22)–(24).
Here, the operator denotes grey multiplication, while the summation represents the grey aggregation of criterion-level contributions. The resulting interval-valued scores , , and preserve uncertainty information and serve as inputs for the subsequent whitenization step.
Step 4. Whitenization of weighted grey scores.
After computing the weighted grey scores from the different normalization schemes, a whitenization step is performed to convert the interval-valued results into crisp scores. To this end, the whitenization operator of Equation (9) is applied at the neutral level
= 0.5, which reduces it to the midpoint transformation defined in Equations (25)–(27).
Step 5. Final aggregation and ranking.
The final MUNRA score for each city alternative is obtained by aggregating the crisp scores from the three normalization procedures: linear, vector, and non-linear. Accordingly, the overall performance score of the alternative
is computed with the aid of Equation (28).
In Equation (28), , , and denote the whitened (crisp) scores obtained from linear, vector, and non-linear normalization procedures, respectively. The aggregation coefficients , , and satisfy the normalization condition . In the absence of any prior preference regarding the relative contribution of the normalization procedures, equal weights are adopted, such that . Finally, the city alternatives are ranked in descending order of , with higher values indicating superior overall economic performance.
4. Case Study: Urban Economic Performance Assessment
A real-world case study is conducted to demonstrate the feasibility and practical relevance of the developed grey group decision-making framework for assessing urban economic performance. The case study focuses on the comparative assessment of selected G7 cities and illustrates how the integrated Grey RANCOM–MUNRA method operates within a complex, data-intensive urban context characterized by uncertainty and heterogeneity.
To implement the case study, a well-defined decision-making problem is formulated. The decision problem is structured around three key components: an expert group, a set of criteria, and a group of city alternatives. The expert team comprises five members who have academic and professional expertise in the fields of economic policy, urban economics, and regional development.
The evaluation criteria comprise six economic performance indicators, all derived from the Global Power City Index (GPCI) framework and representing the core dimensions of urban economic performance. The decision alternatives consist of 16 cities in G7 countries, selected for their global economic significance and the availability of consistent, comparable data across the study period.
This section presents the case study systematically and with methodological consistency. It begins by defining the decision problem the study addresses. It then describes the expert committee involved in the decision-making process. Next, it introduces the economic performance criteria used in the analysis. The next part presents the G7 cities considered as decision alternatives. Finally, the section concludes by implementing the developed grey-based decision support framework.
4.1. Problem Definition
Urban economic performance assessment is an inherently complex MCDM problem because of the multi-dimensional nature of city economies, the heterogeneity of economic indicators, and uncertainty in both data and expert judgments. Cities differ markedly in economic scale, productivity, innovation capacity, labour market dynamics, and global integration, making single-indicator comparisons inadequate and potentially misleading. Consequently, a structured decision-support framework that jointly considers multiple economic dimensions is required to enable transparent and consistent comparative assessment.
In this work, we formulate the decision problem as the comparative assessment and ranking of G7 cities based on their economic performance. The evaluation integrates objective urban economic indicators with expert judgments under uncertainty. Although standardized, internationally comparable data are available, determining the relative importance of economic performance criteria remains nontrivial, particularly when cities exhibit clear trade-offs across economic dimensions.
To overcome these challenges, the problem is embedded in a grey group decision-making environment, where grey interval numbers explicitly represent uncertainty and incomplete information. While the relative importance of economic criteria is determined by considering the independent opinions of each member of the expert committee, city-level performance is evaluated using objective data from GPCI reports for the 2021–2025 period. In this context, this work introduces a hybrid grey RANCOM-MUNRA decision-making methodology that systematically integrates criterion weighting and alternative ranking, enabling the consistent aggregation of expert knowledge and quantitative data while preserving heterogeneity among experts.
4.2. Expert Panel Formation and Elicitation Protocol
The expert committee was formed using purposive sampling after reviewing resumes. Participation required at least 10 years of relevant experience and expertise in fields associated with the assessment problem. We individually invited eight suitable experts via email, summarizing the study’s aim and significance, and five responded by agreeing to participate. Next, we held an individual online meeting with each expert-team member, using standardized information and instructions, to explain the six GPCI-based criteria and the information-gathering procedure. The experts who participated in our research declared no conflict of interest associated with the assessed cities or the organization that provided the data.
Table 4 provides detailed information about the expert-team members.
The data collection process consisted of two independent rounds. Round 1 (2–7 August 2026) provided the sequential rankings required for G-RANCOM, while Round 2 (11–14 August 2026) obtained the level-based evaluations needed for the F-LBWA benchmark model. In both rounds, the experts first participated in an individual online meeting, after which the survey was distributed and collected via email. Each expert filled out the forms individually, and the responses were recorded as
. Therefore, we recorded differences in opinion as separate information and addressed them through the data collection procedures of the relevant methodologies.
Figure A1 summarizes this process, and
Appendix B and
Appendix G provide the survey content and response instructions.
4.3. Definition of Economic Performance Criteria
In this research, the economic performance of G7 cities is formulated as a multi-criteria decision-making problem comprising six main economic criteria that constitute the evaluation’s analytical structure. These criteria are derived from the GPCI’s economic framework and represent the core aspects through which urban economic performance is defined and assessed [
20].
The first criterion, market size , captures the overall scale and productivity of economic activity within a city. It is measured using nominal GDP and GDP per capita . Nominal GDP reflects the absolute magnitude of economic output and indicates the city’s capacity to sustain large-scale production, consumption, and investment activities. In contrast, GDP per capita captures productivity and income levels, thereby distinguishing cities that are large in scale from those that are economically efficient.
Market attractiveness reflects the ability of a city to attract investment, firms, and economic activity over time. This criterion is measured via GDP growth rate and economic freedom . GDP growth rate captures economic dynamism and growth momentum, providing insight into a city’s adaptability and expansion capacity, particularly during periods characterized by economic shocks and subsequent recovery. In addition, economic freedom reflects the institutional and regulatory environment in which economic activities take place, including market openness and regulatory efficiency.
Economic vitality focuses on the structural depth and global integration of the urban economy. It is captured through stock market capitalization and the number of the world’s top 500 companies . Stock market capitalization reflects the scale and sophistication of financial markets, indicating the city’s role in capital mobilization and financial intermediation. Furthermore, the presence of the world’s leading firms captures the concentration of corporate headquarters and high-value-added business activities, which are critical for global economic connectivity and strategic decision-making.
Human capital represents the labour market capacity and skill base of urban economies and is captured through total employment and employees in business support services . Total employment reflects the city’s ability to generate and sustain jobs, thereby indicating economic scale and labour market inclusiveness. At the same time, employees in business support services capture the concentration of specialized labour supporting advanced economic activities, such as professional, managerial, and corporate services.
Business environment captures the operational conditions firms face within the city. It is measured using wage levels , the availability of skilled human resources , and the variety of workplace options . Wage level reflects labour cost structures while simultaneously signalling productivity and skill intensity. Moreover, the availability of skilled human resources measures how easily firms can access qualified labour, a key determinant of innovation capacity and business expansion. Finally, the variety of workplace options reflects the diversity and flexibility of working environments, which increasingly influence firm location decisions in knowledge- and service-based urban economies.
The sixth and final criterion, ease of doing business , reflects the regulatory and risk-related conditions affecting economic activity. This criterion is measured by corporate tax rate and political, economic, and business risk . Corporate tax rate captures the fiscal burden imposed on firms and plays a central role in investment and location decisions. In parallel, political, economic, and business risk reflects the degree of uncertainty faced by economic agents, encompassing macroeconomic stability, institutional reliability, and the overall risk environment.
Of the 13 sub-indicators that make up the six main criteria, 11 are benefit-type indicators in their raw form. The remaining two, the corporate tax rate and the level of political, economic, and business risk, are cost-type indicators by nature. However, the Mori Memorial Foundation applies reverse scoring to these two indicators when computing the published scores. As a result, a lower raw value yields a higher score, and every published sub-indicator and criterion group score is uniformly oriented towards maximization. The sub-indicator scores are then averaged within each criterion group, and the six criterion group scores from the Economy function constitute the criterion values utilized in this research [
20].
Table A1 of
Appendix A presents the definition, the type, and the transformation of each sub-indicator.
4.4. Definition of City Alternatives
In this manuscript, we compare urban economic performance across 16 cities in countries commonly referred to as the G7. The G7 economies frame the analysis on three grounds: (i) they account for a large share of global output and host the metropolitan areas in which their financial, corporate, and labour-market functions are concentrated, (ii) they operate under broadly comparable institutional and regulatory conditions, which keeps observed differences attributable to urban economic performance rather than to national development stage, and (iii) the GPCI reports the six Economy function scores for every G7 city it covers in each edition of the 2021–2025 period, so that each grey interval rests on a full five-year series, and all 16 such cities are retained. Detailed information regarding the selected cities is presented in
Table 5.
4.5. Implementation of the Developed Approach for Criteria Weight Estimation and City-Level Ranking
4.5.1. Criterion Weighting via the G-RANCOM Method
As shown in
Table 4, the expert team combines profiles that differ in seniority and institutional background but share complementary expertise in urban and economic analysis, so the experts receive differentiated grey importance coefficients rather than equal weights. Each member of the expert committee separately ranked the six economic criteria according to the data collection protocol described in
Section 4.2. Five separate ranking functions
were retained as separate inputs and reported in
Table 6. Subsequently, the three-level scale in
Table 2 was applied to each pair of ranked criteria, and a grey comparison matrix
was constructed for each expert utilizing Equation (10). The resulting five matrices are provided in
Table A3,
Table A4,
Table A5,
Table A6 and
Table A7 of
Appendix B.
Table A2 provides the content and guidelines for the first round of the survey. Summing their rows via Equation (11) produces the summed grey criteria values provided in
Table 7. Normalizing these values by Equation (12) yields the individual grey weight vectors given in
Table 8. Next, the individual vectors are integrated through Equation (13), employing the grey expert-importance coefficients in the last two rows of
Table 4, and the aggregate is normalized with the help of Equation (14). The final grey criterion weight values are given in
Table 9.
The five ranking functions are not identical, so their agreement is quantified before aggregation. Kendall’s coefficient of concordance, computed on the mid-ranks of
Table 6 with the standard correction for ties, equals
W = 0.656 (χ
2(5) = 16.40,
p = 0.006), and the mean pairwise Spearman correlation is 0.568, ranging from 0.348 to 0.909. Agreement on the overall importance structure is therefore substantial and statistically significant, while divergence on individual criteria remains. Since consensus was not imposed (
Section 4.2), this divergence is retained as information and carried by the grey intervals of the weight vector, and
Section 5.1 indicates that it does not significantly influence the final ranking.
4.5.2. Ranking City Alternatives Employing the Grey MUNRA Method
Table 10 presents the grey performance matrix
formed for the city alternatives. For each city alternative and criterion, the minimum and maximum GPCI scores observed during 2021–2025 are used as the lower and upper limits of the corresponding grey number, respectively. Accordingly, this grey matrix reports each city’s interval-valued performance across the six criteria, capturing temporal variability and cross-city heterogeneity in urban economic indicators derived from the GPCI datasets. Performance values of city alternatives for each criterion during 2021–2025 are presented in
Table A8 of
Appendix C.
Based on the grey performance matrix given in
Table 10, the normalized grey performance matrices are formed in accordance with the normalization procedures defined in Equations (16)–(21). Specifically,
Table 11 presents the results of linear normalization obtained using Equation (16),
Table 12 reports the vector-normalized grey performance matrix obtained using Equation (18), and
Table 13 presents the non-linear normalization results obtained using Equation (20).
After normalization, the weighted aggregation of the grey performance values is performed using the G-MUNRA technique defined in Equations (22)–(24). For each normalization structure, criterion-level grey values were multiplied by the corresponding grey weights and subsequently aggregated to obtain alternative-level grey scores. These aggregated grey scores are then transformed into crisp values through the whitenization procedure specified in Equations (25)–(27), yielding representative scores for linear, vector, and non-linear normalization. Finally, we computed the overall performance score for each city alternative by aggregating the whitened scores using Equation (28), with equal aggregation coefficients for the three normalization procedures.
Table 14 summarizes the resulting G-MUNRA scores and rankings. Based on
Table 14, the final ranking of city alternatives is
.
4.6. Redundancy and Collinearity of the Criterion Set
The six criteria are drawn from the GPCI’s economic framework [
20] and are conceptually distinct from one another by their very nature. However, because they define different aspects of the same urban economies, a certain statistical relationship between them is expected.
Table 15 displays the Pearson correlation coefficients calculated across 16 city alternatives based on the midpoints of the grey performance intervals in
Table 10. As shown in
Table 15, the highest correlation coefficients between criterion pairs were calculated as 0.767 (
) and 0.696 (
). No pair reaches 0.8, which is considered an indicator of redundancy. This result is not affected by the chosen representative value. When the analysis is repeated using lower bounds, upper bounds, and the five-year averages in
Appendix C instead of midpoints, the strongest pair remains
and
in every case, while the corresponding correlations—0.774, 0.748, and 0.773—remain below the 0.8 threshold. Replacing the metric with a rank-based equivalent also does not alter this conclusion, since the maximum absolute Spearman correlation based on midpoints is 0.729.
Bivariate correlation captures only the relationship between two variables. Therefore, a criterion may be close to a linear combination of the others even though none of the individual correlations is high.
Table 16 addresses this multivariate aspect through auxiliary regressions in which each criterion is regressed on the remaining five criteria. The variance inflation factors (VIF) range from 1.98 for
to 7.25 for
, indicating that none of the criteria approaches the threshold of 10 traditionally used to indicate severe multicollinearity. Since only 16 observations are available for five regressors in each auxiliary regression,
is biassed upward, and the reported factors therefore overstate rather than understate collinearity, with the corresponding values obtained from Adj.
ranging from 1.32 to 4.83.
When the findings from
Table 15 and
Table 16 are considered together, it is observed that no single criterion duplicates the information provided by the others, either individually or collectively. However, the diagnostics do describe this specific sample, and the statistical relationship does not imply preferential dependence. The contribution of each criterion to the initial ranking is further tested with the weighting scenarios described in
Section 5.2. These results support using the additive aggregation of G-MUNRA in the present application.
5. Sensitivity and Comparative Robustness Analyses
The following section presents the sensitivity analysis of the ranking positions derived through the introduced grey hybrid methodology and validates the initial outcomes. For the sensitivity analyses, we explored the influence of variations in the model’s parameters and weight values on the initial ranking positions, first one parameter at a time and then jointly over the entire admissible parameter space. In addition, the validity of the initial outcomes was tested by comparing rankings from the developed grey decision support tool with those from existing grey-based models, deterministic (crisp) models, fuzzy logic-based models, and the GPCI Economy rankings. In all comparisons, we measured agreement between the two rankings using Spearman’s rank correlation coefficient.
5.1. Sensitivity Analysis—Variation in the Model Parameters , and
The introduced decision support model comprises three parameters: the grey importance coefficients , which aggregate expert opinions within G-RANCOM; the whitenization coefficient in the G-MUNRA procedure; and the aggregation coefficients , which govern the contributions of the three normalizations in the G-MUNRA model.
First, we develop 10 scenarios to analyze the effect of the grey importance coefficients, which determine the weights assigned to experts in the initial ranking results. In S1, all experts have the same level of importance, and the value “I” is assigned to all five experts as a linguistic variable. In S2–S5, the most important linguistic term, VI, is assigned to the remaining four experts, respectively, excluding the first expert. In S6–S10, we exclude one expert at a time.
Figure 2 shows the ranking results for all scenarios. As shown in
Figure 2, the 10 scenarios produce six different rankings. The Spearman correlation coefficients (SCCs) between the scenario rankings and the initial ranking results range from 0.9912 to 1.0000. In all scenarios, the first two and last three ranks remain unchanged, and in any given scenario, no city alternative changes its rank by more than one position.
Second, the whitenization coefficient
is set to 0.5 to obtain the initial ranking positions. In this part, ten scenarios are designed to evaluate the influence of
on the initial ranking results of G-MUNRA. In the first scenario, the
parameter is set to 0, while in the last scenario, it increases in increments of 0.10 until it reaches 1. The results of the scenarios associated with the
parameter are illustrated in
Figure 3. Although the 10 scenarios produce three different rankings, the SCCs between the scenario rankings and the initial ranking range from 0.9941 to 1.0000. In all scenarios, the top three and bottom three positions remain unchanged.
Third, nine scenarios are formulated to explore the impact of variations in the aggregation-coefficient vector
, which integrates the outcomes of three different normalizations subject to
. The nine scenarios are formed based on an augmented simplex-centroid design over the unit simplex: three pure vertices (S1–S3), three binary mixtures (S4–S6), and three axial points (S7–S9).
Figure 4 shows the ranking results from the nine formulated scenarios. As shown in
Figure 4, implementing the scenarios yields six separate rankings. The computed correlations between the initial ranking and the scenario rankings range from 0.9588 to 1.0000.
,
,
,
, and
are the five alternatives that retain their ranking positions in all nine scenarios. However, the alternative with the greatest variation in ranking position is
. Across all scenarios, this alternative’s ranking position varies between 4 and 10.
5.2. Sensitivity Analysis—Investigating the Effect of Changing Criterion Weights on the Initial Ranking Results
As a second step in the sensitivity analysis, the grey weights of the criteria are modified. In this context, the lower and upper limits of the grey weight for the first criterion (
) are reduced by the same proportion in four separate scenarios—at 25%, 50%, 75%, and 100%—and the reduced weight is distributed equally among the remaining five criteria in each scenario. The same procedure is repeated sequentially for the remaining five criteria, resulting in 24 scenarios (S1–S24). Because the grey weight of the criterion in question is zero in scenarios S4, S8, S12, S16, S20, and S24, those scenarios exclude this criterion from the assessment.
Figure 5 illustrates that the 24 scenarios produce 19 separate ordering patterns. However, the SCCs between the initial ordering and the scenario-derived orderings range from 0.8882 to 1.0000, with an average correlation of 0.9819. The positions of
and
, which have the highest performance, and
,
, and
, which have the lowest performance, remain unchanged across all scenarios. In contrast,
shifts between the fourth and tenth ranks, while
shifts between the sixth and 12th ranks, indicating that the movements are concentrated in the middle of the ranking.
5.3. Global Sensitivity Analysis
The sensitivity analyses in the previous two subsections vary the model’s parameters one at a time. Hence, these analyses do not account for combinations in which multiple parameters change simultaneously. For this reason, this subsection adds a global analysis. In each of 100,000 replications, the criterion weight vector is drawn from a flat Dirichlet distribution over the entire unit simplex so that every admissible weight vector is equally likely.
The whitenization coefficient in Equations (25)–(27) is drawn uniformly from the unit interval, and the aggregation-coefficient vector in Equation (28) from a flat Dirichlet distribution over its own simplex. The three aggregation coefficients weight the linear, vector, and non-linear normalization scores, respectively, so sampling them also lets the choice of normalization vary. Employing a single procedure corresponds to the vertices (1, 0, 0), (0, 1, 0) and (0, 0, 1); these are tested as separate scenarios in the aggregation-coefficient analysis of
Section 5.1. Each replication employs a crisp weight vector rather than a grey one. The performance matrix in
Table 10 is held fixed, and the whitenization coefficient carries its imprecision. The grey linguistic scale enters only via the weight bounds of
Table 9, which this design discards. We vary it in the second run below. The replications employ a fixed random seed, and the standard error of any frequency presented below does not exceed 0.0016.
Table 17 provides the resulting rank distributions, and
Table A9 of
Appendix D summarizes the complete rank acceptability matrix. The last two columns in
Table 17 are rank acceptability indices in the sense of stochastic multi-criteria acceptability analysis (SMAA). Column
is the first-rank index [
45], whereas column
sums the indices of the first three ranks, in line with SMAA-2 [
46]. Since the measures relating to the criteria are not modelled stochastically, the central weight vectors and confidence factors for this framework are not reported. The indices function as a layer of robustness, not as a ranking tool. The Spearman correlation between the ranking of a replication and the ranking of the G-MUNRA method has a mean of 0.860. It is at least 0.90 in 49.3% of the replications and at least 0.594 in 95% of them. The obtained ordering shows a correlation of 0.977 with the average ranks in
Table 17, so it is close to the ranking expected over the sampled parameter space. The first position is essentially limited to two cities:
and
account for 96.8% of all first ranks, with acceptabilities of 0.738 and 0.230. Apart from
(0.032), no city attains it in more than one replication in a thousand. The three lowest positions are held by the same three cities in 86.9% of the replications. The third and fourth are contested:
and
enter the first three ranks in 48.0% and 46.2% of the replications. The middle of the ranking is far less determinate. The 90% intervals of the cities ranked fifth to 13th span six to 11 positions and overlap almost completely. None of them enters the first three ranks in more than 7% of the replications.
Next, we explore the grey linguistic scale in
Table 2, first on its own. The three grey numbers are generalized to [0,
], [
, 1 −
] and [1 −
, 1], so that
= 1/3 returns
Table 2 itself. The scale degenerates as
approaches 0 or 1/2, which bounds the scan range. The G-RANCOM-MUNRA model is repeated for sixty-one values of
between 0.15 and 0.45. The criterion weights vary significantly, and the lower bound of
’s weight more than doubles over the interval. The ranking does not. Out of 61 calibrations, only two different rankings are obtained. The correlation with the reported ordering never drops below 0.9971, and no city shifts more than one position.
A second run of the same size restricts the weights to the grey bounds of
Table 9. The calibration parameter
of the grey linguistic scale is varied over [0.15, 0.45], so the bounds themselves move between replications, while the whitenization and aggregation coefficients continue to vary. Every uncertain element of the model is therefore perturbed simultaneously. The ranking is then almost fully identified.
is ranked first and
second in every replication, the three lowest positions are fixed, and the modal rank reproduces the reported position for all 16 city alternatives. The rank correlation averages 0.993 and never drops below 0.953.
The analyses above whiten the grey scores before ranking. The interval information can also be utilized directly to establish how firmly it supports each pairwise comparison. The grey aggregate score of a city is formulated before whitenization as the equally weighted sum of the three grey scores given in Equations (22)–(24).
For two grey numbers, the degree of possibility that
is at least as large as
is calculated as follows [
39]:
Here stands for the width of the corresponding interval. A value of 1 means full dominance, while 0.5 indicates indistinguishable intervals. None of the 120 ordered pairs implied by the reported ranking is reversed; the smallest value is 0.507, and 60 pairs are decided with a possibility of at least 0.90. Each of the five highest-ranked cities dominates each of the three lowest-ranked cities with a possibility of 1.000. Adjacent pairs are decided only marginally, with values between 0.51 and 0.76. The single exception is over , decided with a possibility of 1.000.
Both analyses agree on where the framework discriminates and where it does not. Wide gaps are reliable. is far ahead of the other cities, and the highest-ranked cities are far ahead of the lowest-ranked cities under every source of uncertainty examined here, including the complete removal of expert weights. Narrow separations are not reliable. Adjacent pairs are distinguished only marginally, and after the weights are released, cities ranked fifth through 13th freely change positions. Consequently, G-MUNRA’s initial ranking supports comparisons among well-separated cities but not among cities in adjacent positions.
5.4. Validation of Results—Comparison with Existing Grey MCDM Approaches
The comparative robustness analysis is conducted to assess the degree of consistency between the propounded grey-based decision-making methodology and several well-established grey MCDM approaches, namely G-EDAS [
47], G-PIV [
48], G-SAW [
49], G-SRP [
41], and G-WASPAS [
50].
Figure 6 shows that the six ranking techniques produce five diverse rankings (see
Table A10). However, in every method,
ranks first,
ranks tenth, and
ranks 16th. The SCC between the introduced grey decision support algorithm and other grey ranking approaches ranges from 0.9235 to 1.0000, with an average of 0.9706. SCC values are 1.0000 for G-SAW, 0.9853 for G-PIV, 0.9765 for G-SRP, 0.9676 for G-EDAS, and 0.9235 for G-WASPAS. The deviations observed in G-PIV, G-SRP, and G-EDAS are limited to one or two positions, whereas the deviations in G-WASPAS range from one to four positions.
5.5. Validation of Results—Comparison with the Crisp and Fuzzy Methods
To further analyze the consistency and external validity of the introduced grey-based model, the initial ranking outputs are benchmarked against three comparison frameworks: a crisp RANCOM-MUNRA hybrid methodology, a fuzzy LBWA–ROV hybrid methodology, and a fuzzy LBWA–PIV hybrid methodology. By offering alternative approaches to decision information processing and uncertainty modelling, these comparison procedures enable testing the robustness of the developed grey approach across different MCDM paradigms.
In the context of the application of the original (crisp) RANCOM-MUNRA hybrid methodology, a total of six models are designed based on annual data for the 2021–2025 period and average data for the same period. In the crisp-based models, the methodological steps of the RANCOM–MUNRA hybrid procedure are retained, and only the interval-valued inputs are replaced with point values. The results obtained from these models are illustrated in
Figure 7.
Figure 7 combines the rankings of the 16 city alternatives in the crisp-based models for each of the five years, the range of rankings for each alternative, and the rankings of the average-based crisp model and the introduced grey-based approach. Analysis of
Figure 7 shows that all six crisp-based models closely match the grey-based model. The Spearman correlation coefficients range from 0.9824 to 0.9971 for the mean-based and annual models. In all crisp-based models, the number of positions showing complete overlap is between nine and 14. While the deviations remain within a narrow range, the rankings of the alternatives do not change by more than two positions. Additionally, alternatives
,
,
, and
maintain their positions in all six models. However,
, the most volatile alternative, shifts between the 6th and 9th ranks over the five-year period.
Table A11 in
Appendix F reports the scores and rankings of the six crisp models.
Regarding the fuzzy-based methodologies, the Fuzzy Level-Based Weight Assessment (F-LBWA) [
51], the Fuzzy Proximity Index Value (F-PIV) [
52], and the Fuzzy Range of Value (F-ROV) [
53] are utilized. In this context, the F-LBWA procedure calculates the criterion importance weights from a second round of expert evaluation, in which each expert assigned each criterion a level of importance and an intra-level importance value using the questionnaire reported in
Table A12. In accordance with the F-LBWA procedure, the inter-expert values are first transformed into triangular fuzzy numbers. Next, the elasticity coefficient is assigned as
for each expert. Finally, the individual fuzzy weight vectors are aggregated by means of an unweighted Bonferroni mean operator with
. For the F-PIV and F-ROV ranking process, the fuzzy decision matrix is built by representing each input as a triangular fuzzy number, with boundaries corresponding to the minimum and maximum annual observations and a peak at the five-year average. Thus, the grey and fuzzy matrices are based on exactly the same observations, and no linguistic scale or predefined membership function is utilized at any stage. The elicitation responses, the resulting fuzzy weights, and the outputs of the F-PIV and F-ROV procedures are reported in
Table A13,
Table A14 and
Table A15 of
Appendix G. The ranking results obtained by applying these procedures are displayed in
Figure 8. According to
Figure 8, when considering the ranking results of the three models, minor changes are observed in the positions of city alternatives such as
,
,
,
,
, and
. However, the remaining 10 city alternatives show no changes in position.
5.6. Validation of Results—Comparison with the Published GPCI Economy Rankings
The final validity assessment compares the initial ranking with a benchmark derived independently of the proposed hybrid methodology. Given that the six criteria of this study constitute the indicators of the GPCI’s Economy function, the GPCI’s Economy function rankings for the 2021–2025 period (see
Table A16 of
Appendix H) are compared separately with the hybrid model’s rankings. The results of these comparisons are displayed in
Figure 9.
Figure 9a indicates each city’s position in the five published rankings, the range that these positions cover, and the ranking assigned by the introduced hybrid model, while
Figure 9b illustrates the variations between them.
Figure 9a shows that each city’s position in the five published rankings falls within a band of no more than three ranks, and that the proposed hybrid model’s ranking falls within this band for all 16 city alternatives. Additionally, we compute an average correlation of 0.9818 between the rankings, confirming that the developed methodology’s rankings closely align with the published rankings.
6. Discussion, Theoretical Contributions, and Policy Implications
This manuscript proposes and applies a grey-based multi-criteria decision-support tool to compare the economic performance of major cities across G7 economies. Results from the G-RANCOM procedure indicate a differentiated importance structure among the economic criteria used in the assessment. The prominence of market size (), economic vitality (), and business environment () suggests that city-level economic performance in advanced economies is shaped primarily by structural and capacity-related factors, namely economic scale, financial and corporate depth, and the quality of the business ecosystem, including access to skilled labour. In contrast, market attractiveness () and ease of doing business () play a comparatively more limited role in distinguishing cities that operate within broadly similar institutional and regulatory contexts, such as those of the G7 economies.
This weight structure has two policy implications. The first concerns what does not differentiate cities within this group. Market attractiveness and ease of doing business, which capture growth momentum, economic freedom, corporate taxation, and country risk, rank fourth and sixth among the six criteria, and ease of doing business has the smallest weight in the set. This is not because such conditions are unimportant, but because G7 cities already operate under broadly similar institutional and regulatory regimes, so marginal adjustments in tax rates or regulatory burden have limited power to separate one city from another. The second implication concerns what differentiates them. Market size, economic vitality, and business environment jointly account for roughly two-thirds of the total weight, and all three describe accumulated capacity rather than current policy settings, which indicates that positions in this ranking respond to sustained investment in scale, financial and corporate depth, and labour-market quality rather than to short-term measures. One qualification applies to both implications. The weights express the relative importance that an expert panel attaches to each dimension rather than a causal estimate of its effect, so they indicate what informed observers regard as decisive, not what a city should expect to gain from acting on a given dimension.
The G-MUNRA algorithm places New York, London, and Tokyo in the top three positions. An examination of the criterion-level rankings shows these cities lead across several economic dimensions. New York ranks first in market size (), economic vitality (), and business environment (), reflecting its economic scale, the depth of its financial markets, and its highly developed business ecosystem. London ranks first in human capital () and ease of doing business (), which indicates its labour market capacity and the institutional conditions supporting economic activity. Tokyo, third overall, ranks second in economic vitality () and human capital () and third in market size (), reflecting balanced performance across the most heavily weighted criteria. These patterns indicate that the leading cities owe their positions not to exceptional performance in a single dimension but to sustained, concurrent strength across several economically relevant criteria.
At the lower end of the ranking, Osaka, Milan, and Fukuoka occupy the last three positions, and the criterion-level pattern behind these positions is the inverse of the pattern observed at the top. Fukuoka is last of the sixteen cities in market size () and business environment (), the criteria ranked first and third by weight, and 15th in economic vitality (), which carries the second-largest weight. Milan is last in market attractiveness () and ease of doing business () and 14th in business environment (), and its only mid-table position, sixth in human capital (), belongs to one of the two least influential criteria. Osaka is the 15th in business environment (), 14th in market attractiveness (), and 13th in ease of doing business (), and its comparatively better standing in economic vitality () is not sufficient to compensate for them. Low overall positions in this framework therefore arise from simultaneous weakness in the heavily weighted structural criteria rather than from a single deficient indicator.
The global sensitivity analysis reported in
Section 5.3 qualifies these findings. The separation of New York from the remaining alternatives, and that of the highest-ranked from the lowest-ranked cities, remains stable over the entire admissible parameter space, whereas alternatives occupying adjacent positions are separated only marginally. Comparative statements about cities in adjacent positions should therefore be avoided.
The findings of this study offer theoretical implications for both the MCDM literature and the economic assessment of cities. First, the results confirm that urban economic performance is inherently multi-dimensional and cannot be adequately captured through single-indicator or composite-index approaches. The observed rankings indicate that high-performing cities achieve their positions through consistently strong performance across multiple economic criteria rather than dominance in a single dimension, reinforcing a central premise of MCDM theory: overall performance emerges from the joint contribution of interrelated criteria. Second, the study contributes methodologically to decision science by extending RANCOM and MUNRA to the interval grey-number domain. To the best of the authors’ knowledge, these formulations are the first adaptations of RANCOM and MUNRA in which both criterion weighting and alternative ranking use interval-valued information. By embedding grey numbers in both stages of the decision process, the proposed framework extends these methods to environments characterized by data imprecision and variability while preserving their original computational structure. Third, in contrast to the predominantly country-level focus of existing economic performance studies, the study introduces a city-level multi-criteria evaluation framework and provides a systematic MCDM-based comparison of economic performance across major global cities.
The findings also offer insights for urban policymakers and economic strategists in advanced economies. The ranking results indicate that improvements in urban economic performance are unlikely to be achieved through isolated or single-indicator interventions, and that cities benefit from policy approaches addressing several economic dimensions in a coordinated and mutually reinforcing manner. For cities at the lower end of the performance distribution, the results indicate that gradual and balanced progress across several dimensions is more likely to produce durable improvement than an attempt to raise performance rapidly in a single area.
In practical terms, a city administration can apply the framework in four steps. First, the city’s interval scores in
Table 10 are compared with those of the reference group to identify the criteria on which even the upper bound falls below the sample median, distinguishing a structural shortfall from a temporary one. Second, each shortfall is weighted by the corresponding criterion weight in
Table 9, which orders candidate interventions by the ranking movement that closing them would produce rather than by their political visibility. Third, the implied movement is compared with the rank acceptability indices in
Table A9, and where it does not exceed the uncertainty of the current position, the intervention should be justified on grounds other than benchmarking. Fourth, the exercise is repeated as new editions of the index appear, and the width of a city’s intervals is monitored in its own right, since a widening interval signals instability in the underlying position even when the midpoint is unchanged. Applied this way, the framework functions less as a league table than as a diagnostic instrument that shows where a city’s position is weak, how much a given improvement would matter, and whether the resulting movement can be distinguished from measurement uncertainty.
7. Limitations and Future Research
The proposed grey-based framework is based on seven assumptions. Each is stated below together with the risk it carries for the reported ranking, the evidence that bounds its influence, and the extension that would relax it.
The bounds of each grey number are the minimum and the maximum of five annual observations, so interval width is governed by extreme years rather than by measurement error, and the whitened scores of cities with volatile histories are correspondingly less determinate. The rank acceptability indices in
Section 5.3 limit this impact to the middle of the ranking, and intervals constructed from longer series or from percentile-trimmed bounds would mitigate it further.
Before ranking the alternatives, each interval is reduced to a single value, which presents cities with overlapping intervals in a definitive order. The degrees of possibility in
Section 5.3 indicate that no ordered pair implied by the reported ordering is reversed and that adjacent pairs are decided only marginally, and reporting the results in interval-dominance form throughout would remove the need for whitenization.
The criterion weights come from a five-member panel whose importance coefficients the authors assigned. Thus, a different panel or a different assignment rule could yield a different weight vector. The concordance reported in
Section 4.5.1 and the 10 expert-team scenarios in
Section 5.1 keep every city within one position of its reported rank, and a larger panel would allow these coefficients to be estimated rather than assigned.
The criterion set is confined to the GPCI Economy function. Therefore, research and development, accessibility, environment, cultural interaction, and livability dimensions are excluded, and the findings describe economic performance rather than overall urban performance. The scope is stated explicitly in
Section 1, and the external comparison in
Section 5.6 is against the Economy ranking rather than the overall index; extending the indicator system beyond the economic dimension would broaden the conceptual scope.
The study sample consists of the G7 cities covered by the index, a group that is institutionally homogeneous, which limits the transferability of the weight structure to emerging and developing economies. The rule in
Section 4.4 makes the sample exhaustive rather than selective within the G7, and applying the model to cities outside this group would directly test transferability.
The aggregation is additive and therefore assumes preferential independence, under which conceptually related criteria may compensate for one another. The diagnostics in
Section 4.6 report a maximum absolute correlation of 0.767 and a maximum variance inflation factor of 7.25, both below conventional thresholds, and a Choquet integral or an outranking model would be required in settings where interaction is documented.
All performance data come from a single source and enter the model after the index’s standardization and reverse-scoring procedures. Hence, any bias in those procedures propagates to the ranking.
Appendix A presents the orientation and transformation of each sub-indicator, and replicating the analysis with an independent indicator system would allow assessment of source dependence.
8. Conclusions
This study developed an integrated grey-based multi-criteria approach to benchmark urban economic performance under uncertainty and applied it to 16 G7 cities included in the Global Power City Index. Criterion importance was calculated using a grey extension of RANCOM, which converts the ordinal evaluations of a five-member expert committee into interval weights, and cities were ranked using a grey extension of MUNRA, which combines linear, vector, and non-linear normalization.
Two key findings emerge from the analysis. Market size, economic vitality, and the business environment account for approximately two-thirds of the total criterion weight. This suggests that the economic ranking among G7 cities is driven more by accumulated structural capacity than by financial and regulatory conditions, which are already similar across these cities. In the ranking, New York, London, and Tokyo occupy the top three positions, while Osaka, Milan, and Fukuoka occupy the bottom three. This is the result of strengths or weaknesses observed simultaneously across multiple highly weighted criteria, rather than performance in a single dimension in each case.
We analyzed the stability of these positions at three levels. Scenario analyses of the expert-team composition, the whitenization coefficient, the aggregation coefficients, and the criterion weights yield rank correlations between 0.89 and 1.00, and the two highest and three lowest positions remain unchanged under the expert-panel and criterion-weight scenarios. A global analysis of 100,000 replications, in which all model parameters vary jointly over their admissible ranges, assigns New York the first position in 74% of the replications and leaves the three lowest positions unchanged in 87%, whereas cities in adjacent positions are separated only marginally. Comparison with five established grey methods, with crisp and fuzzy counterparts, and with the published GPCI Economy rankings yields Spearman correlations between 0.92 and 1.00. The framework therefore supports comparisons between well-separated cities rather than between cities in neighbouring positions.
This work’s methodological contribution is the formulation of RANCOM and MUNRA in an interval grey environment, allowing criterion weighting and alternative ranking with interval-valued information while preserving the computational structure of the original methods. For urban policymakers, the framework provides a diagnostic tool rather than a league table, since it identifies the criteria on which a city’s position is structurally weak, orders candidate interventions by the ranking movement they would produce, and indicates whether that movement can be distinguished from measurement uncertainty. The assumptions on which these results rest, and the evidence bounding their influence, are set out in
Section 7.