Abstract
Multi-criteria group decision making (MCGDM) under linguistic uncertainty remains a fundamental challenge in applied mathematics, where decision makers seldom assign crisp numerical evaluations and frequently exhibit heterogeneous risk attitudes shaped by behavioural factors. An integrated mathematical framework, hereafter PLR-3WBC (Probabilistic Linguistic Regret-driven Three-Way Bayesian Consensus), is developed to systematically integrate four methodological components that have each been individually validated in the MCGDM literature: representation of decision information with explicit probability mass on linguistic terms; quantification of decision-maker regret and rejoice psychology under linguistic uncertainty; classification of alternatives into three actionable decision regions rather than a single-valued ranking; and group consensus reaching with credal weight aggregation. Each component has demonstrated its effectiveness in its respective domain; the present framework capitalises on their complementary strengths by embedding them within a single pipeline equipped with formal guarantees, an integration that has not been previously reported. The framework integrates five methodological components: probabilistic linguistic term sets (PLTS) for information representation; the Bayesian best–worst method (BBWM) for credal criterion weighting; a regret–rejoice value function adapted to the linguistic domain for behavioural evaluation; three-way decision (3WD) thresholds derived from a loss-function model for actionable classification; and a distance-based consensus reaching process with feedback mechanism for group convergence. A case study on age-friendliness evaluation of twelve aging urban residential communities under an indicator system of five dimensions and eighteen criteria, with four expert decision makers, demonstrates that PLR-3WBC delivers an actionable three-way classification, recovers a transparent group consensus, and produces rankings broadly consistent with classical TOPSIS, VIKOR, PROMETHEE-II, and BWM-TOPSIS (Spearman rank correlation exceeding 0.97), thereby confirming that the integrated framework preserves the ordinal reliability of these established methods, while additionally delivering three outputs that arise from the methodological integration: an actionable three-way classification enabling discrete budget-aligned decisions, credal weight intervals quantifying the depth of expert agreement on criterion importance, and a behavioural reordering of borderline non-dominated alternatives that reflects the loss-averse psychology of the decision panel and would remain hidden under single-method deployment. Sensitivity analyses with respect to the regret aversion coefficient, the loss function parameters, and the consensus threshold confirm that the qualitative classification is stable across a wide parameter envelope, supporting the practical deployment of PLR-3WBC in age-friendly community renewal programmes.
Keywords:
multi-criteria group decision making; probabilistic linguistic term sets; three-way decisions; consensus reaching; age-friendly community MSC:
90B50; 91B06; 03E72; 62C10
1. Introduction
Multi-criteria decision making (MCDM) provides a mathematical foundation for selecting, ranking, or classifying alternatives evaluated against multiple conflicting criteria, and has become a central tool of applied mathematics in operations research, engineering management, environmental planning, and social policy [1,2]. In adjacent fields, Fan et al. [3] have recently demonstrated that reduced-order modelling for high-dimensional trade-off analysis in nuclear engineering design shares a conceptual affinity with the dimensionality-management role of PLTS and three-way decisions, suggesting that the methodological demand for structured dimensionality reduction in multi-criteria evaluation extends well beyond the traditional MCDM community. Three layers of difficulty characterise practical MCDM problems and motivate the present work. The first layer is information uncertainty. In real decision settings, evaluators rarely supply crisp numerical assessments; instead, they prefer linguistic terms such as “good”, “somewhat poor”, or “very high” that admit degrees of membership and, more recently, explicit probability masses that quantify hesitation among adjacent terms [4,5]. The second layer is decision-maker psychology. Classical MCDM aggregations adopt expected-utility logic and ignore behavioural distortions such as regret and rejoice that are empirically robust in choice under uncertainty [6]. The third layer is the disjunction between the smooth real-valued ranking produced by methods such as TOPSIS, VIKOR, or PROMETHEE on the one hand and the discrete actionable classification (accept, reject, or defer) that decision makers actually need on the other hand [1,7,8,9].
The probabilistic linguistic term set (PLTS) of Pang, Wang, and Xu [5] addresses the first layer by attaching a probability vector to a linguistic evaluation, permitting expressions such as “the community is good with probability 0.6 and very good with probability 0.4”. Recent work has substantially extended this paradigm, including PLTS Choquet aggregation [10], score functions based on concentration degree [11], comprehensive surveys of probabilistic linguistic decision making [12], probabilistic uncertain linguistic VIKOR for risk assessment [13], and large-scale group decision making with public participation models [14]. Three-way decision theory of Yao [9] addresses the third layer by partitioning the universe of alternatives into a positive region for acceptance, a negative region for rejection, and a boundary region for deferred decision; recent extensions to Pythagorean, interval-valued fuzzy soft, and decision-theoretic rough set environments have proliferated rapidly [15,16,17]. Behavioural decision theory, in particular regret theory of Bell [6], addresses the second layer by modifying the value function to penalise outcomes that underperform a foregone alternative; its integration with three-way decisions and probabilistic linguistic information has emerged as an active research front in the last three years [18,19,20,21,22]. The Bayesian best–worst method (BBWM) of Mohammadi and Rezaei [23] provides a probabilistic group aggregation of pairwise comparisons that yields credal interval estimates of criterion weights, sharpening the classical best–worst method of Rezaei [24,25] and being recently extended to sustainability, spherical fuzzy, and review contexts [26,27,28]. Group consensus reaching processes [29,30,31,32] provide iterative feedback mechanisms that drive expert evaluations toward agreement.
Each of these five strands has been independently validated as an effective methodological tool in its respective domain: PLTS has been maturely applied in supplier selection [10], urban management [14], and risk assessment [13]; the Bayesian best–worst method has been successfully deployed in sustainability evaluation [26] and green supplier selection [27]; regret theory has been demonstrated to capture realistic behavioural distortions across diverse MCGDM settings [19,21,22]; three-way decisions have been proven effective in Pythagorean fuzzy [15], interval-valued [16], and evidence-theoretic [33] environments; and consensus reaching processes have been well developed with hierarchical feedback [31] and social network mechanisms [30,32]. However, systematic integration of these individually validated methods into a unified pipeline that exploits their complementary strengths for complex MCGDM problems has rarely been attempted. The PLR-3WBC framework proposed here establishes a three-way decision architecture and embeds each method at the stage where its proven capability is most needed (PLTS for information representation, BBWM for credal weighting, regret theory for behavioural adjustment, three-way thresholds for actionable classification, and consensus reaching for group convergence), thereby forming a coherent methodological integration whose effectiveness is validated through a case study with formal theoretical guarantees. The PLR-3WBC framework (Probabilistic Linguistic Regret-driven Three-Way Bayesian Consensus) is proposed here to close this gap. Beyond the methodological contribution, a substantive case study is provided on a problem of considerable demographic urgency: the evaluation of age-friendliness in aging urban residential communities. The World Health Organization [34] and subsequent analyses [35,36,37] have argued that the spatial, social, and service quality of aging-in-place neighbourhoods is a primary determinant of older residents’ wellbeing, and several recent MCDM studies on age-friendly parks, communities, and policy implementation have appeared [36,37,38]. Many countries face an imminent backlog of aging residential stock that requires prioritised retrofit, and the prioritisation problem is a canonical MCDM exercise complicated by all three layers of difficulty noted above.
The contributions of the present study are fourfold. At the methodological level, the study proposes a unified PLR-3WBC framework that integrates PLTS information, BBWM weighting, regret–rejoice value adjustment, three-way decision classification, and consensus reaching within a single mathematical pipeline. At the theoretical level, the study establishes a set of three theoretical results: a representation theorem proving order preservation of the regret-adjusted PLTS score under admissible monotone transformations (Theorem 1); a classification stability theorem giving explicit perturbation bounds under which the three-way classification is invariant (Theorem 2); and a consensus convergence theorem proving finite-time convergence of the feedback-driven consensus reaching process under mild monotonicity (Theorem 3). At the algorithmic level, the study provides a practical algorithm with explicit complexity bound (Algorithm 1 and Proposition 1), suitable for problem instances of realistic size. At the empirical level, the study delivers a comprehensive validation pipeline including an indicator system for age-friendliness evaluation grounded in WHO [34] and Buffel and Phillipson [35], a case study on twelve aging urban communities with four expert decision makers, and an extensive comparative and sensitivity analysis benchmarking PLR-3WBC against classical TOPSIS [1], VIKOR [7], PROMETHEE-II [8], and classical BWM-TOPSIS [24].
| Algorithm 1 PLR-3WBC: Probabilistic Linguistic Regret-driven Three-Way Bayesian Consensus |
|
Among the studies listed in Table 1, the work of Zhu et al. [20] is the closest predecessor, already combining PLTS, regret theory, three-way decisions, and consensus reaching. The present framework differs from [20] in three specific respects. (a) PLR-3WBC replaces the equal or entropy-based point-valued criterion weights of [20] with Bayesian best–worst weights that carry credal posterior distributions, enabling the analyst to quantify the depth of expert agreement on each criterion via credible intervals rather than a single point estimate. (b) The consensus reaching process in PLR-3WBC is equipped with a finite-time convergence guarantee (Theorem 3) with an explicit iteration bound, whereas [20] demonstrates convergence empirically but does not provide a formal bound. (c) The classification stability result (Theorem 2) supplies explicit perturbation bounds under which the three-way partition is invariant, a guarantee absent from [20].
Table 1.
Comparison with representative existing studies across six methodological dimensions.
The overall architecture of the five-stage pipeline and a comparative positioning against representative recent studies are deferred to the opening of Section 2 (Figure 1 and Table 1), so that the formal definitions introduced immediately afterward can be read in context. The remainder of this paper is organised as follows. Section 2 recalls the mathematical preliminaries. Section 3 develops the PLR-3WBC framework. Section 4 presents the case study. Section 5 reports sensitivity and comparative analyses. Section 6 discusses the findings and limitations. Section 7 concludes and outlines future work.
Figure 1.
Overall architecture of the PLR-3WBC framework. The three theoretical guarantees are indicated at their respective stages: Theorem 1 (order preservation) at the regret-adjusted scoring stage, Theorem 2 (classification stability) at the three-way decision stage, and Theorem 3 (finite-time convergence) at the consensus reaching stage.
2. Mathematical Preliminaries
This section recalls the four mathematical objects on which PLR-3WBC is built. Each subsection opens with a brief disciplinary positioning of the method, summarising its advantages over earlier formulations and surveying recent applications, before stating the formal definitions used in the remainder of the work.
Before introducing the formal definitions, the overall architecture of PLR-3WBC is depicted in Figure 1, and a comparative positioning against representative recent studies is provided in Table 1.
Figure 1 depicts the overall architecture and the role of each component on the problem considered here. Probabilistic linguistic term sets resolve the representation of hesitant qualitative judgements about each community; the Bayesian best–worst method converts the experts’ pairwise comparisons into credal weight intervals on the eighteen criteria; the regret–rejoice value function injects the loss-averse psychology of the panel into the evaluation; the three-way decision thresholds convert continuous preferences into the accept, defer, and reject actions that retrofit programme managers require; and the consensus-reaching process reconciles divergent expert opinions before aggregation. The five components are arranged sequentially, the output of each stage feeding the next, so that the integration is realised as a single mathematical pipeline rather than a loose collection of tools. Among the studies listed in Table 1, the work of Zhu et al. [20] is the closest predecessor; the specific distinctions are noted at the end of Section 1.
2.1. Probabilistic Linguistic Term Sets
Among the many extensions of fuzzy sets developed for qualitative evaluation, the probabilistic linguistic term set (PLTS) is distinctive in simultaneously carrying a discrete set of admissible linguistic terms and an explicit probability mass over those terms, which allows hesitation to be quantified more directly than hesitant fuzzy linguistic term sets [4] permit. In the last three years, PLTS-based models have been deployed in domains as varied as supplier selection [10], large-scale public participation in urban management [14], teaching reform plan evaluation [13], sustainable supplier selection across time horizons [39], and large-group emergency decision making under group pressure [40], indicating that the PLTS apparatus is becoming the de facto standard for linguistic group decision problems with probabilistic hesitation. The version of Pang, Wang, and Xu [5] is adopted below.
Let be a discrete linguistic term set with odd cardinality , where denotes the neutral assessment, the highest, and the lowest. The standard choice yields seven terms ranging from (very poor) to (very good) and is adopted throughout the present study.
Definition 1
(Probabilistic linguistic term set, adapted from [5]). A probabilistic linguistic term set on S is an expression
where is the cardinality and is the probability assigned to term . When , the residual probability corresponds to ignorance and is normalised prior to aggregation by the rescaling .
Definition 2
(PLTS expectation and score function). For a normalised PLTS on , the expected linguistic index is
and the normalised score in is
Definition 3
(Hamming distance between PLTS, adapted from [11]). Let and be two normalised PLTS on S, written
and expanded to a common length by appending null terms with zero probability. The Hamming distance is
The Euclidean distance is
Both and equal zero if and only if and are identical after normalisation.
Remark 1
(Scope of the PLTS Hamming distance). The Hamming distance in Definition 3 operates on the product and therefore captures differences in the probability-weighted linguistic indices rather than the full distributional shape. Two PLTSs with identical expected values but different dispersions (for example, and ) may receive after expansion, because the term-by-term products coincide once ordered. This property is consistent with the score function in Definition 2, which aggregates via expectation, but it implies that the consensus metric in Equation (20) measures agreement on expectation-level assessments rather than on the full probability profiles. When distributional shape is decision-relevant, a Wasserstein or Kullback–Leibler divergence on the probability vectors would be more appropriate; the present choice trades distributional resolution for computational tractability in the consensus loop, and the sensitivity analysis of Section 5.5 confirms that the classification is robust to the resulting approximation.
2.2. Bayesian Best–Worst Method
The classical best–worst method [24,25] produces consistent criterion weights with fewer pairwise comparisons than the analytic hierarchy process, and it has therefore become one of the most widely adopted weighting techniques of the last decade, with extensive methodological surveys confirming its consistency-driven appeal [28]. Its Bayesian extension by Mohammadi and Rezaei [23] replaces point-valued weights with credal posterior distributions, allowing the analyst to quantify the depth of group agreement on each criterion via credal ordering. Recent applications include sustainability evaluation of alternative aircraft models [26] and green supplier selection under spherical fuzzy sets [27], confirming that BBWM is the preferred tool when several experts provide pairwise preference data with non-trivial disagreement.
The classical best–worst method (BWM) of Rezaei [24,25] elicits two pairwise comparison vectors per decision maker, namely the best-to-others vector and the others-to-worst vector , each entry being an integer in . The weight vector is recovered by minimising
which is reformulated as a linear program [25]. Mohammadi and Rezaei [23] extend BWM to a Bayesian setting suitable for group decision making. Given K decision makers each providing , the Bayesian best–worst model assumes
with and at the group level. The aggregate weight inherits a posterior distribution that is sampled by Markov chain Monte Carlo, yielding both a point estimate and a credible interval per criterion. The credal ordering relation defined in [23] declares criterion superior to at confidence level if , where denotes the elicited data.
2.3. Regret Theory
Compared with classical expected utility aggregation, regret theory and the closely related prospect theory better describe the asymmetric psychology of decision makers who are sensitive to losses relative to foregone alternatives [6,41]. In the past three years, regret-theoretic adjustments have been integrated into a wide range of MCGDM frameworks, including probabilistic dual hesitant fuzzy decision making with behavioural distance measures [19], quantum group decision making with grey relational degree [21], ELECTRE III under probabilistic interval-valued intuitionistic hesitant fuzzy information [22], large-group emergency decision making with group pressure models [40], and three-way group decisions in dual hesitant fuzzy environments for water-supply alternative selection [42], showing that the behavioural augmentation of standard MCDM aggregation is now a mainstream methodological commitment.
Regret theory [6] modifies the value of an outcome by an additive regret–rejoice term capturing comparison with a foregone alternative. For a realised value v and a foregone reference value , the regret–rejoice value function is
where is the regret aversion coefficient. For , R is positive and represents rejoice; for , R is negative and represents regret. The function is concave in gains and convex in losses, consistent with the empirical evidence on behavioural choice [41]. The reference value is commonly set to the maximum of the remaining alternatives [43]. The behavioural effect of the coefficient is illustrated in Figure 2: as grows, the value function steepens, deepening the convex penalty on outcomes below the reference and flattening the concave reward on outcomes above it, so that a larger encodes stronger loss aversion and sharpens the discrimination between an alternative and its foregone competitors.
Figure 2.
Family of regret–rejoice value functions plotted against the difference for representative values of the regret aversion coefficient .
2.4. Three-Way Decisions
Three-way decision (3WD) theory extends classical two-way (accept/reject) decisions by introducing a third deferred decision region whose threshold is grounded in Bayesian risk minimisation, and it has been shown to outperform two-way decisions when the cost of a type I or type II error is high [9]. Within the past three years, 3WD has been combined with Pythagorean fuzzy ideal solutions [15], interval-valued fuzzy soft sets [16], decision-theoretic rough sets for dilemma reasoning and conflict resolution [17], evidence theory under hesitant fuzzy linguistic input [33], and regret-theoretic models in dual hesitant fuzzy environments [42], confirming that the trisection paradigm is now the preferred device when a discrete and risk-aware action set is required from MCGDM output.
The three-way decision (3WD) model of Yao [9] partitions the universe of alternatives into three regions based on a conditional probability that an alternative with evidence e belongs to the favourable class X, together with two thresholds . The three regions are
Alternatives in are accepted; those in are rejected; those in are deferred for further investigation. The thresholds are derived from a loss function , where are the losses incurred by accepting, deferring, and rejecting an alternative that truly belongs to X, and the analogous losses when the alternative does not belong to X. Under the natural ordering and , Bayesian risk minimisation [9] yields
2.5. Consensus Reaching in Group Decision Making
When several experts evaluate the same alternatives, point aggregation by arithmetic or geometric means is sensitive to outliers and conceals genuine disagreement; iterative consensus reaching processes (CRPs) with explicit feedback mechanisms have therefore become the standard remedy [29]. Recent CRP research integrates social network analysis with herding behaviour modelling [30], hierarchical feedback for hesitant fuzzy linguistic preference relations [31], parallel dynamic feedback under social network analysis [32], hesitant fuzzy CRP for large-scale group decision making [44], and stochastic multi-criteria acceptability for social-network PLTS groups [45]. A distance-based CRP with an identification rule and convex combination feedback is adopted, whose explicit form is given in Section 3.5.
3. The PLR-3WBC Framework
This section develops the PLR-3WBC framework. The construction proceeds in five stages, paralleling the five modules of Figure 1. Stage 1 collects linguistic probabilistic evaluations from the expert panel. Stage 2 produces credal criterion weights via BBWM. Stage 3 constructs a regret-adjusted score matrix. Stage 4 derives 3WD thresholds and classifies alternatives. Stage 5 monitors consensus and triggers feedback if consensus is below the prescribed threshold.
3.1. Stage 1: PLTS Decision Matrix Construction
Let denote the set of m alternatives, the set of n criteria, and the set of K expert decision makers. Each expert supplies an PLTS evaluation matrix , where is a PLTS on the linguistic scale S representing the assessment of alternative on criterion by expert . After normalisation according to Definition 1, the group-level decision matrix is obtained by entry-wise weighted averaging,
where ⊎ denotes the PLTS union with probability re-aggregation [5] and is the relative authority weight of expert , set uniform in the absence of prior information.
3.2. Stage 2: BBWM Credal Weighting
Each expert provides a best-to-others vector and an others-to-worst vector over the n criteria. The Bayesian model in Equation (7) is inferred via 10,000 MCMC iterations with 2000 burn-in samples. The output is a posterior sample of the aggregated weight vector, from which the point estimate and the per-criterion 95 percent credible interval are extracted. The credal ordering relation is computed pairwise, being estimated by the empirical frequency in the posterior sample.
3.3. Stage 3: Regret–Rejoice Adjusted PLTS Score Matrix
Let denote the normalised PLTS score from Equation (3) of alternative on criterion . The reference value on criterion is the criterion-wise maximum across the remaining alternatives,
The regret-adjusted score is
where blends the raw score with the regret–rejoice adjustment and is the regret aversion coefficient. The aggregated regret-adjusted preference of is
The regret aversion coefficient and the blending parameter are the two behavioural parameters of the framework. The value adopted in the case study is drawn from the empirical calibrations reported in the regret-theoretic MCDM literature: Zhu et al. [18] use for probabilistic linguistic environments, Lei et al. [42] adopt in their dual hesitant fuzzy three-way decision study, and Peng and Yang [43] report stable orderings for . The blending parameter assigns 30 percent of the adjusted score to the regret–rejoice term and 70 percent to the raw PLTS score, following the convention in [18,20] that the behavioural correction should moderate rather than dominate the raw assessment. A full sensitivity analysis with respect to is provided in Section 5.1; to address the concern that is also a free parameter, a complementary sensitivity analysis with respect to is added in Section 5.2.
Theorem 1
(Order preservation under monotone transformations of the regret coefficient). Let be defined by Equations (15) and (16) with criterion weights summing to one. For any pair of alternatives such that for all with strict inequality on at least one criterion, the inequality holds for every and every . Consequently, the regret-adjusted ranking is Pareto consistent: a Pareto dominant alternative is never demoted below a Pareto dominated one by the regret adjustment.
Proof.
Fix and . By Equation (14), and similarly for . When for all j with strict inequality on some criterion, examination of the four cases for the relative position of together with the monotonicity of in its first argument shows that for every j, with strict inequality on at least one criterion. Since for all j, summation yields . □
Remark 2
(Behavioural reordering of non-dominated alternatives). Theorem 1 guarantees that Pareto consistency is preserved by the regret adjustment but does not preclude reordering of mutually non-dominated alternatives, which is precisely the behavioural content of the framework: among alternatives that are incomparable in the dominance order, the regret-adjusted score reflects the loss-averse psychology of the decision panel.
3.4. Stage 4: Three-Way Classification with Loss-Function Thresholds
A loss function satisfying and is elicited from the consortium of decision makers. The thresholds follow from Equation (12). To map the regret-adjusted preference to the unit interval for subsequent three-way classification, a normalised preference degree is defined as
with and . This min–max linear rescaling is the standard normalisation in the MCDM literature for converting heterogeneous preference values to a common scale [1,46,47]; it preserves the rank order and the relative spacing of the values and is therefore compatible with the Pareto consistency guaranteed by Theorem 1. Importantly, is a normalised preference degree, not a frequentist conditional probability in the classical statistical sense; the three-way thresholds derived from Equation (12) partition the preference scale in the same way that Yao’s model [9] partitions a probability scale, and the loss function semantics are preserved because the monotonicity of the partition depends only on the order and separation of , not on a probabilistic interpretation. Two alternative calibrations were examined as a robustness check. A logistic calibration , with fitted to match the quartiles of and the sample mean, produces the same three-way classification for all twelve communities under the baseline parameters; an isotonic regression calibration, trained on a leave-one-out split, likewise preserves the partition. These checks are reported in Section 5.3. The three-way classification is then
Theorem 2
Proof.
The calibration in Equation (17) is Lipschitz with constant . A perturbation of magnitude in the preference vector induces a perturbation of at most in the conditional probability. If this perturbation is no larger than , the conditional probability remains on the same side of both thresholds, preserving the classification. □
3.5. Stage 5: Consensus Reaching Process
Define the group consensus level as
which lies in with corresponding to perfect agreement. A consensus threshold is fixed by the analyst (typically ). If , the feedback mechanism identifies the expert-criterion-alternative triple with the largest contribution to disagreement,
and recommends a revised PLTS for expert on cell , computed as the convex combination with the PLTS median across the remaining experts and the feedback intensity. The feedback intensity is adopted following the mainstream range reported in the consensus reaching literature: Dong et al. [29] recommend moderate intensities that balance convergence speed against respect for expert autonomy; Li et al. [31] adopt in their hesitant fuzzy linguistic consensus model; and Zhou et al. [32] report stable consensus trajectories for . A value at the lower end of this range is preferred here because the four-expert panel spans diverse professional backgrounds, and aggressive feedback () risks overriding substantive disciplinary judgements. The sensitivity of the final classification to is bounded indirectly by the consensus threshold analysis of Section 5.5, which shows that the ranking is stable across a wide range of final values induced by different -convergence speeds.
Theorem 3
(Finite-time convergence of the consensus reaching process). Let denote the group disagreement at iteration t, and let the convex-combination feedback rule be applied with intensity , using the identification rule of Equation (21) and the PLTS median as target.
- (i)
- (Monotonicity.) for all .
- (ii)
- (Geometric contraction.) Define the disagreement concentration ratiowhere is the per-cell disagreement contribution. The identification rule guarantees by the pigeonhole principle; in practice, disagreement concentration yields κ substantially above this lower bound. Under this definition, for all .
- (iii)
- (Iteration bound.) The consensus threshold is reached withiniterations.
Proof.
(i) The convex combination moves toward the component-wise median. Because is convex in its first argument and the median minimises the sum of Hamming distances, the total pairwise disagreement sum is non-increasing under this operation, giving .
(ii) The identification rule of Equation (21) selects the triple with maximum . Since is the average of all cell contributions, the maximum cell contributes at least , so . The convex combination reduces by a factor of at least , yielding the stated geometric bound.
(iii) Iterating the contraction inequality and solving for the first t with gives the ceiling expression. □
Remark 3
(On the assumptions underlying Theorem 3). Once the contractive property of the feedback rule is established, finite-time convergence follows by a standard geometric-series argument, which might suggest that the theorem embeds its conclusion partially within its assumptions. The following clarification is therefore warranted. The contraction in part (ii) is not assumed; it is derived from two independent and verifiable premises: (a) the convexity of the Hamming distance in its first argument, which is a metric property of Definition 3, and (b) the identification rule of Equation (21), which selects the cell of maximum disagreement and thereby guarantees the pigeonhole lower bound . The substantive content of the theorem is therefore the explicit and computable iteration bound (23), which depends on the observable quantity κ and the analyst-chosen parameters ζ and , enabling a practitioner to predict—before running the consensus loop—how many feedback rounds will be needed. This distinguishes the result from a bare existence-of-convergence statement. The quantity κ is defined by Equation (22) as a deterministic infimum over the iteration history; its empirical value in the case study is a measurement, not a free parameter, and the close agreement between the predicted contraction factor and the observed ratio provides empirical support for the tightness of the bound.
3.6. Algorithm
Algorithm 1 consolidates the five stages.
Proposition 1
(Computational complexity). Algorithm 1 runs in time and space, where is the average PLTS cardinality. For typical problem sizes (, , , , ), the algorithm completes in under five seconds on a standard desktop processor.
Proof.
The consensus loop has at most iterations; each iteration recomputes the pairwise distance sum and the cell-wise disagreement contributions at cost . The MCMC sampling for BBWM costs for S posterior samples over K experts and n criteria, with quadratic per-step cost from the Dirichlet update. The remaining stages are . The space cost follows similarly. □
Remark 4
(Implementation details). The PLR-3WBC algorithm was implemented in Python 3.11. The BBWM posterior sampling used the PyMC 5.10 library with the No-U-Turn sampler [23]; PLTS operations, regret-adjusted scoring, consensus iteration, and three-way classification were coded in NumPy 1.26. The comparative methods (TOPSIS, VIKOR, PROMETHEE-II, BWM-TOPSIS) were implemented from their original formulations using the same NumPy environment. All experiments were executed on a desktop workstation with an Intel Core i7-13700 processor and 32 GB RAM running Ubuntu 22.04. The wall-clock time for the full PLR-3WBC pipeline, including MCMC samples and four consensus iterations, was 3.8 s; repeating the pipeline ten times yielded a mean time of 3.9 s with standard deviation 0.2 s.
Proposition 2
(Pareto optimality of the acceptance region). Let be the acceptance region produced by Equation (18). Under Theorem 1, no Pareto dominated alternative is in unless every alternative that dominates it is also in .
Proof.
Suppose Pareto dominates and . By Theorem 1, , hence by Equation (17), so . □
4. Case Study: Age-Friendliness Evaluation of Aging Urban Residential Communities
This section applies PLR-3WBC to a realistic case study involving twelve aging urban residential communities in a provincial capital city in eastern China, evaluated by a panel of four expert decision makers from different professional backgrounds. The case is constructed to be representative rather than identifying; community identifiers are anonymised and panel composition is reported only at the role level. All evaluations reported in this study are based on real field data collected from twelve existing residential communities in a single provincial capital city in eastern China. The city name is withheld to preserve the anonymity of the participating communities and experts, in compliance with the data-sharing agreement signed with the municipal housing authority.
4.1. Indicator System Construction
The indicator system follows the World Health Organization age-friendly cities framework [34], the manifesto of Buffel and Phillipson [35], and recent MCDM-based age-friendliness studies that emphasise multi-dimensional indicators across physical, social, service, safety, and digital aspects [36,37,38], adapted to the residential community scale. Five dimensions and eighteen criteria populate the system, summarised in Table 2.
Table 2.
Indicator system for age-friendliness evaluation of aging urban residential communities.
The expert panel comprises four members: a senior urban planner, a gerontology researcher, a community service administrator, and a retrofit design engineer. Each expert provided PLTS evaluations on the seven-term linguistic scale , with denoting “very poor” and “excellent”. In addition, each expert provided BWM pairwise comparison vectors for the eighteen criteria. A panel of four experts is consistent with the Bayesian best–worst method literature, where Mohammadi and Rezaei [23] demonstrated BBWM with panels of three to five experts, and subsequent BBWM applications have routinely employed four-to-six-expert panels [26,27]. The PLTS consensus literature likewise reports case studies with three to five experts as the standard configuration for problems of this scale [18,20,42]; the feedback-driven consensus process of Section 3.5 is specifically designed to reconcile inter-expert divergence and thus compensates for the small panel size by enforcing agreement before aggregation.
4.2. Data Collection and Initial Disagreement
The twelve communities, denoted , were not drawn at random but selected by a stratified purposive scheme designed to span the principal sources of heterogeneity in the city’s aging residential stock. Three stratification variables were controlled: construction era (early 1980s, the 1990s, and the early-to-mid 2000s), spatial form (multi-storey walk-up, mid-rise with partial elevator coverage, and enclosed gated estate), and demographic intensity, measured by the share of residents aged 60 and above. Within each stratum the communities were chosen so that the sample jointly covers floor areas from 1.8 to 6.4 hectares and aging ratios from 14 to 38 percent of total residents, the empirical range observed across the district’s retrofit-eligible stock. The design therefore targets representativeness of variation rather than statistical estimation of a population mean: by deliberately including the extremes of construction era, building form, and density, the sample exercises the full operating range of the indicator system and stresses the classification thresholds, which is the property relevant to validating a prioritisation method rather than conducting a census.
Evaluations were gathered over six weeks through a two-channel protocol. First, each of the four experts conducted on-site field visits to all twelve communities, recording objective evidence against every criterion (ramp slopes and handrail continuity for C1.2, illuminance readings for C1.3, walking distance to the nearest medical facility for C3.1, camera-coverage maps for C4.1, and so on) against a standardised checklist. Second, the field evidence was combined with structured interviews of community administrators and resident representatives to capture the service-level criteria (the C2 group, C3.4, and the C5 group) that are not directly observable. Each expert then translated the pooled evidence for every alternative-criterion cell into a probabilistic linguistic term set on the seven-term scale S, assigning probability mass to adjacent terms wherever the evidence was mixed or incomplete; the residual mass permitted by Definition 1 was used to record explicit ignorance and was renormalised before aggregation.
The initial group consensus level computed from Equation (20) is , below the operating threshold , which triggers the feedback mechanism described in Section 3.5. Two limitations of the sample are acknowledged and addressed within the framework. The sample is confined to a single provincial capital, so the absolute weights and thresholds are city-specific; the sensitivity analysis of Section 5 therefore reports how the classification responds to changes in the behavioural and loss parameters rather than treating any single setting as canonical. The sample is also modest in size relative to the city’s full stock; this is mitigated by the purposive coverage of the heterogeneity strata above and by the consensus-reaching process, which suppresses the influence of any single outlying evaluation before classification. To keep the presentation transparent, the four illustrative communities of Table 3 are an excerpt drawn from that full data set rather than the entire evidence base.
Table 3.
Excerpt of the initial group PLTS matrix (communities , , , on criteria C1.1, C2.2, C3.1, C4.1, C5.2). PLTS notation uses linguistic indices.
A representative excerpt of the initial group PLTS matrix on five criteria for four illustrative communities is reported in Table 3 after the first expert aggregation step but before consensus iteration.
4.3. Consensus Reaching Iteration
Application of the feedback mechanism in Equation (21) with intensity converges in four iterations to a final consensus level . The trajectory of is reported in Table 4. As the table shows, the consensus level rises monotonically from to , first exceeding the threshold at iteration 3 (); the per-iteration increment decays steadily from to , the diminishing-returns pattern characteristic of a contractive feedback process and the empirical basis for the contraction estimate reported below.
Table 4.
Group consensus level across feedback iterations, with feedback intensity .
To verify that the feedback mechanism produces substantively plausible revisions rather than merely driving numerical convergence, the two largest revisions across the four iterations are examined. In Iteration 1, expert (retrofit design engineer) revised cell from toward the group median , yielding a revised PLTS of . Community is a 1980s walk-up with partially repaired ramps; the original assessment of “very poor” reflected the engineer’s strict structural standard, whereas the remaining experts, who weighed observed handrail continuity alongside slope measurement, assigned a moderately less severe evaluation. The revised cell thus shifts one-half step upward on the linguistic scale while preserving the negative orientation, consistent with the mixed field evidence. In Iteration 2, expert (gerontology researcher) revised cell from toward , yielding . Community reports a home-care service penetration rate that the gerontologist initially evaluated favourably on the basis of official statistics, whereas the field-visit evidence from the remaining experts indicated lower actual utilisation; the revision accordingly anchors the assessment closer to neutral, a correction consistent with the discrepancy between nominal and effective service coverage observed during the site visits. Both revisions preserve the qualitative direction of the original assessment and adjust magnitude within one linguistic step, supporting the plausibility of the feedback-driven consensus process.
The empirical contraction factor , with , is approximately per iteration. The disagreement concentration ratio of Equation (22) is estimated empirically as , well above the pigeonhole lower bound , reflecting the fact that inter-expert disagreement is concentrated in a small number of criterion-alternative cells rather than spread uniformly. Substituting into Theorem 3(ii) gives a predicted contraction factor of , in close agreement with the empirical value; the iteration bound of Theorem 3(iii) yields , consistent with the threshold first being exceeded at Iteration 3 ().
4.4. Bayesian Best–Worst Weights
The four experts identified C3.3 (emergency response button coverage) as the best criterion and C5.3 (digital information service) as the worst, and provided BWM comparison vectors. The BBWM model in Equation (7) was fitted using MCMC samples after a 2000 burn-in. The aggregate weights and 95 percent credible intervals are reported in Table 5, ordered by point estimate. As the table shows, the weights span roughly a 3.5-fold range, from C3.3 (emergency response button coverage, ) down to C5.3 (digital information service, ). The credal ordering between adjacent mid-ranked criteria is weak (posterior probability of correct ordering between 0.55 and 0.70), as discussed below. In a Bayesian weighting context, the meaningful distinction is between well-separated weight clusters, here the top cluster (C3.3, C1.1, C3.1, C4.2) and the bottom cluster (C5.1, C4.1, C5.3), rather than a total rank order, a point that single-point weighting methods cannot surface [23,28]. The health-and-care dimension D3 supplies three of the four most heavily weighted criteria (C3.3, C3.1, C3.4), and together with the leading physical-environment criterion C1.1 () the safety-critical and accessibility criteria dominate the weight budget, whereas the smart-and-digital criteria C5.1–C5.3 occupy the lower tail. The credible intervals are tight at the two extremes and widen across the mid-ranked band, the uncertainty structure examined next.
Table 5.
Bayesian best–worst aggregate weights with 95 percent credible intervals for the eighteen age-friendliness criteria. Criteria are listed in descending order of posterior mean weight. Adjacent credible intervals overlap for mid-ranked criteria, so the meaningful output of Bayesian weighting is the identification of well-separated weight clusters rather than a total order; see Figure 3, the accompanying discussion, and the general recommendation of [23,28].
The full posterior densities are displayed as violin plots in Figure 3: the distributions for the top-ranked criteria C3.3, C1.1, and C3.1 sit well to the right and barely overlap the cluster of low-weight criteria around C5.3, so the credal ordering between distant criteria is sharp—the posterior probability that exceeds . By contrast, the 95 percent credible intervals of the mid-ranked criteria C1.4, C2.1, C5.2, C4.3, and C2.3 overlap pairwise, and the corresponding adjacent-pair credal probabilities lie between and . The precise ordering within this mid-ranked band therefore carries less information than the overall weight magnitude, while the broad ordering between distant criteria is well-identified.
Figure 3.
Posterior density violin plots of the eighteen criterion weights from the Bayesian best–worst analysis with MCMC samples after a 2000 burn-in.
The dimension-level structure of the twelve communities is summarised by the radar profiles of Figure 4, which aggregate the criterion scores within each of the five dimensions D1–D5. Communities and trace large, well-balanced pentagons that score highly on every dimension, whereas and contract sharply on the heavily weighted health (D3) and physical environment (D1) dimensions; the intermediate communities form partially indented profiles that are strong on some dimensions and weak on others. These shapes foreshadow the three-way classification obtained in Section 4.5, where balance across the high-weight dimensions, rather than excellence on any single one, proves decisive.
Figure 4.
Five-dimensional radar profiles of the twelve communities across the dimensions D1 physical, D2 social, D3 health, D4 safety, and D5 digital.
4.5. Three-Way Classification Results
With regret aversion coefficient , blending parameter , and loss function (using the natural ordering , , , , , ), the thresholds derived from Equation (12) are
The regret-adjusted preferences and three-way classifications are reported in Table 6, which lists, for each community, the regret-adjusted preference , its calibrated conditional probability , the resulting region, and the within-class rank. The preferences span : the six communities with enter POS, the three with enter NEG, and the three intermediate communities () form the boundary region awaiting further evidence.
Table 6.
Regret-adjusted preferences , calibrated conditional probabilities , three-way classifications, and within-class rankings for the twelve communities, , , , .
As shown in Figure 5, the classification offers immediate managerial value: six communities (, , , , , ) fall in the POS region and warrant immediate retrofit, three (, , ) lie in the BND region and require further investigation, and three (, , ) occupy the NEG region and are not prioritised under the current budget cycle. The within-class ranking, also annotated in Figure 5, provides a fine-grained order for sequencing the retrofit programme.
Figure 5.
Three-way classification of the twelve aging communities into positive, boundary, and negative regions with within-class ranking labels.
5. Sensitivity and Comparative Analysis
5.1. Sensitivity to the Regret Aversion Coefficient
The regret aversion coefficient controls the behavioural strength of the regret–rejoice adjustment. The regret-adjusted preference is evaluated across the grid
with corresponding to the risk-neutral baseline. The Spearman rank correlation of the resulting orderings with the operating ranking at is reported in Table 7.
Table 7.
Sensitivity of the ranking to the regret aversion coefficient .
As Table 7 reports, the ranking exhibits high stability across the entire envelope, with Spearman correlation against the operating value remaining above 0.93. At (risk neutral), two borderline communities shift class (one from BND to POS, one from BND to NEG); at (high regret aversion), two communities again shift class. This is the behavioural reordering of non-dominated alternatives noted after Theorem 1. Specifically, at (risk neutral), community moves from POS to BND and community moves from BND to NEG; at (high regret aversion), community moves from BND to POS and community moves from NEG to BND. All four shifts involve adjacent-class transitions of communities whose values lie within 0.03 of the threshold, which is consistent with the classification stability bound of Theorem 2: the classification gaps for these communities are narrow, rendering them sensitive to the regret-induced score perturbation.
To address whether any reordering violates the Pareto consistency guaranteed by Theorem 1, a pairwise dominance check was conducted across all community pairs. No community is Pareto dominated on all eighteen criteria by any other community; every pair is mutually non-dominated on at least one criterion. Consequently, the Pareto consistency condition of Theorem 1 is trivially satisfied (there are no dominated–dominator pairs whose relative ordering could be reversed), and all observed rank changes across the grid are permissible reorderings of non-dominated alternatives, which is the intended behavioural content of the regret adjustment. Proposition 2 further guarantees that no Pareto dominated alternative can enter POS without its dominator; since no dominance relationships exist in this case study, the guarantee holds vacuously but would become operative in problem instances with Pareto dominated alternatives.
5.2. Sensitivity to the Blending Parameter
The blending parameter controls the relative weight of the regret–rejoice term in the adjusted score of Equation (15). Setting reduces PLR-3WBC to a regret-free framework, while would assign the entire adjusted score to the behavioural term. The regret-adjusted preference is evaluated across the grid
with held constant. The Spearman rank correlation with the baseline ordering at and the number of three-way class changes are reported in Table 8.
Table 8.
Sensitivity of the ranking to the blending parameter , with fixed.
The ranking is stable across the entire envelope, with Spearman correlation remaining above 0.94. At , where the regret term is absent, two borderline communities shift class ( from POS to BND and from BND to NEG), replicating the pattern observed at in Table 7, as expected, since both configurations annihilate the behavioural correction. At , two communities again shift, indicating that an excessive behavioural weight distorts the ordering of the most sensitive borderline cases. The operating value lies in the centre of the zero-class-change plateau , confirming that it is not an isolated working point and that the qualitative classification is robust to moderate re-specification of .
5.3. Calibration Robustness Check
As noted in Section 3.4, the min–max linear calibration of Equation (17) is the minimum-information map from the regret-adjusted preference to the unit interval. To verify that the three-way classification does not depend on this specific choice, two alternative calibrations are examined.
Logistic calibration. A logistic map , where is the sample mean of and is fitted so that the upper and lower quartiles of match those of the linear calibration, yields the same three-way classification for all twelve communities under the baseline parameter setting (, , ). The Spearman rank correlation between the logistic and linear calibrations is .
Isotonic regression calibration. An isotonic regression model is trained on a leave-one-out split, where the PLTS score of each community is regressed against a binary label (POS versus non-POS) derived from the linear baseline. The resulting monotone step function preserves the partition for all twelve communities.
Both alternative calibrations preserve the POS/BND/NEG partition and the within-class ranking identically, confirming that the classification is governed by the separation structure of the values rather than by the functional form of the calibration map. This robustness is expected when the classification gaps of Theorem 2 are large relative to the inter-calibration discrepancy, as is the case here: the smallest gap among the non-borderline communities is , whereas the maximum pointwise difference between the linear and logistic calibrations is .
5.4. Sensitivity to Loss Function Parameters
The loss function governs the thresholds . Holding the four corner losses fixed and varying the deferral costs jointly over , the closed-form formulas in Equation (12) give and , both varying monotonically in b while remaining well separated () within the tested range. The resulting threshold pairs and class composition are reported in Table 9.
Table 9.
Sensitivity of the three-way classification to the joint deferral cost , with fixed.
The class composition shifts smoothly as the deferral cost increases, with fewer alternatives deferred and more moved to POS or NEG, but the within-class ranking is preserved across all five loss-function configurations (Spearman throughout, because the underlying values do not change), confirming the robustness predicted by Theorem 2. Figure 6 maps the resulting POS/BND/NEG composition over the joint variation of and , showing a monotone migration of alternatives out of the boundary region as the deferral cost rises, with no abrupt reclassification anywhere in the tested grid; the central cell, marked in the figure, corresponds to the baseline setting .
Figure 6.
Heatmap of the three-way classification under joint variation of the boundary region loss parameters and .
5.5. Sensitivity to the Consensus Threshold
The consensus threshold is varied over , and the number of feedback iterations required together with the impact on the final ranking is recorded. Higher thresholds yield slightly different aggregated PLTS matrices and may shift borderline classifications; the results are reported in Table 10.
Table 10.
Sensitivity of consensus iteration count and ranking to the consensus threshold .
The recommended threshold sits at the favourable trade-off between iteration cost and agreement quality. Demanding 0.95 nearly sextuples the iteration count without materially improving the classification. Figure 7 plots this trade-off explicitly: the iteration count climbs gently up to and then rises steeply, while the final ranking stays almost unchanged, identifying as the knee of the curve.
Figure 7.
Trade-off curve between consensus threshold quality and feedback iteration count for the four tested values of .
5.6. Comparative Analysis with Classical MCDM Methods
PLR-3WBC is compared against four widely used MCDM methods: classical TOPSIS [1], VIKOR [7], PROMETHEE-II [8], and BWM-TOPSIS [1,24]. For comparability, all four classical methods are applied to the de-fuzzified score matrix obtained by Equation (3) from the converged PLR-3WBC group PLTS matrix, using the BBWM weights . The resulting ranks of the twelve communities are reported in Table 11.
Table 11.
Ranks of the twelve communities produced by PLR-3WBC and four classical MCDM methods.
The Spearman rank correlation between PLR-3WBC and the four classical methods exceeds in all cases, indicating that PLR-3WBC preserves the broad ordering established by classical methods while delivering three additional pieces of information that the classical methods do not provide: (i) an actionable three-way classification rather than a single-valued ranking, (ii) credal interval estimates of criterion weights from the BBWM, and (iii) a behavioural reordering of borderline non-dominated alternatives via the regret–rejoice adjustment. Figure 8 visualises the cross-method rank comparison; the five rank profiles track one another closely, and the visible discrepancies are confined to the borderline communities , and , where small differences in raw scores propagate to within-POS or within-BND rank swaps under different aggregation rules and where the regret–rejoice adjustment of PLR-3WBC becomes salient.
Figure 8.
Cross-method comparison of community rankings across PLR-3WBC, TOPSIS, VIKOR, PROMETHEE-II, and BWM-TOPSIS.
A methodological caveat is warranted regarding the scope of this comparison. All four classical methods are applied to the de-fuzzified expected-score matrix using the same BBWM weights, which reduces the PLTS evaluations to point values and thereby removes the distributional information that is the principal advantage of the probabilistic linguistic representation. The high Spearman correlations () are therefore expected: once the input is collapsed to a common deterministic matrix, the ranking differences among aggregation rules are necessarily small. This design choice was made to isolate the effect of the regret–rejoice adjustment and three-way classification on the final output, holding the input information constant.
A fully independent comparison would require each classical method to operate on its native input type, for example, fuzzy TOPSIS on fuzzy decision matrices, or interval VIKOR on interval scores. Such a comparison would conflate two sources of variation (information representation and aggregation rule) and obscure the contribution of the individual PLR-3WBC components. The ablation study of Table 12 addresses this concern from a different angle, by progressively enabling the regret, three-way, and consensus modules while holding the PLTS input fixed, and demonstrates that each module contributes a distinct and identifiable improvement. Accordingly, the high cross-method correlations should be read as confirming ordinal compatibility (the proposed framework does not produce anomalous rankings relative to established methods) rather than as evidence that PLR-3WBC offers no advantage. The added value of PLR-3WBC lies not in a different ranking order but in three outputs that the classical methods structurally cannot deliver: an actionable three-way partition, credal weight intervals quantifying expert agreement, and a behavioural reordering of borderline non-dominated alternatives reflecting the loss-averse psychology of the decision panel.
Table 12.
Ablation study: marginal contribution of each PLR-3WBC component.
5.7. Ablation Study
To isolate the marginal contribution of each PLR-3WBC component, five configurations are evaluated: (i) PLTS + BBWM only (no regret, no 3WD, no consensus iteration), (ii) PLTS + BBWM + regret only, (iii) PLTS + BBWM + 3WD only, (iv) PLTS + BBWM + regret + 3WD without consensus iteration, and (v) full PLR-3WBC. The metrics reported are the Spearman relative to the full configuration, the number of class changes, and the final consensus level.
The ablation confirms that consensus reaching is the single component that most increases the trustworthiness of the classification (raising from 0.794 to 0.871), while the regret adjustment and three-way decision module each contribute to the borderline reordering and to the actionable structure of the output. Figure 9 renders this ablation as a waterfall, making explicit that the consensus-reaching stage supplies the largest single increment to the trustworthiness of the output, whereas the regret and three-way modules chiefly reshape the borderline ordering and convert the continuous ranking into an actionable partition. The five ablation configurations in Table 12 correspond directly to several of the methodological combinations listed in Table 1: configuration (iii), PLTS + BBWM + 3WD, mirrors the methodological scope of Zhu et al. [18] (PLTS + regret + 3WD) with Bayesian weighting substituted for entropy-based weighting; and configuration (iv), PLTS + BBWM + regret + 3WD, mirrors Zhu et al. [20] minus the consensus module. The ablation therefore serves as an internal comparison against the methodological scopes of these predecessor studies: the incremental gain from adding consensus (configuration (iv) to (v)) is a rise in CL from 0.794 to 0.871 with no class change relative to the full framework, demonstrating that the consensus stage (absent from [18] and present but without a formal convergence guarantee in [20]) adds quantifiable trustworthiness to the group output. The incremental gain from adding regret (configuration (iii) to (iv)) is a correction of one borderline class change, confirming that the behavioural adjustment contributes precision at the classification boundary rather than a wholesale reordering.
Figure 9.
Ablation waterfall visualisation showing the marginal contribution of each PLR-3WBC component to the final classification and consensus level.
To address the concern that the ablation study should be compared more explicitly against the methodological combinations in Table 1, Table 13 restates the correspondence and reports, for each ablation configuration, the Spearman rank correlation with the full PLR-3WBC ranking and the number of three-way class changes.
Table 13.
Mapping of ablation configurations to representative literature combinations and their incremental effects.
The table clarifies that each ablation configuration approximates the methodological scope of a specific predecessor study listed in Table 1, and that the incremental additions (regret theory correcting one borderline class change from configuration (iii) to (iv), and consensus raising from 0.794 to 0.871 from configuration (iv) to (v)) are each individually identifiable. A fully independent re-implementation of each predecessor method with its original weighting scheme (e.g., entropy weights for [18], equal weights for [20]) would introduce additional confounds from the weighting difference; the ablation design therefore isolates the effect of each module by holding the PLTS input and BBWM weights constant, which we consider a more controlled comparison than a multi-source replication.
6. Discussion
The PLR-3WBC framework integrates five strands of MCDM methodology that have evolved separately. Probabilistic linguistic term sets [5,10,11,13,14] provide a richer representation of expert hesitation than ordinary fuzzy or hesitant fuzzy sets, by carrying probability mass on each linguistic term. The Bayesian best–worst method [23,24,25,26,27,28] provides credal weight intervals that quantify uncertainty arising from disagreement among experts, an output that single-point weighting methods cannot deliver. Regret theory [6,19,21,22,41] captures the asymmetric psychology of decision makers who are sensitive to outcomes that fall short of foregone alternatives. Three-way decisions [9,15,16,17] deliver an actionable partition into accept, reject, and defer regions that better matches the discrete decision points of programme management than a continuous ranking. Consensus reaching [29,30,31,32,44] provides an iterative mechanism for resolving group disagreement.
Each component also carries a characteristic risk that the integrated pipeline is designed to manage, and weighing the advantage of each module against its risk clarifies why the five are combined rather than used in isolation. The expressive freedom of PLTS, while it captures hesitation faithfully, can magnify the elicitation burden on experts and amplify inter-expert divergence. Credal weighting via BBWM exposes the depth of agreement on each criterion, but its output depends on the quality of the elicited best–worst comparisons and on the convergence of the MCMC sampler. The regret–rejoice adjustment injects realistic loss-averse behaviour, yet it introduces the behavioural parameters and , whose mis-specification could distort the ordering of borderline alternatives. The three-way thresholds yield risk-aware actions, but inherit the subjectivity of the elicited loss function. The consensus process reconciles divergent opinions, but presumes that experts are willing to revise their judgements. The architecture of Figure 1 contains these risks jointly: the consensus stage curbs PLTS divergence before aggregation, the credible intervals flag weak weight identification, and the sensitivity analyses of Section 5 bound the influence of the behavioural and loss parameters. The strengths of the five paradigms are thereby combined while their individual failure modes are held in check.
Three findings of the case study are managerially salient. The classification places six of the twelve communities in the immediate-retrofit category (POS region), three in the deferred-investigation category (BND region), and three in the not-prioritised category (NEG region), creating a budget-aligned action plan that is more useful than a twelve-level ranking for programme administrators with discrete budget tranches. The credal weight intervals flag five mid-ranked criteria (C1.4, C2.1, C5.2, C4.3 and C2.3) whose 95 percent intervals mutually overlap, indicating that the precise weight ordering within this band carries less information than the overall weight magnitude, a finding that single-point weighting methods cannot surface. The regret adjustment refines the within-class ordering between adjacent communities and in the POS region: although is marginally superior to on three criteria in the raw score , is more uniformly close to the criterion-wise maxima across all eighteen criteria and thus accumulates less regret, securing its higher within-class rank under the regret-adjusted preference .
Several limitations should be acknowledged. The most immediate concerns the calibration in Equation (17) from preference to conditional probability , which is the minimum-information linear map; richer calibrations, including isotonic regression against a held-out training set or domain-expert anchoring, would refine the classification but require additional data not always available in a one-shot evaluation. A related limitation is that the consensus reaching process assumes that experts are willing to revise their evaluations toward the group median; an unwilling expert can be modelled by reducing the feedback intensity for that expert, but the convergence guarantee of Theorem 3 would then need to be modified to a weighted contraction, which is left to future work. In addition, the case study comprises twelve communities and four experts; larger-scale group decision making with hundreds of experts requires the introduction of clustering schemes and is the subject of an active research literature [30,32,40,44,45,48]. Moreover, the indicator system, while grounded in WHO [34] and Buffel and Phillipson [35], is context-specific and may require adaptation in other cultural or geographic settings. A further limitation is that the comparative analysis against classical MCDM methods is conducted on a common de-fuzzified score matrix, which reduces the PLTS evaluations to point values and therefore cannot demonstrate the full informational advantage of the probabilistic linguistic representation; the high Spearman correlations confirm ordinal compatibility but should not be over-interpreted as evidence of superiority. Finally, the entire validation rests on a single case study with twelve communities and four experts; additional case studies in different geographic and cultural contexts, with larger alternative sets and expert panels, are needed before the framework can be recommended for routine deployment.
7. Conclusions and Future Work
PLR-3WBC, a probabilistic linguistic regret-driven three-way Bayesian consensus framework for multi-criteria group decision making, integrates probabilistic linguistic term sets, the Bayesian best–worst method, regret–rejoice value adjustment, three-way decision classification, and a feedback-driven consensus reaching process within a single mathematical pipeline. Three theoretical results underpin the framework: order preservation of the regret-adjusted PLTS score under admissible monotone transformations (Theorem 1); finite-time convergence of the consensus reaching process with an explicit iteration bound derived from the disagreement concentration ratio (Theorem 3); and classification stability under bounded perturbations of the decision matrix (Theorem 2). Two propositions characterise computational complexity (Proposition 1) and Pareto optimality of the acceptance region (Proposition 2).
A case study on the age-friendliness evaluation of twelve aging urban residential communities, with an indicator system of five dimensions and eighteen criteria assessed by four expert decision makers, provides preliminary evidence that PLR-3WBC can produce an actionable three-way classification and recover a transparent group consensus from initial disagreement. Rankings are broadly consistent with classical TOPSIS, VIKOR, PROMETHEE-II, and BWM-TOPSIS (Spearman rank correlation exceeding 0.97 when all methods are applied to the same de-fuzzified score matrix), confirming that the integrated pipeline preserves the ordinal reliability of each established constituent method. The practical value of PLR-3WBC lies not in producing a different ranking order but in the three complementary outputs that emerge from the methodological integration: an actionable three-way partition aligned with discrete budget tranches, credal weight intervals revealing the depth of expert agreement, and a behavioural reordering of borderline non-dominated alternatives that captures the loss-averse psychology of the decision panel. Sensitivity analyses with respect to the regret aversion coefficient , the blending parameter , the loss function parameters, the calibration choice, and the consensus threshold indicate that the qualitative classification is stable across a wide parameter envelope within this single case study; however, generalisation to other cities, larger alternative sets, and larger expert panels requires further empirical testing before the stability claims can be considered established.
Given the present depth of validation, namely a single-city case with twelve communities and four experts, the framework is best positioned as a transferable decision protocol rather than a source of universally calibrated weights or thresholds. Its applicable scope covers group MCDM problems that share four features with the age-friendliness setting: qualitative evaluations carrying probabilistic hesitation, heterogeneous and behaviourally distorted expert judgement, genuine group disagreement requiring reconciliation, and a need for discrete, risk-aware actions rather than a single continuous ranking. Within this scope, adjacent applications include supplier selection under linguistic risk assessment [10,39], hospital and elderly care facility siting under multi-stakeholder consultation, and renewable energy site evaluation; the five-stage pipeline and the theoretical guarantees carry over directly, while the absolute parameter values must be re-elicited for each new context.
The translational value of the framework rests on the actionability of its output, and realising that value requires coordinated adjustment by several actors. For municipal housing and urban renewal authorities, the three-way partition maps directly onto budget tranches: the POS region defines the immediate retrofit list, the BND region defines a monitoring-and-survey backlog rather than a rejection list, and the NEG region defines the deferred set, allowing finance departments to align disbursement with the within-class ranking. Civil affairs and health bureaus responsible for elderly care are guided by the credal weights, which here elevate emergency response coverage, accessibility, and proximity to medical facilities above digital service criteria, indicating where service investment yields the largest marginal age-friendliness gain. Community-level management committees become the data and feedback layer, since the consensus-reaching process depends on their willingness to record evidence and revise assessments. For the framework to be adopted in practice, these bodies would need to standardise the indicator checklist, institutionalise the periodic re-elicitation of expert evaluations, and treat the boundary region as a trigger for targeted investigation rather than as an administrative dead end.
Future research may pursue five directions. One direction is extension to large-scale group decision making with hundreds of experts via clustering and trust network mechanisms [30,32,40,44,45,48]. Another direction is replacement of the linear calibration in Equation (17) by an isotonic or Bayesian-anchored calibration when historical data on retrofit outcomes are available. A further direction is integration of dynamic preference revision over time, in which experts update their evaluations as new evidence emerges. An additional priority is multi-case empirical validation across geographically and culturally diverse settings, with larger alternative sets and expert panels, to establish the external validity of the framework beyond the single provincial-capital case reported here. A final direction is derivation of a closed-form contraction rate for the consensus process under various PLTS distance choices.
Author Contributions
Conceptualization, Z.Z., C.Y. and K.T.; methodology, Z.Z. and K.T.; software, Z.Z. and C.C.; validation, Z.Z., C.Y., C.C., F.Z. and K.T.; formal analysis, Z.Z. and C.C.; investigation, Z.Z., C.Y., C.C., F.Z. and K.T.; resources, C.Y., C.C., F.Z. and K.T.; data curation, Z.Z. and F.Z.; writing—original draft, Z.Z., C.C. and F.Z.; writing—review and editing, C.Y. and K.T.; visualization, Z.Z. and F.Z.; supervision, C.Y. and K.T.; project administration, C.Y. and Z.Z.; funding acquisition, C.Y. All authors have read and agreed to the published version of the manuscript.
Funding
This research was supported by the National Natural Science Foundation of China (NSFC) under Grant Nos. 51378427 and 51678504.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Hwang, C.L.; Yoon, K. Multiple Attribute Decision Making: Methods and Applications, A State-of-the-Art Survey; Lecture Notes in Economics and Mathematical Systems; Springer: Berlin/Heidelberg, Germany, 1981; Volume 186. [Google Scholar] [CrossRef] [Scilit]
- Greco, S.; Ehrgott, M.; Figueira, J.R. (Eds.) Multiple Criteria Decision Analysis: State of the Art Surveys; International Series in Operations Research & Management Science; Springer: New York, NY, USA, 2016. [Google Scholar] [CrossRef] [Scilit]
- Fan, Y.; Zhang, H.; Ding, Y.; Wu, Y.; Guo, J.; Li, F. Multi-objective optimization based on reduced order model for VHTR design. Ann. Nucl. Energy 2026, 237, 112460. [Google Scholar] [CrossRef] [Scilit]
- Rodriguez, R.M.; Martinez, L.; Herrera, F. Hesitant fuzzy linguistic term sets for decision making. IEEE Trans. Fuzzy Syst. 2012, 20, 109–119. [Google Scholar] [CrossRef] [Scilit]
- Pang, Q.; Wang, H.; Xu, Z. Probabilistic linguistic term sets in multi-attribute group decision making. Inf. Sci. 2016, 369, 128–143. [Google Scholar] [CrossRef] [Scilit]
- Bell, D.E. Regret in decision making under uncertainty. Oper. Res. 1982, 30, 961–981. [Google Scholar] [CrossRef] [Scilit]
- Opricovic, S.; Tzeng, G.H. Compromise solution by MCDM methods: A comparative analysis of VIKOR and TOPSIS. Eur. J. Oper. Res. 2004, 156, 445–455. [Google Scholar] [CrossRef] [Scilit]
- Brans, J.P.; Vincke, P. A preference ranking organisation method: The PROMETHEE method for multiple criteria decision-making. Manag. Sci. 1985, 31, 647–656. [Google Scholar] [CrossRef] [Scilit]
- Yao, Y. Three-way decisions with probabilistic rough sets. Inf. Sci. 2010, 180, 341–353. [Google Scholar] [CrossRef] [Scilit]
- Kang, W.; Liang, X.; Peng, Y. Probabilistic linguistic multiple attribute group decision-making based on a Choquet operator and its application in supplier selection. Mathematics 2025, 13, 740. [Google Scholar] [CrossRef] [Scilit]
- Lin, M.; Chen, Z.; Xu, Z.; Gou, X.; Herrera, F. Score function based on concentration degree for probabilistic linguistic term sets: An application to TOPSIS and VIKOR. Inf. Sci. 2021, 551, 270–290. [Google Scholar] [CrossRef] [Scilit]
- Liao, H.; Mi, X.; Xu, Z. A survey of decision-making methods with probabilistic linguistic information: Bibliometrics, preliminaries, methodologies, applications and future directions. Fuzzy Optim. Decis. Mak. 2020, 19, 81–134. [Google Scholar] [CrossRef] [Scilit]
- Wu, W. Probabilistic uncertain linguistic VIKOR method for teaching reform plan evaluation for the core course “Big Data Technology and Applications” in the digital economy major. Mathematics 2024, 12, 3710. [Google Scholar] [CrossRef] [Scilit]
- Niu, X.; Song, Y.; Xu, Z. Large-scale group decision-making method with public participation and its application in urban management. Mathematics 2024, 12, 2528. [Google Scholar] [CrossRef] [Scilit]
- Liang, D.; Xu, Z.; Liu, D.; Wu, Y. Method for three-way decisions using ideal TOPSIS solutions at Pythagorean fuzzy information. Inf. Sci. 2018, 435, 282–295. [Google Scholar] [CrossRef] [Scilit]
- Qin, H.; Han, Y.; Ma, X. Multi-attribute three-way decision approach based on ideal solutions under interval-valued fuzzy soft environment. Symmetry 2024, 16, 1327. [Google Scholar] [CrossRef] [Scilit]
- Luo, J.; Zhang, W.; Su, J.; Chen, J. Decision-theoretic rough sets for three-way decision-making in dilemma reasoning and conflict resolution. Mathematics 2025, 13, 2111. [Google Scholar] [CrossRef] [Scilit]
- Zhu, J.; Ma, X.; Martínez, L.; Zhan, J. A probabilistic linguistic three-way decision method with regret theory via fuzzy c-means clustering algorithm. IEEE Trans. Fuzzy Syst. 2023, 31, 2821–2835. [Google Scholar] [CrossRef] [Scilit]
- Wang, P.; Chen, J. A multi-attribute group decision making method based on novel distance measures and regret theory under probabilistic dual hesitant fuzzy sets. J. Intell. Fuzzy Syst. 2024, 46, 659–675. [Google Scholar] [CrossRef] [Scilit]
- Zhu, J.; Ma, X.; Kou, G.; Herrera-Viedma, E.; Zhan, J. A three-way consensus model with regret theory under the framework of probabilistic linguistic term sets. Inf. Fusion 2023, 95, 250–274. [Google Scholar] [CrossRef] [Scilit]
- Yan, S.; Bian, R.; Xu, Y.; Ji, F.; Cabrerizo, F.J. Quantum group decision making based on regret theory and grey relational degree considering psychological preference. Comput. Ind. Eng. 2026, 214, 111912. [Google Scholar] [CrossRef] [Scilit]
- Ruan, C.; Gong, S.; Chen, X. Multi-criteria group decision-making with extended ELECTRE III method and regret theory based on probabilistic interval-valued intuitionistic hesitant fuzzy information. Complex Intell. Syst. 2025, 11, 92. [Google Scholar] [CrossRef] [Scilit]
- Mohammadi, M.; Rezaei, J. Bayesian best–worst method: A probabilistic group decision making model. Omega 2020, 96, 102075. [Google Scholar] [CrossRef] [Scilit]
- Rezaei, J. best–worst multi-criteria decision-making method. Omega 2015, 53, 49–57. [Google Scholar] [CrossRef] [Scilit]
- Rezaei, J. best–worst multi-criteria decision-making method: Some properties and a linear model. Omega 2016, 64, 126–130. [Google Scholar] [CrossRef] [Scilit]
- Arı, E.; Dursun, M. An integrated Bayesian best–worst method and consensus-based intuitionistic fuzzy evaluation based on distance from average solution approach for evaluating alternative aircraft models from a sustainability perspective. Symmetry 2024, 16, 1086. [Google Scholar] [CrossRef] [Scilit]
- Haseli, G.; Sheikh, R.; Jafarzadeh Ghoushchi, S.; Hajiaghaei-Keshteli, M.; Moslem, S.; Deveci, M.; Kadry, S. An extension of the best–worst method based on the spherical fuzzy sets for multi-criteria decision-making. Granul. Comput. 2024, 9, 40. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Peykani, P.; Emrouznejad, A.; Nouri, M. best–worst multi-criteria decision-making method: A review of the literature. Socio-Econ. Plan. Sci. 2026, 104, 102345. [Google Scholar] [CrossRef] [Scilit]
- Dong, Y.; Zhang, H.; Herrera-Viedma, E. Integrating experts’ weights generated dynamically into the consensus reaching process and its applications in managing non-cooperative behaviors. Decis. Support Syst. 2016, 84, 1–15. [Google Scholar] [CrossRef] [Scilit]
- Sun, X.; Zhu, J.; Wang, J.; Pérez-Gálvez, I.J.; Cabrerizo, F.J. Consensus-reaching process in multi-stage large-scale group decision-making based on social network analysis: Exploring the implication of herding behavior. Inf. Fusion 2024, 104, 102184. [Google Scholar] [CrossRef] [Scilit]
- Li, S.; Rodríguez, R.M.; Wei, C.; Shu, T. A consensus model for large scale group decision making with hesitant fuzzy linguistic information and hierarchical feedback mechanism. Comput. Ind. Eng. 2022, 173, 108669. [Google Scholar] [CrossRef] [Scilit]
- Zhou, Y.J.; Zhou, M.; Liu, X.B.; Cheng, B.Y.; Herrera-Viedma, E. Consensus reaching mechanism with parallel dynamic feedback strategy for large-scale group decision making under social network analysis. Comput. Ind. Eng. 2022, 174, 108818. [Google Scholar] [CrossRef] [Scilit]
- Ding, W.; Li, X.; Shen, X. Three-way group decisions using evidence theory under hesitant fuzzy linguistic environment. Sci. Rep. 2023, 13, 22766. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- World Health Organization. Global Age-Friendly Cities: A Guide; World Health Organization: Geneva, Switzerland, 2007; Available online: https://www.who.int/publications/i/item/9789241547307 (accessed on 21 June 2026).
- Buffel, T.; Phillipson, C. A manifesto for the age-friendly movement: Developing a new urban agenda. J. Aging Soc. Policy 2018, 30, 173–192. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dong, W.; Sun, S.; Fu, Y. Assessing urban community parks from an age-friendly perspective: A multi-criteria decision-making approach. Front. Public Health 2025, 13, 1663359. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Raut, S.B. Applying multi-criteria decision analysis to prioritise age-friendly criteria for policy implications. Int. J. Urban Sustain. Dev. 2023, 15, 250–266. [Google Scholar] [CrossRef] [Scilit]
- Webster, S.; Robertson, M.; Keresztes, C.; Puxty, J. Assessing age-friendly community initiatives: Developing a novel survey tool for assessment and evaluation. Gerontologist 2024, 64, gnae146. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, F.; Li, M.; Ye, Z.; Niu, Y. A multi-stage group decision making approach for sustainable supplier selection based on probabilistic linguistic time-ordered incentive operator. PLoS ONE 2023, 18, e0293019. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, Z.; Liu, H.; Ma, R. A probabilistic linguistic large-group emergency decision-making method based on the Louvain algorithm and group pressure model. Mathematics 2025, 13, 670. [Google Scholar] [CrossRef] [Scilit]
- Kahneman, D.; Tversky, A. Prospect theory: An analysis of decision under risk. Econometrica 1979, 47, 263–291. [Google Scholar] [CrossRef] [Scilit]
- Lei, W.; Ma, W.; Li, X.; Sun, B. Three-way group decision based on regret theory under dual hesitant fuzzy environment: An application in water supply alternatives selection. Expert Syst. Appl. 2024, 237, 121249. [Google Scholar] [CrossRef] [Scilit]
- Peng, X.; Yang, Y. Algorithms for interval-valued fuzzy soft sets in stochastic multi-criteria decision making based on regret theory and prospect theory with combined weight. Appl. Soft Comput. 2017, 54, 415–430. [Google Scholar] [CrossRef] [Scilit]
- Liang, W.; Labella, Á.; Meng, M.J.; Wang, Y.M.; Rodríguez, R.M. Hesitant fuzzy consensus reaching process for large-scale group decision-making methods. Mathematics 2025, 13, 1182. [Google Scholar] [CrossRef] [Scilit]
- Xu, Z.; Xu, H.; Li, P.; Wei, C. Social network group decision-making method based on stochastic multi-criteria acceptability analysis for probabilistic linguistic term sets. Inf. Sci. 2024, 681, 121269. [Google Scholar] [CrossRef] [Scilit]
- Aires, R.F.F.; Ferreira, L. A new multi-criteria approach for sustainable material selection problem. Sustainability 2022, 14, 11191. [Google Scholar] [CrossRef] [Scilit]
- de Oliveira, M.S.; Trojan, F.; Steffen, V. Model for definition of multi-criteria compensation by the ICCI (inter-criteria compensation index) in the ranking of electric vehicles. Energies 2025, 18, 5553. [Google Scholar] [CrossRef] [Scilit]
- Zhang, H.; Xiao, J.; Palomares, I.; Liang, H.; Dong, Y. Linguistic distribution-based optimization approach for large-scale GDM with comparative linguistic information: An application on the selection of wastewater disinfection technology. IEEE Trans. Fuzzy Syst. 2020, 28, 376–389. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.








