Score Cloud Analysis for Rule-Aware Ranking Robustness Under Discrete Judgment Uncertainty
Abstract
1. Introduction
2. Materials and Methods
2.1. Structural Representation of a Rule-Defined Ranking
2.2. Wushu Taolu as a Synthetic Case Study
2.3. Deterministic Score-Cloud Diagnostics

| Label | Synthetic Rule |
|---|---|
| Red | At least one of the following holds: ; ; A/C single-call tie/rank activity is present; or the official score margin is zero. |
| Amber | The pair is not red, and at least one of the following holds: ; ; or deterministic screening envelope overlap/tie activity is present. |
| Green | Neither the red nor the amber condition holds under the declared synthetic screening convention. |
2.4. Single-Call Rank Sensitivity
2.5. Group B Sensitivity and Discarded Mark Re-Entry
2.6. Scenario-Conditional Rank Clouds
2.7. Choosing a Perturbation Model
2.8. Diagnostic Quantities
3. Results: Synthetic Stress Test
3.1. Design
3.2. Baseline Synthetic Case
3.3. Rank Cloud and Score Cloud Behavior
3.4. Sensitivity to Perturbation Scale and Dependence
3.5. Transferability Beyond the Case Study
4. Discussion
4.1. What Is New and What Is Not
4.2. Synthetic Demonstration Versus Empirical Validation
4.3. Bayesian Interpretation
4.4. Robustness Is Not Correctness
4.5. Applicability Beyond Wushu Taolu
5. Limitations
6. Conclusions
Supplementary Materials
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Goldstein, H.; Spiegelhalter, D.J. League Tables and Their Limitations: Statistical Issues in Comparisons of Institutional Performance. J. R. Stat. Soc. Ser. A Stat. Soc. 1996, 159, 385–443. [Google Scholar] [CrossRef] [Scilit]
- Lockwood, J.R.; Louis, T.A.; McCaffrey, D.F. Uncertainty in Rank Estimation: Implications for Value-Added Modeling Accountability Systems. J. Educ. Behav. Stat. 2002, 27, 255–270. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Saisana, M.; Saltelli, A.; Tarantola, S. Uncertainty and Sensitivity Analysis Techniques as Tools for the Quality Assessment of Composite Indicators. J. R. Stat. Soc. Ser. A Stat. Soc. 2005, 168, 307–323. [Google Scholar] [CrossRef] [Scilit]
- Saltelli, A.; Annoni, P.; Azzini, I.; Campolongo, F.; Ratto, M.; Tarantola, S. Variance Based Sensitivity Analysis of Model Output. Design and Estimator for the Total Sensitivity Index. Comput. Phys. Commun. 2010, 181, 259–270. [Google Scholar] [CrossRef] [Scilit]
- Borgonovo, E. A New Uncertainty Importance Measure. Reliab. Eng. Syst. Saf. 2007, 92, 771–784. [Google Scholar] [CrossRef] [Scilit]
- JCGM 100:2008(E); Evaluation of Measurement Data: Guide to the Expression of Uncertainty in Measurement. BIPM: Sèvres, France, 2008. [CrossRef] [Scilit]
- JCGM 101:2008; Evaluation of Measurement Data: Supplement 1 to the Guide to the Expression of Uncertainty in Measurement: Propagation of Distributions Using a Monte Carlo Method. BIPM: Sèvres, France, 2008. [CrossRef] [Scilit]
- Cohen, J. A Coefficient of Agreement for Nominal Scales. Educ. Psychol. Meas. 1960, 20, 37–46. [Google Scholar] [CrossRef] [Scilit]
- Fleiss, J.L. Measuring Nominal Scale Agreement among Many Raters. Psychol. Bull. 1971, 76, 378–382. [Google Scholar] [CrossRef] [Scilit]
- Shrout, P.E.; Fleiss, J.L. Intraclass Correlations: Uses in Assessing Rater Reliability. Psychol. Bull. 1979, 86, 420–428. [Google Scholar] [CrossRef] [PubMed]
- Koo, T.K.; Li, M.Y. A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research. J. Chiropr. Med. 2016, 15, 155–163. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Brennan, R.L. Generalizability Theory; Springer: New York, NY, USA, 2001. [Google Scholar]
- Zitzewitz, E. Nationalism in Winter Sports Judging and Its Lessons for Organizational Decision Making. J. Econ. Manag. Strategy 2006, 15, 67–99. [Google Scholar] [CrossRef] [Scilit]
- Heiniger, S.; Mercier, H. Judging the Judges: Evaluating the Accuracy and National Bias of International Gymnastics Judges. J. Quant. Anal. Sport. 2021, 17, 289–305. [Google Scholar] [CrossRef] [Scilit]
- Tukey, J.W. A Survey of Sampling from Contaminated Distributions. In Contributions to Probability and Statistics; Olkin, I., Ghurye, S.G., Hoeffding, W., Madow, W.G., Mann, H.B., Eds.; Stanford University Press: Stanford, CA, USA, 1960; pp. 448–485. [Google Scholar]
- Huber, P.J.; Ronchetti, E.M. Robust Statistics, 2nd ed.; Wiley: Hoboken, NJ, USA, 2009. [Google Scholar] [CrossRef] [Scilit]
- International Wushu Federation. The International Wushu Federation (IWUF) Released the Latest Wushu Competition Rules and Judging Methods, Which Will Be Officially Implemented in 2025. 2024. Available online: https://www.iwuf.org/en/news/gjwl/2024/0914/8955.html (accessed on 30 May 2026).
- International Wushu Federation. Wushu Taolu Competition Rules and Judging Methods (2024); International Wushu Federation: Lausanne, Switzerland, 2024; Available online: https://www.iwuf.org/xhimg/soft/240912/WUSHU-TAOLU-COMPETITION-RULES-AND-JUDGING-METHODS-2024.pdf (accessed on 30 May 2026).
- Pearl, J. Causal Inference in Statistics: An Overview. Stat. Surv. 2009, 3, 96–146. [Google Scholar] [CrossRef] [Scilit]
- Bareinboim, E.; Correa, J.D.; Ibeling, D.; Icard, T. On Pearl’s Hierarchy and the Foundations of Causal Inference. In Probabilistic and Causal Inference: The Works of Judea Pearl; Dechter, R., Geffner, H., Halpern, J.Y., Eds.; ACM Books: New York, NY, USA, 2022; pp. 507–556. [Google Scholar] [CrossRef] [Scilit]
- Morris, T.P.; White, I.R.; Crowther, M.J. Using Simulation Studies to Evaluate Statistical Methods. Stat. Med. 2019, 38, 2074–2102. [Google Scholar] [CrossRef] [Scilit] [PubMed]





| Related Approach | Usual Target | What Is Retained Here | Added Reporting Object |
|---|---|---|---|
| Uncertainty propagation | Distribution of a measured or model-derived quantity | Monte Carlo propagation through the complete rule-defined map | Rank frequencies after official tie-break recomputation |
| Finite-difference sensitivity | Response to a specified input perturbation | Exact recomputation under thresholds, caps, trimming, truncation, and tie-break rules | Directional score exposures and single-call rank sensitivity |
| Rank uncertainty | Uncertainty in an ordering or league position | Pairwise inversion and rank-persistence frequencies | Rule-aware Rank Cloud with official tie-breaks |
| Rater reliability | Agreement, severity, or variance components | Recognition that recorded marks are rater-mediated | No reliability, severity, or variance component is estimated without repeated empirical ratings |
| Robust statistics | Resistance to outliers or contamination | Trimmed and winsorized summaries motivate Overall Performance (Group B) sensitivity | Re-entry of discarded marks into tie-breaks can be diagnosed |
| Supported by the Method | Not Supported Without Empirical Calibration |
|---|---|
| Sensitivity of the official scoring and ranking rule under declared perturbations. | Probability that a judge, athlete, panel, or federation made an error. |
| Deterministic exposure and scenario-conditional Rank Cloud frequencies. | True underlying athletic ranking or true performance order. |
| Comparison of alternative perturbation scenarios in a synthetic proof-of-concept. | Prevalence of fragility in real Wushu competitions. |
| Identification of pairs whose ranking relation is active under the specified stress test. | Operational policy thresholds or official review triggers. |
| Model Choice | Use When | Required Information | Main Risk |
|---|---|---|---|
| Independent calls | No evidence of shared ambiguity or panel-level dependence is available. | Declared perturbation probabilities for each decision class. | May understate uncertainty if calls are correlated. |
| Intra-athlete random effect | Several marginal calls may share routine-level ambiguity. | Repeated review or elicited intraclass dependence. | Sensitivity depends on assumed dependence. |
| Panel-level effect | Rater-scale or panel interpretation may affect several marks. | Repeated panels, judge logs, or calibration exercises. | Confounding with athlete quality. |
| Bayesian predictive model | Expert beliefs or empirical data can be encoded as priors/posteriors. | Prior distributions or calibration data. | Results inherit prior and model assumptions. |
| Quantity | Definition/Status | Required Inputs | Interpretation | Not Interpreted as |
|---|---|---|---|---|
| ; proposed aggregate diagnostic | Marginal A/C calls and signed finite differences | Aggregate contested-point exposure; and are the directional quantities | Error probability | |
| Equation (23); reporting construct | , , , , , and | Ranking-active pairwise screening exposure | Calibrated risk | |
| FMR | Pairwise exposure divided by official margin; proposed diagnostic | Relevant pairwise exposure and ; regularization is used for synthetic screening and plotting | Exposure relative to the official margin | Probability |
| Single-call activity | , including tie-active equality; proposed diagnostic | Largest ranking-relevant marginal A/C call and the tie-aware ranking rule | Whether one A/C call can activate a tie or rank change | Proof of a wrong call |
| Score Cloud | Image of under admissible perturbation states; adapted uncertainty-set display | Scoring function, admissible states, and score inputs | Rule-aware score support under declared states | True-score interval |
| Rank Cloud | Distribution of under sampled perturbations; adapted rank uncertainty display | Ranking function, perturbation distribution, and tie-break variables | Scenario-conditional rank frequency; finite-Monte Carlo uncertainty is reported separately | True-rank probability |
| Equation (31); adapted Monte Carlo summary | Simulated ranks after recomputing | Scenario-conditional pairwise rank inversion frequency, reported with Equation (34) | Empirical judging-error rate |
| Quantity | Boundedness | Monotonicity/Sensitivity | Zero Margin Behavior | Model Dependence |
|---|---|---|---|---|
| Bounded by the finite sum of declared marginal A/C units and rule caps | Nondecreasing when same-direction marginal units are added | Not a ratio; defined independently of pairwise margin | No | |
| Bounded by the included A/C, B screening, and directional Head Judge components | Nondecreasing in every included exposure component | Defined at zero margin; its raw ratio to margin is not | No | |
| Raw FMR | Unbounded as | Increases with exposure and decreases with nonzero margin | Undefined when | No |
| Epsilon-screened FMR | Bounded above by maximum admissible exposure divided by the chosen floor | Used for declared synthetic screening and plotting at zero or near-zero margins | Uses as a screening/display convention; the raw ratio remains undefined at zero | No |
| Score Cloud | Bounded by the image of admissible perturbation states under score caps | Under fixed , nested admissible sets imply nested exact Score Clouds; non-nested changes imply no monotonicity | Defined at athlete level | Optional weights only |
| Rank Cloud | Frequencies sum to one over represented rank or rank/tie states | Changes with perturbation probabilities, dependence, and rule changes | May allocate mass to tie-break or residual-tie states when represented | Yes |
| Design Element | Baseline Value or Rule | Purpose |
|---|---|---|
| Competitors | synthetic competitors | Creates adjacent-pair relations across several official margins. |
| Monte Carlo replications | Gives maximum plug-in binomial Monte Carlo standard error (MCSE) near 0.0022. | |
| A/C marginal calls | Baseline outcome-reversal probability 0.28 for marginal calls; robust 3–0 calls have zero outcome-reversal probability | Controls one-vote decision perturbation activity. |
| Group B baseline model | Five two-decimal marks; centered Gaussian mark perturbations; mapping to the centesimal grid; clipping to the synthetic legal range ; sorting, trimming, , and recomputation | Preserves the rule-aware mark-level B state and tie-break input. |
| Deterministic B screening width | Pre-specified : 0.019 (low variation), 0.030 (moderate), and 0.055–0.070 (high); independent of the Monte Carlo Gaussian scale | Supplies a transparent deterministic screening width for TSE and Figure 2. |
| Head Judge perturbation | Signed total-score changes of where declared; baseline activation probability 0.12 | Tests directional Head Judge exposure in pairwise ranking relations. |
| Synthetic labels | Green/amber/red labels defined by declared FMR, single-call, envelope-contact, zero margin, and rank frequency rules | Provides descriptive stress test classes, not operational thresholds. |
| Scenario Element | Declared Values | Purpose |
|---|---|---|
| A/C scale multipliers | Low 0.35, baseline 1.00, high 1.40; scaled probabilities clipped to | Tests sensitivity to the declared call-perturbation scale. |
| Group B mark scales | Athlete-level standard deviations: 0.0076 (low), 0.025 (moderate), and approximately 0.048–0.060 (high); scenario multipliers 0.50, 1.00, and 1.50 | Creates low-, moderate-, and high-variation mark profiles. |
| Group B dependence | Within-athlete mark correlations 0.10 (low), 0.20 (moderate), and 0.65 (high) | Tests dependence among marks entering the trimmed component. |
| Head Judge scale | Activation probabilities 0.05 (low), 0.12 (baseline), and 0.18 (high) | Varies the occurrence of declared signed Head Judge states. |
| Intra-athlete random effect | Independent baseline; correlated scenario uses a Beta random effect with the scaled marginal-call mean and concentration at | Represents shared ambiguity among calls for the same athlete. |
| Retained-block comparison | Gaussian perturbation of the retained B component, clipped to ; the lower discarded mark is held at baseline | Provides a compact surrogate comparison, not the baseline model. |
| Pair | Margin | TSE | FMRtot,ε | Single-Call | Label | ||||
|---|---|---|---|---|---|---|---|---|---|
| A01–A02 | 0.003 | 1.300 | 0.300 | 0.116 | 0.000 | 1.716 | 572.00 | yes | red |
| A03–A04 | 0.170 | 0.000 | 0.000 | 0.038 | 0.000 | 0.038 | 0.22 | no | green |
| A06–A07 | 0.000 | 1.300 | 0.300 | 0.139 | 0.000 | 1.739 | 1739.00 | tie-break active | red |
| A10–A11 | 0.130 | 0.100 | 0.000 | 0.049 | 0.100 | 0.249 | 1.92 | no | amber |
| Pair | Score Inversion | Score Tie | Rank Inversion | MCSE | 95% Monte Carlo Interval |
|---|---|---|---|---|---|
| A01–A02 | 0.4837 | 0.0036 | 0.4853 | 0.0022 | [0.4809, 0.4897] |
| A03–A04 | 0.0000 | 0.0000 | 0.0000 | 0.0000 | |
| A06–A07 | 0.4877 | 0.0032 | 0.4890 | 0.0022 | [0.4846, 0.4934] |
| A10–A11 | 0.0377 | 0.0068 | 0.0428 | 0.0009 | [0.0411, 0.0446] |
| Scenario | Red | Amber | Green | Pairs Changing Class | Mean Rank Inversion Frequency | Maximum Rank Inversion Frequency |
|---|---|---|---|---|---|---|
| Low perturbation | 11 | 10 | 8 | 0 | 0.096 | 0.490 |
| Baseline stress | 11 | 10 | 8 | 0 | 0.162 | 0.640 |
| High perturbation | 12 | 9 | 8 | 1 | 0.197 | 0.758 |
| Intra-athlete random effect | 11 | 10 | 8 | 0 | 0.141 | 0.520 |
| Retained-block surrogate | 11 | 10 | 8 | 0 | 0.158 | 0.625 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the author. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Spoto, S.E. Score Cloud Analysis for Rule-Aware Ranking Robustness Under Discrete Judgment Uncertainty. Stats 2026, 9, 87. https://doi.org/10.3390/stats9050087
Spoto SE. Score Cloud Analysis for Rule-Aware Ranking Robustness Under Discrete Judgment Uncertainty. Stats. 2026; 9(5):87. https://doi.org/10.3390/stats9050087
Chicago/Turabian StyleSpoto, Sebastiano Ettore. 2026. "Score Cloud Analysis for Rule-Aware Ranking Robustness Under Discrete Judgment Uncertainty" Stats 9, no. 5: 87. https://doi.org/10.3390/stats9050087
APA StyleSpoto, S. E. (2026). Score Cloud Analysis for Rule-Aware Ranking Robustness Under Discrete Judgment Uncertainty. Stats, 9(5), 87. https://doi.org/10.3390/stats9050087
