RESC: A Rubric-Guided Error–Score Consistency Framework for Action Quality Assessment in Vocational Fitter Training
Abstract
1. Introduction
- 1.
- We propose RESC, a rubric-guided AQA framework for small-sample vocational training. It combines deep pose representations with instructor scoring rules to improve cross-subject scoring and provide traceable error evidence and model-derived deduction contributions for the final score.
- 2.
- We build a video acquisition system in an authentic training workshop and construct the Fitter-AQA dataset. The dataset contains 134 operation videos from 25 participants with different skill levels and provides instructor-consensus scores, error categories, and severity labels for small-sample AQA research in realistic industrial training environments.
- 3.
- We systematically evaluate scoring, ranking, and error diagnosis under strict leave-one-subject-out cross-validation. Comparisons with feature-learning, elastic-alignment, and rubric-guided methods, together with multiple ablation studies, demonstrate the value of structured rubric priors under limited-data conditions.
2. Related Work
2.1. Deep-Learning-Based Action Quality Assessment
2.2. Limited-Sample Learning and Cross-Subject Generalization
2.3. Pose-Aware and Fine-Grained Process Modeling
2.4. Rubric-Driven Interpretable Scoring
3. Materials and Methods
3.1. Data Acquisition System
3.2. Dataset Construction and Quality Control
3.3. Task Definition and Overview of RESC
3.4. Rubric-Guided Pose Geometry
3.5. Error Prototype Diagnosis Module
3.6. Ordered Severity and Adaptive Deduction Gating
3.7. Nonnegative Deductions and Conditional Error–Score Consistency
3.8. Training Objective and Inference
4. Experiments and Results
4.1. Training Environment and Evaluation Metrics
4.2. Comparison with Representative Baselines
4.2.1. Elastic-Alignment Quality Assessment
4.2.2. End-to-End Feature-Learning Methods
4.2.3. Structured Scoring Methods
4.2.4. Controlled Same-Input Comparison and Statistical Stability
4.3. Ablation Studies
4.3.1. Feature-Evidence Ablation
4.3.2. Component Ablation
4.3.3. Counterfactual Directional-Consistency Analysis
4.4. Error Diagnosis
Ordinal Severity Evaluation
4.5. Visualization of Interpretable Scoring
5. Discussion
6. Future Work
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| AQA | Action quality assessment |
| AP | Average precision |
| KRR | Kernel ridge regression |
| LOSO | Leave-one-subject-out |
| QWK | Quadratic-weighted Cohen’s kappa |
Appendix A. RESC Configuration
| Setting | Value or Selection Rule |
|---|---|
| Raw feature dimension | 438. |
| PCA output and retention rule | 12 dimensions, predefined for all folds; fitted within each fold using training participants only. |
| Encoder input dimension | 22, comprising 12 PCA components and 10 directly retained rubric-specific descriptors. |
| Encoder dimensions | 22–128–128. |
| Adaptive-gate dimensions | 134–64–6, using the concatenated 128-dimensional latent representation and six prototype similarities; sigmoid output. |
| Activation | GELU in the encoder and adaptive gate. |
| Normalization | LayerNorm after each encoder linear layer. |
| Dropout | 0.18, applied to the encoder. |
| Optimizer | AdamW. |
| Initial learning rate | 0.001. |
| Weight decay | 0.0002. |
| Batch size | 16. |
| Maximum training epochs | 240. |
| Early stopping | Patience 28, based on the inner-validation objective. |
| Early-stopping objective | Weighted inner-validation loss: deduction score 1.6; error classification 0.75; ordinal severity 0.65. |
| Random seed | 42. |
| Score-loss coefficients | Deduction-chain score 2.3; auxiliary direct score 0.2. |
| Diagnostic-loss coefficients | Quality level 0.35; error classification 1.15. |
| Severity-loss coefficients | Ordinal severity 0.9; expected-severity absolute error 0.2. |
| Structural-loss coefficients | Score consistency 0.55; deduction-weight regularization 0.05. |
| Feature Block | Scalar Descriptors and Aggregation | Dim. |
|---|---|---|
| General pose | Detection confidence and body scale (4); trunk angle (1); bilateral knee angle/flexion (4); bilateral elbow angle/relative height (4); wrist, shoulder, hip, and ankle midpoint coordinates (8); normalized inter-joint distances (7). Seven temporal statistics per scalar plus four sequence summaries. | 200 |
| Trunk lean | Five shoulder–hip–knee–ankle alignment and line-deviation scalars, each summarized by nine temporal statistics. | 45 |
| Knee/support | Six lower-limb angle, flexion, and support-line scalars, each summarized by nine temporal statistics. | 54 |
| Left-elbow raise | Three relative-height and joint-angle scalars, each summarized by nine temporal statistics. | 27 |
| Stroke amplitude | Six wrist-position scalars with nine temporal statistics, plus 12 path-length, range, and displacement summaries. | 66 |
| Hand posture | Four normalized wrist–elbow distance scalars, each summarized by nine temporal statistics. | 36 |
| Narrow stance | One normalized stance-distance scalar summarized by ten temporal and validity statistics. | 10 |
| Total | 438 |
| Baseline | MAE | RMSE | PLCC | SRCC |
|---|---|---|---|---|
| USDL | 2.224 (–5.344) | 3.525 (–7.400) | 0.164 (–0.429) | 0.182 (–0.402) |
| NS-AQA-adapted | 2.259 (0.428–3.863) | 2.189 (0.389–3.926) | 0.100 (0.019–0.198) | 0.182 (0.008–0.398) |
References
- Liu, J.; Wang, H.; Stawarz, K.; Li, S.; Fu, Y.; Liu, H. Vision-based human action quality assessment: A systematic review. Expert Syst. Appl. 2025, 263, 125642. [Google Scholar] [CrossRef] [Scilit]
- Tran, D.; Bourdev, L.; Fergus, R.; Torresani, L.; Paluri, M. Learning Spatiotemporal Features with 3D Convolutional Networks. In Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2015; pp. 4489–4497. [Google Scholar] [CrossRef] [Scilit]
- Carreira, J.; Zisserman, A. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2017; pp. 4724–4733. [Google Scholar] [CrossRef] [Scilit]
- Yan, S.; Xiong, Y.; Lin, D. Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition. Proc. AAAI Conf. Artif. Intell. 2018, 32, 912. [Google Scholar] [CrossRef] [Scilit]
- Doughty, H.; Damen, D.; Mayol-Cuevas, W. Who’s Better? Who’s Best? Pairwise Deep Ranking for Skill Determination. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2018; pp. 6057–6066. [Google Scholar] [CrossRef] [Scilit]
- Yu, X.; Rao, Y.; Zhao, W.; Lu, J.; Zhou, J. Group-aware Contrastive Regression for Action Quality Assessment. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2021; pp. 7899–7908. [Google Scholar] [CrossRef] [Scilit]
- Xu, J.; Rao, Y.; Yu, X.; Chen, G.; Zhou, J.; Lu, J. FineDiving: A Fine-grained Dataset for Procedure-aware Action Quality Assessment. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2022; pp. 2939–2948. [Google Scholar] [CrossRef] [Scilit]
- Xu, J.; Yin, S.; Zhao, G.; Wang, Z.; Peng, Y. FineParser: A Fine-Grained Spatio-Temporal Action Parser for Human-Centric Action Quality Assessment. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2024; pp. 14628–14637. [Google Scholar] [CrossRef] [Scilit]
- Sturm, F.; Hergenroether, E.; Reinhardt, J.; Vojnovikj, P.S.; Siegel, M. Challenges of the Creation of a Dataset for Vision Based Human Hand Action Recognition in Industrial Assembly. In Lecture Notes in Networks and Systems; Springer: Cham, Switzerland, 2023; pp. 1079–1098. [Google Scholar] [CrossRef] [Scilit]
- Zheng, H.; Lee, R.; Lu, Y. HA-ViD: A Human Assembly Video Dataset for Comprehensive Assembly Knowledge Understanding. In Proceedings of the Advances in Neural Information Processing Systems 36; Neural Information Processing Systems Foundation, Inc. (NeurIPS): San Diego, CA, USA, 2023; pp. 67069–67081. [Google Scholar] [CrossRef] [Scilit]
- Sturm, F.; Trat, M.; Sathiyababu, R.; Allipilli, H.; Menz, B.; Hergenroether, E.; Siegel, M. Self-supervised representation learning for robust fine-grained human hand action recognition in industrial assembly lines. Mach. Vis. Appl. 2025, 36, 19. [Google Scholar] [CrossRef] [Scilit]
- Tamantini, C.; Cordella, F.; Lauretti, C.; Zollo, L. The WGD—A Dataset of Assembly Line Working Gestures for Ergonomic Analysis and Work-Related Injuries Prevention. Sensors 2021, 21, 7600. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sener, F.; Chatterjee, D.; Shelepov, D.; He, K.; Singhania, D.; Wang, R.; Yao, A. Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2022; pp. 21064–21074. [Google Scholar] [CrossRef] [Scilit]
- Pirsiavash, H.; Vondrick, C.; Torralba, A. Assessing the Quality of Actions. In Lecture Notes in Computer Science; Springer International Publishing: Berlin/Heidelberg, Germany, 2014; pp. 556–571. [Google Scholar] [CrossRef] [Scilit]
- Parmar, P.; Morris, B.T. Learning to Score Olympic Events. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW); IEEE: Piscataway, NJ, USA, 2017; pp. 76–84. [Google Scholar] [CrossRef] [Scilit]
- Parmar, P.; Morris, B. Action Quality Assessment Across Multiple Actions. In Proceedings of the 2019 IEEE Winter Conference on Applications of Computer Vision (WACV); IEEE: Piscataway, NJ, USA, 2019; pp. 1468–1476. [Google Scholar] [CrossRef] [Scilit]
- Parmar, P.; Morris, B.T. What and How Well You Performed? A Multitask Learning Approach to Action Quality Assessment. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2019; pp. 304–313. [Google Scholar] [CrossRef] [Scilit]
- Feichtenhofer, C.; Fan, H.; Malik, J.; He, K. SlowFast Networks for Video Recognition. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2019; pp. 6201–6210. [Google Scholar] [CrossRef] [Scilit]
- Bertasius, G.; Wang, H.; Torresani, L. Is Space-Time Attention All You Need for Video Understanding? In Proceedings of the 38th International Conference on Machine Learning; PMLR: Red Hook, NY, USA, 2021; Volume 139, pp. 813–824. [Google Scholar]
- Wang, S.; Yang, D.; Zhai, P.; Chen, C.; Zhang, L. TSA-Net: Tube Self-Attention Network for Action Quality Assessment. In Proceedings of the 29th ACM International Conference on Multimedia; ACM: New York, NY, USA, 2021; pp. 4902–4910. [Google Scholar] [CrossRef] [Scilit]
- Bai, Y.; Zhou, D.; Zhang, S.; Wang, J.; Ding, E.; Guan, Y.; Long, Y.; Wang, J. Action Quality Assessment with Temporal Parsing Transformer. In Lecture Notes in Computer Science; Springer Nature: Cham, Switzerland, 2022; pp. 422–438. [Google Scholar] [CrossRef] [Scilit]
- Liu, J.; Wang, H.; Zhou, W.; Stawarz, K.; Corcoran, P.; Chen, Y.; Liu, H. Adaptive Spatiotemporal Graph Transformer Network for Action Quality Assessment. IEEE Trans. Circuits Syst. Video Technol. 2025, 35, 6628–6639. [Google Scholar] [CrossRef] [Scilit]
- Tang, Y.; Ni, Z.; Zhou, J.; Zhang, D.; Lu, J.; Wu, Y.; Zhou, J. Uncertainty-Aware Score Distribution Learning for Action Quality Assessment. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2020; pp. 9836–9845. [Google Scholar] [CrossRef] [Scilit]
- Zhang, B.; Chen, J.; Xu, Y.; Zhang, H.; Yang, X.; Geng, X. Auto-encoding score distribution regression for action quality assessment. Neural Comput. Appl. 2024, 36, 929–942. [Google Scholar] [CrossRef] [Scilit]
- Fu, W.; Fang, W.; Huang, J.; Zhu, K.; Chen, R.; Feng, C. Skeleton-Based Action Quality Assessment with Anomaly-Aware DTW Optimization for Intelligent Sports Education. Sensors 2025, 25, 7160. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yanik, E.; Schwaitzberg, S.; Yang, G.; Intes, X.; Norfleet, J.; Hackett, M.; De, S. One-shot skill assessment in high-stakes domains with limited data via meta learning. Comput. Biol. Med. 2024, 174, 108470. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhu, X.; Toisoul, A.S.; Perez-Rua, J.M.; Zhang, L.; Martinez, B.; Xiang, T. Few-shot Action Recognition with Prototype-centered Attentive Learning. In Proceedings of the British Machine Vision Conference 2021; British Machine Vision Association: Durham, UK, 2021. [Google Scholar] [CrossRef] [Scilit]
- Li, S.; Liu, H.; Qian, R.; Li, Y.; See, J.; Fei, M.; Yu, X.; Lin, W. TA2N: Two-Stage Action Alignment Network for Few-Shot Action Recognition. Proc. AAAI Conf. Artif. Intell. 2022, 36, 1404–1411. [Google Scholar] [CrossRef] [Scilit]
- Gao, Y.; Vedula, S.S.; Reiley, C.E.; Ahmidi, N.; Varadarajan, B.; Lin, H.C.; Tao, L.; Zappella, L.; Bejar, B.; Yuh, D.D.; et al. JHU-ISI Gesture and Skill Assessment Working Set (JIGSAWS): A Surgical Activity Dataset for Human Motion Modeling. In Proceedings of the MICCAI Workshop on Modeling and Monitoring of Computer Assisted Interventions, Boston, MA, USA, 14–18 September 2014. [Google Scholar]
- Kasa, K.; Burns, D.; Goldenberg, M.G.; Selim, O.; Whyne, C.; Hardisty, M. Multi-Modal Deep Learning for Assessing Surgeon Technical Skill. Sensors 2022, 22, 7328. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Benmansour, M.; Malti, A.; Jannin, P. Deep neural network architecture for automated soft surgical skills evaluation using objective structured assessment of technical skills criteria. Int. J. Comput. Assist. Radiol. Surg. 2023, 18, 929–937. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, D.; Li, Q.; Jiang, T.; Wang, Y.; Miao, R.; Shan, F.; Li, Z. Towards Unified Surgical Skill Assessment. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2021; pp. 9517–9526. [Google Scholar] [CrossRef] [Scilit]
- Zhao, W.; Wang, L.; Li, Y.; Liu, X.; Zhang, Y.; Yan, B.; Li, H. A Multi-Scale and Multi-Stage Human Pose Recognition Method Based on Convolutional Neural Networks for Non-Wearable Ergonomic Evaluation. Processes 2024, 12, 2419. [Google Scholar] [CrossRef] [Scilit]
- Senjaya, W.F.; Yahya, B.N.; Lee, S.L. Sensor-Based Motion Tracking System Evaluation for RULA in Assembly Task. Sensors 2022, 22, 8898. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shao, D.; Zhao, Y.; Dai, B.; Lin, D. FineGym: A Hierarchical Video Dataset for Fine-Grained Action Understanding. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2020; pp. 2613–2622. [Google Scholar] [CrossRef] [Scilit]
- Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. In Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2017; pp. 618–626. [Google Scholar] [CrossRef] [Scilit]
- Kim, B.; Wattenberg, M.; Gilmer, J.; Cai, C.; Wexler, J.; Viegas, F.; Sayres, R. Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV). In Proceedings of the 35th International Conference on Machine Learning; PMLR: Red Hook, NY, USA, 2018; Volume 80, pp. 2668–2677. [Google Scholar]
- Chen, C.; Li, O.; Tao, C.; Barnett, A.J.; Su, J.; Rudin, C. This Looks Like That: Deep Learning for Interpretable Image Recognition. In Proceedings of the Advances in Neural Information Processing Systems 32, Vancouver, BC, Canada, 8–14 December 2019. [Google Scholar]
- Koh, P.W.; Nguyen, T.; Tang, Y.S.; Mussmann, S.; Pierson, E.; Kim, B.; Liang, P. Concept Bottleneck Models. In Proceedings of the 37th International Conference on Machine Learning; PMLR: Red Hook, NY, USA, 2020; Volume 119, pp. 5338–5348. [Google Scholar]
- Han, R.; Zhou, K.; Atapour-Abarghouei, A.; Liang, X.; Shum, H.P. FineCausal: A Causal-Based Framework for Interpretable Fine-Grained Action Quality Assessment. In Proceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW); IEEE: Piscataway, NJ, USA, 2025; pp. 6008–6017. [Google Scholar] [CrossRef] [Scilit]
- Matsuyama, H.; Kawaguchi, N.; Lim, B.Y. IRIS: Interpretable Rubric-Informed Segmentation for Action Quality Assessment. In Proceedings of the 28th International Conference on Intelligent User Interfaces; ACM: New York, NY, USA, 2023; pp. 368–378. [Google Scholar] [CrossRef] [Scilit]
- Majeedi, A.; Gajjala, V.R.; GNVV, S.S.S.N.; Li, Y. RICA2: Rubric-Informed, Calibrated Assessment of Actions. In Lecture Notes in Computer Science; Springer Nature: Cham, Switzerland, 2024; pp. 143–161. [Google Scholar] [CrossRef] [Scilit]
- Rai, A.; Kovashka, A. Rubric-Constrained Figure Skating Scoring. In Proceedings of the 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV); IEEE: Piscataway, NJ, USA, 2025; pp. 9105–9113. [Google Scholar] [CrossRef] [Scilit]
- Okamoto, L.; Parmar, P. Hierarchical NeuroSymbolic Approach for Comprehensive and Explainable Action Quality Assessment. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW); IEEE: Piscataway, NJ, USA, 2024; pp. 3204–3213. [Google Scholar] [CrossRef] [Scilit]
- You, S.; Ding, D.; Canini, K.; Pfeifer, J.; Gupta, M. Deep Lattice Networks and Partial Monotonic Functions. In Proceedings of the Advances in Neural Information Processing Systems 30, Long Beach, CA, USA, 4–9 December 2017. [Google Scholar]
- Wehenkel, A.; Louppe, G. Unconstrained Monotonic Neural Networks. In Proceedings of the Advances in Neural Information Processing Systems 32, Vancouver, BC, Canada, 8–14 December 2019. [Google Scholar]
- Runje, D.; Shankaranarayana, S.M. Constrained Monotonic Neural Networks. In Proceedings of the 40th International Conference on Machine Learning; PMLR: Red Hook, NY, USA, 2023; Volume 202, pp. 29338–29353. [Google Scholar]
- Sakoe, H.; Chiba, S. Dynamic programming algorithm optimization for spoken word recognition. IEEE Trans. Acoust. Speech Signal Process. 1978, 26, 43–49. [Google Scholar] [CrossRef] [Scilit]
- Keogh, E.J.; Pazzani, M.J. Derivative Dynamic Time Warping. In Proceedings of the 2001 SIAM International Conference on Data Mining. Society for Industrial and Applied Mathematics, Chicago, IL, USA, 5–7 April 2001; pp. 1–11. [Google Scholar] [CrossRef] [Scilit]
- Jeong, Y.S.; Jeong, M.K.; Omitaomu, O.A. Weighted dynamic time warping for time series classification. Pattern Recognit. 2011, 44, 2231–2240. [Google Scholar] [CrossRef] [Scilit]
- Zhao, J.; Itti, L. shapeDTW: Shape Dynamic Time Warping. Pattern Recognit. 2018, 74, 171–184. [Google Scholar] [CrossRef] [Scilit]
- Cuturi, M.; Blondel, M. Soft-DTW: A Differentiable Loss Function for Time-Series. In Proceedings of the 34th International Conference on Machine Learning; PMLR: Red Hook, NY, USA, 2017; Volume 70, pp. 894–903. [Google Scholar]
- Marteau, P.F. Time Warp Edit Distance with Stiffness Adjustment for Time Series Matching. IEEE Trans. Pattern Anal. Mach. Intell. 2009, 31, 306–318. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Herrmann, M.; Webb, G.I. Amercing: An intuitive and effective constraint for dynamic time warping. Pattern Recognit. 2023, 137, 109333. [Google Scholar] [CrossRef] [Scilit]
- Shi, L.; Zhang, Y.; Cheng, J.; Lu, H. Two-Stream Adaptive Graph Convolutional Networks for Skeleton-Based Action Recognition. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2019; pp. 12018–12027. [Google Scholar] [CrossRef] [Scilit]
- Chen, Y.; Zhang, Z.; Yuan, C.; Li, B.; Deng, Y.; Hu, W. Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action Recognition. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2021; pp. 13339–13348. [Google Scholar] [CrossRef] [Scilit]









| (a) Error Definitions and Pose Evidence | ||
| Error category | Instructional definition | Pose evidence |
| Trunk lean | Insufficient or excessive forward trunk lean, or pronounced trunk sway during filing. | Angle and temporal variation of the right shoulder–hip–knee–ankle alignment. |
| Knee posture | Knee hyperextension, insufficient or excessive flexion, or unstable lower-body support. | Left hip–knee–ankle angle and its temporal variation. |
| Left-elbow raise | Noncompliant left-elbow height or failure to maintain a stable elbow position. | Relative height and joint angle of the left shoulder, elbow, and wrist. |
| Stroke amplitude | The filing stroke is too short and the effective reciprocating range is insufficient. | Horizontal trajectory lengths of both wrists. |
| Hand posture | Grip, force application, or coordination between the two hands does not comply with the operating requirements. | Elbow–wrist configuration, inter-wrist distance, and wrist-position proxy features. |
| Narrow stance | Foot spacing is too narrow to provide a stable base for standing and force application. | Inter-ankle distance, its body-scale-normalized value, and temporal variation. |
| (b) Common ordinal severity scale | ||
| Severity level | Operational definition | |
| 0—Absent | No observable violation of the corresponding rubric criterion. | |
| 1—Mild | A small or occasional deviation with limited influence on execution. | |
| 2—Moderate | A clear or recurrent deviation that affects stability or compliance. | |
| 3—Severe | A pronounced or persistent deviation that substantially affects action quality. | |
| Method | Core Interpretation | Adaptation to Fitter-AQA | Training Settings |
|---|---|---|---|
| RICA2-adapted | Models rubric criteria as graph nodes and propagates criterion evidence to an uncertainty-aware score. | The graph was rebuilt with six error nodes and one score root. Outputs were aligned to error evidence, severity, and continuous score using the common I3D input. | One-layer Transformer, hidden dimension 64, four heads, dropout 0.15; AdamW, learning rate 0.001, weight decay 0.002; maximum 260 epochs, patience 35; 24 Monte Carlo samples; seeds 17, 42, and 73. |
| RCS-adapted | Uses text-defined rubric queries to extract criterion evidence and aggregate criterion scores. | The six fitter rubric descriptions replaced the original task criteria, and temporal attention learned criterion-relevant segments directly from video-level supervision. | Same temporal backbone and optimizer as RICA2-adapted; maximum 260 epochs, patience 35; attention regularization 0.02; seeds 17, 42, and 73. |
| IRIS-adapted | Uses rubric-item queries to obtain item-level evidence and subscores before global scoring. | Six video-level error items replaced action-segment criteria, and weak item localization learned item-relevant evidence from video-level labels. | Same temporal backbone and optimizer as RICA2-adapted; maximum 260 epochs, patience 35; attention regularization 0.02; seeds 17, 42, and 73. |
| NS-AQA-adapted | Represents criteria as neural symbols and combines ordered severity states through nonnegative rules. | Six rubric-specific pose groups generated four-level severity symbols and rule-based score deductions. | Per-criterion MLP, hidden dimension 32, dropout 0.15; AdamW, learning rate 0.002, weight decay 0.002; maximum 500 epochs, patience 50; seeds 17, 42, and 73. |
| Method | MAE ↓ | RMSE ↓ | PLCC ↑ | SRCC ↑ |
|---|---|---|---|---|
| CDTW-KRR [48] | 8.991 | 11.951 | 0.488 | 0.418 |
| DDTW-KRR [49] | 8.506 | 10.286 | 0.692 | 0.652 |
| WDTW-KRR [50] | 9.161 | 12.220 | 0.463 | 0.407 |
| ShapeDTW-KRR [51] | 9.100 | 12.110 | 0.473 | 0.407 |
| Soft-DTW-KRR [52] | 8.362 | 11.282 | 0.544 | 0.459 |
| TWED-KRR [53] | 7.730 | 10.158 | 0.666 | 0.533 |
| ADTW-KRR [54] | 9.070 | 12.102 | 0.474 | 0.413 |
| RESC | 4.450 | 5.870 | 0.902 | 0.852 |
| Method | MAE ↓ | RMSE ↓ | PLCC ↑ | SRCC ↑ |
|---|---|---|---|---|
| ST-GCN [4] | 9.178 | 11.783 | 0.491 | 0.467 |
| 2s-AGCN [55] | 9.011 | 11.807 | 0.550 | 0.498 |
| USDL [23] | 6.674 | 9.395 | 0.738 | 0.670 |
| CoRe [6] | 8.352 | 10.468 | 0.630 | 0.632 |
| CTR-GCN [56] | 9.371 | 11.938 | 0.535 | 0.482 |
| TPT [21] | 8.832 | 11.886 | 0.567 | 0.586 |
| RESC | 4.450 | 5.870 | 0.902 | 0.852 |
| Method | MAE ↓ | RMSE ↓ | PLCC ↑ | SRCC ↑ |
|---|---|---|---|---|
| IRIS-adapted [41] | 8.627 | 11.199 | 0.563 | 0.577 |
| RICA2-adapted [42] | 8.074 | 10.830 | 0.616 | 0.607 |
| NS-AQA-adapted [44] | 6.709 | 8.059 | 0.802 | 0.670 |
| RCS-adapted [43] | 8.569 | 10.868 | 0.591 | 0.584 |
| RESC | 4.450 | 5.870 | 0.902 | 0.852 |
| Method | MAE ↓ | RMSE ↓ | PLCC ↑ | SRCC ↑ |
|---|---|---|---|---|
| Same-input Ridge | 7.286 (4.943–9.591) | 9.162 (6.432–11.407) | 0.745 (0.588–0.887) | 0.710 (0.408–0.859) |
| Same-input MLP | 6.732 (3.769–11.380) | 10.007 (5.124–15.466) | 0.691 (0.302–0.926) | 0.579 (0.058–0.816) |
| Auxiliary direct-score head | 4.580 (3.095–6.396) | 6.264 (4.628–8.041) | 0.890 (0.800–0.941) | 0.833 (0.629–0.876) |
| Full RESC | 4.450 (3.199–6.139) | 5.870 (4.367–7.701) | 0.902 (0.837–0.944) | 0.852 (0.665–0.887) |
| Input Evidence | MAE ↓ | RMSE ↓ | PLCC ↑ | SRCC ↑ |
|---|---|---|---|---|
| General pose statistics | 6.749 | 9.264 | 0.714 | 0.644 |
| Rubric geometry | 7.012 | 9.291 | 0.712 | 0.721 |
| General pose statistics + rubric geometry | 4.450 | 5.870 | 0.902 | 0.852 |
| Variant | MAE ↓ | RMSE ↓ | PLCC ↑ | SRCC ↑ |
|---|---|---|---|---|
| w/o error prototypes | 7.546 | 8.614 | 0.696 | 0.645 |
| w/o ordinal severity | 5.769 | 6.771 | 0.797 | 0.744 |
| w/o adaptive deduction gate | 5.808 | 5.963 | 0.802 | 0.767 |
| w/o monotonicity | 4.503 | 6.076 | 0.880 | 0.804 |
| w/o consistency | 4.507 | 5.899 | 0.908 | 0.831 |
| RESC | 4.450 | 5.870 | 0.902 | 0.852 |
| Rubric Item | Adjacent Interventions | Item-Deduction Consistency (%) | Final-Score Consistency (%) | Maximum Reverse Score Change (Points) |
|---|---|---|---|---|
| Trunk lean | 4763 | 100.00 | 100.00 | 0.0000 |
| Knee posture | 4699 | 100.00 | 100.00 | 0.0000 |
| Left-elbow raise | 4841 | 100.00 | 100.00 | 0.0000 |
| Stroke amplitude | 5169 | 100.00 | 100.00 | 0.0000 |
| Hand posture | 4882 | 100.00 | 100.00 | 0.0000 |
| Narrow stance | 4955 | 100.00 | 95.40 | 0.0016 |
| Overall | 29,309 | 100.00 | 99.22 | 0.0016 |
| Metric | Value |
|---|---|
| Discrete-level MAE | 0.496 |
| Exact ordinal accuracy (%) | 72.76 |
| Within-one-level accuracy (%) | 85.82 |
| Quadratic-weighted kappa (QWK) | 0.631 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Mu, W.; Wang, S.; Huang, J. RESC: A Rubric-Guided Error–Score Consistency Framework for Action Quality Assessment in Vocational Fitter Training. Sensors 2026, 26, 5575. https://doi.org/10.3390/s26175575
Mu W, Wang S, Huang J. RESC: A Rubric-Guided Error–Score Consistency Framework for Action Quality Assessment in Vocational Fitter Training. Sensors. 2026; 26(17):5575. https://doi.org/10.3390/s26175575
Chicago/Turabian StyleMu, Wenfeng, Shigang Wang, and Jiajia Huang. 2026. "RESC: A Rubric-Guided Error–Score Consistency Framework for Action Quality Assessment in Vocational Fitter Training" Sensors 26, no. 17: 5575. https://doi.org/10.3390/s26175575
APA StyleMu, W., Wang, S., & Huang, J. (2026). RESC: A Rubric-Guided Error–Score Consistency Framework for Action Quality Assessment in Vocational Fitter Training. Sensors, 26(17), 5575. https://doi.org/10.3390/s26175575

