TWEEF: Trustworthiness Estimation and Enhancement Framework for Machine Learning Models
Abstract
1. Introduction
2. Related Works
3. Framework Design and Development
3.1. Architecture
3.1.1. Core Component: TrustworthyClassifier
3.1.2. Estimator Layer
- Trustworthiness Estimator: Integrates the outputs of all selected metric dimensions to produce a unified trustworthiness score, reflecting the model’s overall assessment under one of the , or aggregation mechanisms described in Section 3.2.
- Performance Estimator: Computes classical performance metrics such as Accuracy, Precision, Recall, and F1-Score.
- Interpretability Estimator: Computes quantitative interpretability metrics based on a feature-attribution approach (by using the Shapley values [48] method) including Monotonicity, Non-Sensitivity, and Effective Complexity [49]. Monotonicity evaluates the ordinal consistency between feature attributions and their expected predictive impact using Spearman’s rank correlation coefficient, and is scaled to the interval by taking its absolute value. Non-Sensitivity assesses whether only functionally irrelevant features are assigned zero attribution, and is computed as the normalized proportion of attributes whose zero-attribution assignments disagree with their expected predictive relevance. Effective Complexity measures the minimum number of dominant features (with higher absolute attributions) required to preserve predictive performance within a predefined tolerance (set up to 0.95% of total attributions). Then, Effective Complexity is normalized by dividing the number of dominant features by the total number of input features. Interpretability metric normalization ensures comparability and consistent aggregation with performance and fairness dimensions.
3.1.3. Processor Layer
3.1.4. Metrics Layer
3.1.5. Utility Layer
3.2. Aggregation Mechanisms
3.2.1. Linguistic Weighted Average (LWA)
3.2.2. Gaussian Weighted Aggregation (GWA)
3.2.3. Subjective Logic (SL)-Based Aggregation
3.3. Framework Operation Flow
4. Framework Validation
4.1. Experiment 1: German Credit Dataset
4.2. Experiment 2: COMPAS Dataset
4.3. Experiment 3: Adult Income Dataset
4.4. Experiment 4: Exploring Metric Weights
5. Conclusions
6. Limitations and Future Work
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Wing, J.M. Trustworthy AI. Commun. ACM 2021, 64, 64–71. [Google Scholar] [CrossRef]
- Mehrabi, N.; Morstatter, F.; Saxena, N.; Lerman, K.; Galstyan, A. A Survey on Bias and Fairness in Machine Learning. ACM Comput. Surv. 2021, 54, 115. [Google Scholar] [CrossRef]
- Zhou, J.; Gandomi, A.H.; Chen, F.; Holzinger, A. Evaluating the quality of machine learning explanations: A survey on methods and metrics. Electronics 2021, 10, 593. [Google Scholar] [CrossRef]
- Markus, A.F.; Kors, J.A.; Rijnbeek, P.R. The role of explainability in creating trustworthy artificial intelligence for health care: A comprehensive survey of the terminology, design choices, and evaluation strategies. J. Biomed. Inform. 2021, 113, 103655. [Google Scholar] [CrossRef]
- Barocas, S.; Hardt, M.; Narayanan, A. Fairness and Machine Learning: Limitations and Opportunities; MIT Press: Cambridge, MA, USA, 2023. [Google Scholar]
- Burkart, N.; Huber, M.F. A Survey on the Explainability of Supervised Machine Learning. J. Artif. Intell. Res. 2021, 70, 245–317. [Google Scholar] [CrossRef]
- Sharma, S.; Henderson, J.; Ghosh, J. Certifai: A common framework to provide explanations and analyse the fairness and robustness of black-box models. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, New York, NY, USA, 7–9 February 2020; pp. 166–172. [Google Scholar]
- Bellamy, R.K.; Dey, K.; Hind, M.; Hoffman, S.C.; Houde, S.; Kannan, K.; Lohia, P.; Martino, J.; Mehta, S.; Mojsilović, A.; et al. AI Fairness 360: An extensible toolkit for detecting and mitigating algorithmic bias. IBM J. Res. Dev. 2019, 63, 1–4. [Google Scholar] [CrossRef]
- Xu, Z. Linguistic Decision Making, 1st ed.; Springer: Berlin/Heidelberg, Germany, 2013. [Google Scholar] [CrossRef]
- Thiebes, S.; Lins, S.; Sunyaev, A. Trustworthy Artificial Intelligence. Electron. Mark. 2021, 31, 447–464. [Google Scholar] [CrossRef]
- Kowald, D.; Schedl, M.; Lex, E. Establishing and evaluating trustworthy AI: Overview and research challenges. arXiv 2024, arXiv:2406.03012. [Google Scholar] [CrossRef]
- Yin, M.; Vaughan, J.W.; Wallach, H. Understanding the Effect of Accuracy on Trust in Machine Learning Models. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, Glasgow, UK, 4–9 May 2019; ACM: New York, NY, USA, 2019. [Google Scholar] [CrossRef]
- Chatzimparmpas, A.; Martins, R.M.; Jusufi, I.; Kucher, K.; Rossi, F.; Kerren, A. The State of the Art in Enhancing Trust in Machine Learning Models with the Use of Visualizations. Comput. Graph. Forum 2020, 39, 713–756. [Google Scholar] [CrossRef]
- Kusner, M.J.; Loftus, J.; Russell, C.; Silva, R. Counterfactual fairness. Adv. Neural Inf. Process. Syst. 2017, 30. [Google Scholar] [CrossRef]
- Saleiro, P.; Kuester, B.; Hinkson, L.; London, J.; Stevens, A.; Anisfeld, A.; Rodolfa, K.T.; Ghani, R. Aequitas: A bias and fairness audit toolkit. arXiv 2018, arXiv:1811.05577. [Google Scholar]
- Arya, V.; Bellamy, R.K.E.; Chen, P.Y.; Dhurandhar, A.; Hind, M.; Hoffman, S.C.; Houde, S.; Liao, Q.V.; Luss, R.; Mojsilović, A.; et al. AI Explainability 360: An Extensible Toolkit for Understanding Data and Machine Learning Models. J. Mach. Learn. Res. 2020, 21, 1–6. [Google Scholar]
- Nicolae, M.I.; Sinn, M.; Tran, M.N.; Buesser, B.; Rawat, A.; Wistuba, M.; Zantedeschi, V.; Baracaldo, N.; Chen, B.; Ludwig, H.; et al. Adversarial Robustness Toolbox v1.0.0. arXiv 2019, arXiv:1807.01069. [Google Scholar] [CrossRef]
- Verma, S.; Rubin, J. Fairness definitions explained. In Proceedings of the FairWare ’18 International Workshop on Software Fairness, Gothenburg, Sweden, 29 May 2018; pp. 1–7. [Google Scholar] [CrossRef]
- Dwork, C.; Hardt, M.; Pitassi, T.; Reingold, O.; Zemel, R. Fairness through awareness. In Proceedings of the ITCS ’12 3rd Innovations in Theoretical Computer Science Conference, Cambridge, MA, USA, 8–10 January 2012; pp. 214–226. [Google Scholar] [CrossRef]
- Grgic-Hlaca, N.; Zafar, M.B.; Gummadi, K.P.; Weller, A. The Case for Process Fairness in Learning: Feature Selection for Fair Decision making. In Proceedings of the NIPS Symposium on Machine Learning and the Law, Barcelona, Spain, 9 December 2016; Volume 1, p. 11. [Google Scholar]
- Berk, R.; Heidari, H.; Jabbari, S.; Kearns, M.; Roth, A. Fairness in Criminal Justice Risk Assessments: The State of the Art. Sociol. Methods Res. 2021, 50, 3–44. [Google Scholar] [CrossRef]
- Feldman, M.; Friedler, S.A.; Moeller, J.; Scheidegger, C.; Venkatasubramanian, S. Certifying and Removing Disparate Impact. In Proceedings of the KDD ’15 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Sydney, Australia, 10–13 August 2015; pp. 259–268. [Google Scholar] [CrossRef]
- Kamiran, F.; Calders, T. Classifying without discriminating. In Proceedings of the 2009 2nd International Conference on Computer, Control and Communication, Karachi, Pakistan, 17–18 February 2009; IEEE: Piscataway, NJ, USA, 2009; pp. 1–6. [Google Scholar]
- Kamiran, F.; Calders, T. Classification with no discrimination by preferential sampling. In Proceedings of the Machine Learning confer-ence of Belgium and The Netherlands, Leuven, Belgium, 27–28 May 2010; Citeseer: Princeton, NJ, USA, 2010; pp. 1–6. [Google Scholar]
- Kamiran, F.; Calders, T. Data preprocessing techniques for classification without discrimination. Knowl. Inf. Syst. 2012, 33, 1–33. [Google Scholar] [CrossRef]
- Zemel, R.; Wu, Y.; Swersky, K.; Pitassi, T.; Dwork, C. Learning fair representations. In Proceedings of the International Conference on Machine Learning, PMLR, Atlanta, GA, USA, 17–19 June 2013; pp. 325–333. [Google Scholar]
- Kamishima, T.; Akaho, S.; Asoh, H.; Sakuma, J. Fairness-aware classifier with prejudice remover regularizer. In Proceedings of the Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Bristol, UK, 24–28 September 2012; Springer: Berlin/Heidelberg, Germany, 2012; pp. 35–50. [Google Scholar]
- Hardt, M.; Price, E.; Srebro, N. Equality of Opportunity in Supervised Learning. In Proceedings of the Advances in Neural Information Processing Systems; Lee, D., Sugiyama, M., Luxburg, U., Guyon, I., Garnett, R., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2016; Volume 29. [Google Scholar]
- Pleiss, G.; Raghavan, M.; Wu, F.; Kleinberg, J.; Weinberger, K.Q. On fairness and calibration. arXiv 2017, arXiv:1709.02012. [Google Scholar] [CrossRef]
- Goldstein, A.; Kapelner, A.; Bleich, J.; Pitkin, E. Peeking inside the black box: Visualizing statistical learning with plots of individual conditional expectation. J. Comput. Graph. Stat. 2015, 24, 44–65. [Google Scholar] [CrossRef]
- Van Looveren, A.; Klaise, J. Interpretable counterfactual explanations guided by prototypes. arXiv 2019, arXiv:1907.02584. [Google Scholar]
- Lundberg, S.M.; Lee, S.I. A unified approach to interpreting model predictions. In Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, 4–9 December 2017. [Google Scholar]
- Freitas, A.A. Comprehensible classification models: A position paper. ACM SIGKDD Explor. Newsl. 2014, 15, 1–10. [Google Scholar] [CrossRef]
- Ribeiro, M.T.; Singh, S.; Guestrin, C. “Why should i trust you?” Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 1135–1144. [Google Scholar]
- Huysmans, J.; Dejaeger, K.; Mues, C.; Vanthienen, J.; Baesens, B. An empirical evaluation of the comprehensibility of decision table, tree and rule based predictive models. Decis. Support Syst. 2011, 51, 141–154. [Google Scholar] [CrossRef]
- Zhou, J.; Li, Z.; Hu, H.; Yu, K.; Chen, F.; Li, Z.; Wang, Y. Effects of influence on user trust in predictive decision making. In Proceedings of the Extended Abstracts of the 2019 CHI Conference on Human Factors in Computing Systems, Glasgow, UK, 4–9 May 2019; pp. 1–6. [Google Scholar]
- Zhou, J.; Arshad, S.Z.; Yu, K.; Chen, F. Correlation for user confidence in predictive decision making. In Proceedings of the 28th Australian Conference on Computer-Human Interaction, Launceston, Australia, 29 November–2 December 2016; pp. 252–256. [Google Scholar]
- Molnar, C.; Casalicchio, G.; Bischl, B. Interpretable machine learning–a brief history, state-of-the-art and challenges. In Proceedings of the Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Ghent, Belgium, 14–18 September 2020; Springer: Berlin/Heidelberg, Germany, 2020; pp. 417–431. [Google Scholar]
- Slack, D.; Friedler, S.A.; Scheidegger, C.; Roy, C.D. Assessing the local interpretability of machine learning models. arXiv 2019, arXiv:1902.03501. [Google Scholar] [CrossRef]
- Lakkaraju, H.; Kamar, E.; Caruana, R.; Leskovec, J. Interpretable & explorable approximations of black box models. arXiv 2017, arXiv:1707.01154. [Google Scholar] [CrossRef]
- Sundararajan, M.; Taly, A.; Yan, Q. Axiomatic attribution for deep networks. In Proceedings of the International Conference on Machine Learning, PMLR, Sydney, Australia, 6–11 August 2017; pp. 3319–3328. [Google Scholar]
- AL-Sukeinee, R.J.; Khudeyer, R.S. Review: Deep Learning and Fuzzy Logic Applications. Eng. Technol. J. 2024, 9, 4231–4240. [Google Scholar] [CrossRef]
- SS Júnior, J.; Mendes, J.; Souza, F.; Premebida, C. Survey on deep fuzzy systems in regression applications: A view on interpretability. Int. J. Fuzzy Syst. 2023, 25, 2568–2589. [Google Scholar] [CrossRef]
- Fernandez-Peralta, R. A Comprehensive Survey of Fuzzy Implication Functions. arXiv 2025, arXiv:2503.05702. [Google Scholar] [CrossRef]
- Ouifak, H.; Idri, A. A comprehensive review of fuzzy logic based interpretability and explainability of machine learning techniques across domains. Neurocomputing 2025, 647, 130602. [Google Scholar] [CrossRef]
- Cheng, D.; Zhou, X.; Xu, Y. Fuzzy Fusion of Wireless Sensor Network Data for Environmental Monitoring. Expert Syst. Appl. 2012, 39, 11756–11765. [Google Scholar]
- Das, A.K.; Chakraborty, B.; Goswami, S.; Chakrabarti, A. A fuzzy set based approach for effective feature selection. Fuzzy Sets Syst. 2022, 449, 187–206. [Google Scholar] [CrossRef]
- Cohen, S.; Ruppin, E.; Dror, G. Feature selection based on the Shapley value. In Proceedings of the 19th International Joint Conference on Artificial Intelligence (IJCAI-05), Edinburgh, UK, 30 July–5 August 2005; pp. 665–670. [Google Scholar]
- Nguyen, A.; Martínez, M.R. On quantitative aspects of model interpretability. arXiv 2020, arXiv:2007.07584. [Google Scholar] [CrossRef]
- Nori, H.; Jenkins, S.; Koch, P.; Caruana, R. InterpretML: A Unified Framework for Machine Learning Interpretability. arXiv 2019, arXiv:1909.09223. [Google Scholar] [CrossRef]
- Ünver, M. Gaussian Aggregation Operators and Applications to Intuitionistic Fuzzy Classification. J. Classif. 2025, 42, 596–623. [Google Scholar] [CrossRef]
- Jøsang, A. Subjective Logic: A Formalism for Reasoning Under Uncertainty; Springer International Publishing: Cham, Switzerland, 2016. [Google Scholar] [CrossRef]
- Esposito, C.; Galli, A.; Moscato, V.; Sperlí, G. Multi-criteria assessment of user trust in Social Reviewing Systems with subjective logic fusion. Inf. Fusion 2022, 77, 1–18. [Google Scholar] [CrossRef]
- Papenmeier, A.; Englebienne, G.; Seifert, C. How model accuracy and explanation fidelity influence user trust. arXiv 2019, arXiv:1907.12652. [Google Scholar] [CrossRef]







| Dataset | Original Samples | Final Samples | Final Features | Class Balance |
|---|---|---|---|---|
| German Credit | 1000 | 1000 | 48 | 30.0% |
| COMPAS | 7214 | 6172 | 401 | 45.3% |
| Adult | 48,842 | 45,222 | 104 | 24.1% |
| Dimension | Metric(s) |
|---|---|
| Performance | Accuracy, Precision, Recall, F1-Score |
| Fairness | Statistical Parity, Treatment Equality, Equalized Odds, Equal Opportunity |
| Interpretability | Monotonicity, Non-Sensitivity, Effective Complexity |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Ugalde, J.; Salas, R.; Torres, R.; Velandia, D.; Bariviera, A.F.; Estevez, P.A.; Godoy, M.P. TWEEF: Trustworthiness Estimation and Enhancement Framework for Machine Learning Models. Appl. Sci. 2026, 16, 1077. https://doi.org/10.3390/app16021077
Ugalde J, Salas R, Torres R, Velandia D, Bariviera AF, Estevez PA, Godoy MP. TWEEF: Trustworthiness Estimation and Enhancement Framework for Machine Learning Models. Applied Sciences. 2026; 16(2):1077. https://doi.org/10.3390/app16021077
Chicago/Turabian StyleUgalde, Jonathan, Rodrigo Salas, Romina Torres, Daira Velandia, Aurelio F. Bariviera, Pablo A. Estevez, and Maria Paz Godoy. 2026. "TWEEF: Trustworthiness Estimation and Enhancement Framework for Machine Learning Models" Applied Sciences 16, no. 2: 1077. https://doi.org/10.3390/app16021077
APA StyleUgalde, J., Salas, R., Torres, R., Velandia, D., Bariviera, A. F., Estevez, P. A., & Godoy, M. P. (2026). TWEEF: Trustworthiness Estimation and Enhancement Framework for Machine Learning Models. Applied Sciences, 16(2), 1077. https://doi.org/10.3390/app16021077

