The Point-of-Balanced Performance in Binary Classification: How to Simplify the Performance Metrics Ecosystem
Abstract
1. Introduction
2. Modeling of the Binary Classification Problem
3. Performance Metrics for Binary Classification
4. The Concept of Equivalent Performance Metrics
5. The Point-of-Balanced Performance (PoBP)
5.1. The Classification Threshold of the Point-of-Balanced Performance
5.2. Performance Metrics That Become Equal at the Point-of-Balanced Performance
5.3. Performance Metrics That Become Equivalent at the Point-of-Balanced Performance
6. PoBP at the Receiver Operating Characteristic Curve
7. Methods for the Identification of the Point-of-Balanced Performance
7.1. Methods for the Identification of the Point-of-Balanced Performance During Inference
7.2. Evaluation of the PoBP Identification Methods During Inference
8. Performance at the PoBP vs. Typical Characteristic Points
9. Discussion and Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Appendix A. Binary Classification Problems and Applied Classifiers
Appendix A.1. Binary Classification Problems
| Problem | Problem Title | Number of Features | Dataset Size | Prevalence |
|---|---|---|---|---|
| Problem 1 | Bank Analysis [35] | 20 | 41,188 | 89% |
| Problem 2 | German Data [36] | 24 | 1000 | 70% |
| Problem 3 | Adult Income [37] | 14 | 32,561 | 76% |
| Problem 4 | The Rain Problem [38] | 22 | 145,461 | 76% |
| Problem 5 | Student Depression [39] | 16 | 27,901 | 41% |
Appendix A.2. Applied Binary Classifiers
Appendix B
Appendix B.1. Matthews Correlation Coefficient (MCC) Analysis
Appendix B.2. Cohen’s Kappa Analysis
Appendix C. Pairs of Equivalent Metrics
Appendix C.1. Equivalence of Jaccard Index (TS) and F1-Score
Appendix C.2. Equivalence of Balanced Accuracy (BA) and Informedness (BM)
Appendix C.3. Equivalence of EI to FiC
Appendix C.4. Equivalence of Prevalence Threshold to LR+
Appendix D. Equivalence of Metrics at the PoBP
Appendix D.1. Equivalence of Accuracy with PPV at the Point-of-Balanced Performance
Appendix D.2. Equivalence of TNR to TPR at the Point-of-Balanced Performance
Appendix D.3. Equivalence of Informedness to TPR at the Point-of-Balanced Performance
Appendix D.4. Equivalence of P4 to TPR at the Point-of-Balanced Performance
Appendix D.5. Equivalence of LR+ to TPR at the Point-of-Balanced Performance
Appendix D.6. Equivalence of LR− to TPR at the Balanced Point of Performance
Appendix D.7. Equivalence of DOR to TPR at the Point-of-Balanced Performance
Appendix D.8. Equivalence of G-Mean Score with TPR at the Point-of-Balanced Performance
References
- Alpaydin, E. Introduction to Machine Learning; The MIT Press: Cambridge, MA, USA, 2010; ISBN 978-0-262-01243-0. [Google Scholar]
- Sarker, I.H. Machine Learning: Algorithms, Real-World Applications and Research Directions. SN Comput. Sci. 2021, 2, 160. [Google Scholar] [CrossRef] [Scilit]
- Garg, A.; Roth, D. Understanding probabilistic classifiers. In Proceedings of the ECML 2001 12th European Conference on Machine Learning, LNAI 2167, Freiburg, Germany, 5–7 September 2001; pp. 179–191. [Google Scholar]
- Stehman, S.V. Selecting and interpreting measures of thematic classification accuracy. Remote Sens. Environ. 1997, 62, 77–89. [Google Scholar] [CrossRef] [Scilit]
- Ting, K.M. Confusion Matrix. In Encyclopedia of Machine Learning and Data Mining; Springer: Berlin/Heidelberg, Germany, 2010. [Google Scholar]
- Uddin, S.; Khan, A.; Hossain, M.; Moni, M.A. Comparing different supervised machine learning algorithms for disease prediction. BMC Med. Inform. Decis. Mak. 2019, 19, 281. [Google Scholar] [CrossRef] [Scilit]
- Chicco, D.; Jurman, G. The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genom. 2020, 21, 6. [Google Scholar] [CrossRef] [Scilit]
- Cao, C.; Chicco, D.; Hoffman, M.M. The MCC-F1 curve: A performance evaluation technique for binary classification. arXiv 2020, arXiv:2006.11278. [Google Scholar] [CrossRef] [Scilit]
- Chicco, D.; Jurman, G. The Matthews correlation coefficient (MCC) should replace the ROC AUC as the standard metric for assessing binary classification. BioData Min. 2023, 16, 4. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Chicco, D.; Warrens, M.; Jurman, G. The Matthews Correlation Coefficient (MCC) is More Informative Than Cohen’s Kappa and Brier Score in Binary Classification Assessment. IEEE Access 2021, 9, 78368–78381. [Google Scholar] [CrossRef] [Scilit]
- Chicco, D.; Jurman, G. A statistical comparison between Matthews correlation coefficient (MCC), prevalence threshold, and Fowlkes–Mallows index. J. Biomed. Inform. 2023, 144, 104426. [Google Scholar] [CrossRef] [Scilit]
- Itaya, Y.; Tamura, J.; Hayashi, K.; Yamamoto, K. Asymptotic Properties of Matthews Correlation Coefficient. Stat. Med. 2025, 44, e10303. [Google Scholar] [CrossRef] [Scilit]
- Zhu, Q. On the performance of Matthews correlation coefficient (MCC) for imbalanced dataset. Pattern Recognit. Lett. 2020, 136, 71–80. [Google Scholar] [CrossRef] [Scilit]
- Shao, G.; Tang, L.; Liao, J. Overselling overall map accuracy misinforms about research reliability. Landsc. Ecol. 2019, 34, 2487–2492. [Google Scholar] [CrossRef] [Scilit]
- Leeflang, M.M.; Bossuyt, P.M.; Irwig, L. Diagnostic test accuracy may vary with prevalence: Implications for evidence-based diagnosis. J. Clin. Epidemiol. 2009, 62, 5–12. [Google Scholar] [CrossRef] [Scilit]
- Powers, D. The problem with kappa. In Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics (EACL ‘12); Association for Computational Linguistics: Avignon, France, 2012; pp. 345–355. [Google Scholar]
- Powers, D.M.W. Evaluation: From Precision, Recall and F-Measure to ROC, Informedness, Markedness & Correlation. arXiv 2011, arXiv:2010.16061. [Google Scholar]
- Larner, A.J. Efficiency Index for Binary Classifiers: Concept, Extension, and Application. Mathematics 2023, 11, 2435. [Google Scholar] [CrossRef] [Scilit]
- Kuncheva, L.I.; Arnaiz-González, Á.; Díez-Pastor, J.-F.; Gunn, I.A.D. Instance selection improves geometric mean accuracy: A study on imbalanced data classification. Prog. Artif. Intell. 2019, 8, 215–228. [Google Scholar] [CrossRef] [Scilit]
- Marra, A. G4 & the balanced metric family—A novel approach to solving binary classification problems in medical device validation & verification studies. BioData Min. 2024, 17, 43. [Google Scholar] [CrossRef] [Scilit]
- Lipton, Z.C.; Elkan, C.; Naryanaswamy, B. Optimal Thresholding of Classifiers to Maximize F1 Measure. In Proceedings of the Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD; Springer: Berlin/Heidelberg, Germany, 2014; Volume 8725, pp. 225–239. [Google Scholar]
- Canbek, G.; Sagiroglu, S.; Temizel, T.T.; Baykal, N. Binary classification performance measures/metrics: A comprehensive visualized roadmap to gain new insights. In 2017 International Conference on Computer Science and Engineering (UBMK); IEEE: New York, NY, USA, 2017; pp. 821–826. [Google Scholar] [CrossRef] [Scilit]
- Hernández-Orallo, J.; Flach, P.; Ferri, C. A unified view of performance metrics: Translating threshold choice into expected classification loss. J. Mach. Learn. Res. 2012, 13, 2813–2869. [Google Scholar]
- Christen, P.; Hand, D.J.; Kirielle, N. A Review of the F-Measure: Its History, Properties, Criticism, and Alternatives. ACM Comput. Surv. 2023, 56, 73. [Google Scholar] [CrossRef] [Scilit]
- Hand, D.J.; Christen, P.; Kirielle, N. F*: An interpretable transformation of the F-measure. Mach Learn. 2021, 110, 451–456. [Google Scholar] [CrossRef] [Scilit] [PubMed] [PubMed Central]
- Kim, H.; Park, Y. Calibrating F1 Scores for Fair Performance Comparison of Binary Classification Models with Application to Student Dropout Prediction. IEEE Access 2025, 13, 136554–136567. [Google Scholar] [CrossRef] [Scilit]
- Walauskis, M.A.; Khoshgoftaar, T.M. Choosing the Right Metrics: A Study of Performance Measurement for Binary Classification in Imbalanced and Big Data. Int. FLAIRS Conf. Proc. 2025, 38. [Google Scholar] [CrossRef] [Scilit]
- Canbek, G.; Taskaya Temizel, T.; Sagiroglu, S. BenchMetrics: A systematic benchmarking method for binary classification performance metrics. Neural Comput. Applic 2021, 33, 14623–14650. [Google Scholar] [CrossRef] [Scilit]
- Shirdel, M.; Di Mauro, M.; Liotta, A. Worthiness Benchmark: A novel concept for analyzing binary classification evaluation metrics. Inf. Sci. 2024, 678, 120882. [Google Scholar] [CrossRef] [Scilit]
- Fawcett, T. An introduction to ROC analysis. Pattern Recognit. Lett. 2006, 27, 861–874. [Google Scholar] [CrossRef] [Scilit]
- Hanley, J.A.; McNeil, B.J. The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology 1982, 143, 29–36. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Refaeilzadeh, P.; Tang, L.; Liu, H. Cross-Validation. In Encyclopedia of Database Systems; Liu, L., Özsu, M.T., Eds.; Springer: Boston, MA, USA, 2009. [Google Scholar] [CrossRef] [Scilit]
- Markoulidakis, I.; Markoulidakis, G. Probabilistic Confusion Matrix: A Novel Method for Machine Learning Algorithm Generalized Performance Analysis. Technologies 2024, 12, 113. [Google Scholar] [CrossRef] [Scilit]
- Niculescu-Mizil, A.; Caruana, R. Predicting good probabilities with supervised learning. In Proceedings of the 22nd International Conference on Machine Learning (ICML’ 05), Bonn, Germany, 7–11 August 2005; Association for Computing Machinery: New York, NY, USA, 2005; pp. 625–632. [Google Scholar]
- Moro, S.; Cortez, P.; Rita, P. A data-driven approach to predict the success of bank telemarketing. Decis. Support Syst. 2014, 62, 22–31. [Google Scholar] [CrossRef] [Scilit]
- Hofmann, H. Statlog (German Credit Data). Machine Learning Repository. 1994, Volume 53. Available online: https://archive.ics.uci.edu/dataset/144/statlog+german+credit+data (accessed on 20 February 2026).
- Becker, B.; Kohavi, R. Adult [Dataset]. UCI Machine Learning Repository. 1996. Available online: https://archive.ics.uci.edu/dataset/2/adult (accessed on 20 February 2026).
- Australia Meteorology Government Bureau. Available online: http://www.bom.gov.au/ (accessed on 20 February 2026).
- Garcia-Ceja, E.; Riegler, M.; Jakobsen, P.; Tørresen, J.; Nordgreen, T.; Oedegaard, K.J.; Fasmer, O.B. A Motor Activity Database of Depression Episodes in Unipolar and Bipolar Patients. In Proceedings of the MMSys’18 9th ACM on Multimedia Systems Conference, Amsterdam, The Netherlands, 12–15 June 2018; Available online: https://dl.acm.org/doi/pdf/10.1145/3204949.3208125 (accessed on 20 February 2026).
- Quinlan, J.R. Induction of decision trees. Mach. Learn. 1986, 1, 81–106. [Google Scholar] [CrossRef] [Scilit]
- Rokach, L.; Maimon, O.Z. Data Mining with Decision Trees: Theory and Applications; World Scientific: Singapore, 2008; Volume 69. [Google Scholar]
- Basak, D.; Srimanta, P.; Patranabis, D.C. Support Vector Regression. Neural Inf. Process.-Lett. Rev. 2007, 11, 203–224. [Google Scholar]
- Abe, S. Support Vector Machines for Pattern Classification, 2nd ed.; Advances in Computer Vision and Pattern Recognition; Springer: London, UK, 2010. [Google Scholar] [CrossRef] [Scilit]
- Cortes, C.; Vapnik, V. Support-vector networks. Mach. Learn. 1995, 20, 273–297. [Google Scholar] [CrossRef] [Scilit]
- Pal, M. Random forest classifier for remote sensing classification. Int. J. Remote Sens. 2005, 26, 217–222. [Google Scholar] [CrossRef] [Scilit]
- Breiman, L. Random forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef] [Scilit]
- Walczak, S.; Cerpa, N. Artificial Neural Networks, Encyclopedia of Physical Science and Technology, 3rd ed.; Academic Press: Cambridge, MA, USA, 2003; pp. 631–645. ISBN 9780122274107. [Google Scholar] [CrossRef] [Scilit]
- Doulamis, A.; Doulamis, N.; Kollias, S. On-line retrainable neural networks: Improving the performance of neural networks in image analysis problems. IEEE Trans. Neural Netw. 2000, 11, 137–155. [Google Scholar] [CrossRef] [PubMed]
- Haykin, S. Neural Networks: A Comprehensive Foundation; Prentice-Hall Inc.: Upper Anhe, NJ, USA, 2007. [Google Scholar]
- Doulamis, A.; Doulamis, N.; Protopapadakis, E.; Voulodimos, A. Combined Convolutional Neural Networks and Fuzzy Spectral Clustering for Real Time Crack Detection in Tunnels. In Proceedings of the 2018 25th IEEE International Conference on Image Processing (ICIP), Athens, Greece, 7–10 October 2018; pp. 4153–4157. [Google Scholar] [CrossRef] [Scilit]
- Albawi, S.; Mohammed, T.A.; Al-Zawi, S. Understanding of a convolutional neural network. In Proceedings of the 2017 International Conference on Engineering and Technology (ICET), Antalya, Turkey, 21–23 August 2017; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Chauhan, R.; Ghanshala, K.K.; Joshi, R.C. Convolutional Neural Network (CNN) for Image Detection and Recognition. In Proceedings of the 2018 1st International Conference on Secure Cyber Computing and Communication (ICSCCC), Jalandhar, India, 15–17 December 2018; pp. 278–282. [Google Scholar] [CrossRef] [Scilit]
- Haouari, B.; Amor, N.B.; Elouedi, Z.; Mellouli, K. Naïve possibilistic network classifiers. Fuzzy Sets Syst. 2009, 160, 3224–3238. [Google Scholar] [CrossRef] [Scilit]
- Hosmer, D.W., Jr.; Lemeshow, S.; Sturdivant, R.X. Applied Logistic Regression; JohnWiley & Sons: Hoboken, NJ, USA, 2013; Volume 398. [Google Scholar]
- Chen, T.; Guestrin, C. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd ACM Sigkdd International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 785–794. [Google Scholar]





| Actual Positive | Actual Negative | Total | |
|---|---|---|---|
| Predicted Positive | TP | FP | TP + FP = PP |
| Predicted Negative | FN | TN | FN + TN = PN |
| Total | TP + FN = P | FP + TN = N | P + N = PP + PN = S |
| A/A | Metric Name | Symbol and Definition |
|---|---|---|
| 1 | Accuracy | |
| 2 | Fraction of Inefficiency | |
| 3 | True Positive Rate (Recall, Sensitivity) | |
| 4 | True Negative Rate (Specificity) | |
| 5 | Balanced Accuracy | |
| 6 | False Positive Rate (Fall-Out Rate) | |
| 7 | False Negative Rate (Miss Rate) | |
| 8 | Positive Predictive Value (Precision) | |
| 9 | Negative Predictive Value | |
| 10 | False Discovery Rate | |
| 11 | False Omission Rate | |
| 12 | F1-Score | |
| 13 | Fβ-Score | |
| 14 | Cohen’s kappa (see Appendix B) | |
| 15 | Youden’s J Statistic (Informedness) | |
| 16 | Markedness | |
| 17 | Matthews Correlation Coefficient (see Appendix B) | |
| 18 | Fowlkes–Mallows Index | |
| 19 | Jaccard Index (Threat Score) | |
| 20 | Prevalence Threshold | |
| 21 | P4 metric | |
| 22 | Positive Likelihood Ratio | |
| 23 | Negative Likelihood Ratio | |
| 24 | Diagnostics Odd Ratio | |
| 25 | Efficiency Index | |
| 26 | Inefficiency Index | |
| 27 | G-Mean Score | |
| 28 | G4 Score |
| ANN | LogReg | RF | DT | NB | CNN | XGB | |
|---|---|---|---|---|---|---|---|
| Accuracy | 0.916 | 0.906 | 0.904 | 0.905 | 0.905 | 0.913 | 0.918 |
| Fraction of Inefficiency | 0.084 | 0.094 | 0.096 | 0.095 | 0.095 | 0.087 | 0.082 |
| True Positive Rate (Recall, Sensitivity) | 0.569 | 0.377 | 0.276 | 0.536 | 0.344 | 0.429 | 0.561 |
| True Negative Rate (Specificity) | 0.959 | 0.978 | 0.982 | 0.953 | 0.976 | 0.974 | 0.964 |
| Balanced Accuracy | 0.764 | 0.678 | 0.629 | 0.745 | 0.660 | 0.701 | 0.762 |
| False Positive Rate (Fall-Out Rate) | 0.041 | 0.022 | 0.018 | 0.047 | 0.024 | 0.026 | 0.036 |
| False Negative Rate (Miss Rate) | 0.431 | 0.623 | 0.724 | 0.464 | 0.656 | 0.571 | 0.439 |
| Positive Predictive Value (Precision) | 0.638 | 0.700 | 0.654 | 0.602 | 0.645 | 0.678 | 0.669 |
| Negative Predictive Value | 0.946 | 0.920 | 0.917 | 0.940 | 0.922 | 0.931 | 0.944 |
| False Discovery Rate | 0.362 | 0.300 | 0.346 | 0.398 | 0.355 | 0.322 | 0.331 |
| False Omission Rate | 0.054 | 0.080 | 0.083 | 0.060 | 0.078 | 0.069 | 0.056 |
| F1-Score | 0.602 | 0.490 | 0.388 | 0.567 | 0.449 | 0.525 | 0.610 |
| Fβ-Score (β = 2) | 0.623 | 0.597 | 0.513 | 0.588 | 0.549 | 0.607 | 0.644 |
| Cohen’s kappa | 0.555 | 0.443 | 0.345 | 0.514 | 0.402 | 0.480 | 0.564 |
| Informedness | 0.528 | 0.355 | 0.258 | 0.490 | 0.320 | 0.403 | 0.525 |
| Markedness | 0.585 | 0.620 | 0.570 | 0.542 | 0.567 | 0.608 | 0.613 |
| Mathews’ Correlation Coefficient | 0.556 | 0.469 | 0.384 | 0.515 | 0.426 | 0.495 | 0.567 |
| Fowlkes–Mallows Index | 0.603 | 0.514 | 0.425 | 0.568 | 0.471 | 0.539 | 0.612 |
| Jaccard Index (Threat Score) | 0.430 | 0.325 | 0.241 | 0.396 | 0.289 | 0.356 | 0.439 |
| Prevalence Threshold | 0.211 | 0.195 | 0.204 | 0.228 | 0.209 | 0.197 | 0.202 |
| P4 metric | 0.184 | 0.162 | 0.138 | 0.177 | 0.152 | 0.169 | 0.186 |
| Positive Likelihood Ratio | 13.998 | 17.098 | 15.292 | 11.487 | 14.391 | 16.559 | 15.586 |
| Negative Likelihood Ratio | 0.449 | 0.637 | 0.737 | 0.486 | 0.672 | 0.586 | 0.456 |
| Diagnostics Odd Ratio | 31.165 | 26.849 | 20.748 | 23.623 | 21.419 | 28.235 | 34.214 |
| Efficiency Index | 10.856 | 9.632 | 9.463 | 9.497 | 9.564 | 10.450 | 11.162 |
| Inefficiency Index | 0.092 | 0.104 | 0.106 | 0.105 | 0.105 | 0.096 | 0.090 |
| G-Mean Score | 0.739 | 0.607 | 0.521 | 0.715 | 0.580 | 0.646 | 0.735 |
| G4 Score | 0.758 | 0.698 | 0.635 | 0.734 | 0.668 | 0.716 | 0.764 |
| Groups of Equivalent Metrics | Equations |
|---|---|
| Accuracy (Acc) Fraction of Inefficiency (FiC) Efficiency Index (EI) Inefficiency Index (InI) | |
| Balanced Accuracy (BA) Informedness (BM) | |
| Recall (TPR) Miss Rate (FNR) | |
| Specificity (TNR) Fall Out Rate (FPR) | |
| Precision (PPV) False Discovery Rate (FDR) | |
| Negative Predictive Value (NPV) False Omission Rate (FOR) | |
| F1-Score (F1) Jaccard Index (TS) | |
| Prevalence Threshold (PT) Positive Likelihood Ratio (LR+) | |
| Groups of Metrics | Equality at the Point-of-Balanced Performance |
|---|---|
| True Positive Rate (TPR) Precision (PPV) F1-score (F1) Fβ-score (Fβ) Fowlkes–Mallows Index (FM) | |
| Informedness (BM) Markedness (MK) Matthews Correlation Coefficient (MCC) Cohen’s kappa (k) | |
| True Negative Rate (TNR) Negative Predictive Value (NPV) | |
| G-Mean Score G4 Score |
| ANN | LogReg | RF | DT | NB | CNN | XGB | |
|---|---|---|---|---|---|---|---|
| Accuracy | 0.9135 | 0.8932 | 0.8829 | 0.9035 | 0.8914 | 0.9108 | 0.9144 |
| Fraction of Inefficiency | 0.0865 | 0.11 | 0.12 | 0.10 | 0.11 | 0.089 | 0.0856 |
| True Positive Rate | 0.61 | 0.55 | 0.47 | 0.59 | 0.51 | 0.60 | 0.63 |
| True Negative Rate | 0.9515 | 0.94 | 0.93 | 0.95 | 0.94 | 0.950 | 0.9517 |
| Balanced Accuracy | 0.78 | 0.75 | 0.70 | 0.77 | 0.73 | 0.78 | 0.79 |
| False Positive Rate | 0.049 | 0.06 | 0.07 | 0.055 | 0.06 | 0.050 | 0.048 |
| False Negative Rate | 0.39 | 0.45 | 0.53 | 0.41 | 0.49 | 0.40 | 0.37 |
| Positive Predictive Value | 0.61 | 0.55 | 0.47 | 0.59 | 0.51 | 0.60 | 0.63 |
| Negative Predictive Value | 0.951 | 0.94 | 0.93 | 0.945 | 0.94 | 0.950 | 0.952 |
| False Discovery Rate | 0.39 | 0.45 | 0.53 | 0.41 | 0.49 | 0.40 | 0.37 |
| False Omission Rate | 0.049 | 0.061 | 0.066 | 0.055 | 0.061 | 0.050 | 0.048 |
| F1-Score | 0.61 | 0.55 | 0.47 | 0.59 | 0.51 | 0.60 | 0.63 |
| Fβ-Score (β = 2) | 0.61 | 0.55 | 0.47 | 0.59 | 0.51 | 0.60 | 0.63 |
| Cohen’s kappa | 0.56 | 0.49 | 0.40 | 0.53 | 0.45 | 0.55 | 0.58 |
| Informedness | 0.56 | 0.49 | 0.40 | 0.53 | 0.45 | 0.55 | 0.58 |
| Markedness | 0.57 | 0.49 | 0.40 | 0.53 | 0.45 | 0.55 | 0.58 |
| Matthews Correlation Coefficient | 0.56 | 0.49 | 0.40 | 0.53 | 0.45 | 0.55 | 0.58 |
| Fowlkes–Mallows Index | 0.61 | 0.55 | 0.47 | 0.59 | 0.51 | 0.60 | 0.63 |
| Jaccard Index | 0.44 | 0.38 | 0.30 | 0.41 | 0.35 | 0.43 | 0.46 |
| Prevalence Threshold | 0.220 | 0.249 | 0.273 | 0.234 | 0.256 | 0.224 | 0.217 |
| P4 metric | 0.186 | 0.174 | 0.156 | 0.181 | 0.166 | 0.185 | 0.189 |
| Positive Likelihood Ratio | 12.62 | 9.14 | 7.10 | 10.73 | 8.42 | 12.03 | 12.98 |
| Negative Likelihood Ratio | 0.41 | 0.47 | 0.57 | 0.44 | 0.52 | 0.42 | 0.39 |
| Diagnostics Odd Ratio | 31.00 | 19.27 | 12.44 | 24.50 | 16.30 | 28.86 | 33.10 |
| Efficiency Index | 10.56 | 8.36 | 7.54 | 9.36 | 8.21 | 10.21 | 10.69 |
| Inefficiency Index | 0.095 | 0.120 | 0.133 | 0.107 | 0.122 | 0.098 | 0.094 |
| G-Mean Score | 0.76 | 0.72 | 0.66 | 0.74 | 0.70 | 0.76 | 0.77 |
| G4 Score | 0.76 | 0.72 | 0.66 | 0.74 | 0.70 | 0.76 | 0.77 |
| Problem | Classifier | Test Set | Set Size | Method 1 | Method 2 | Method 3 |
|---|---|---|---|---|---|---|
| 1 | SVM | Unbiased | 8238 | 0% | 2.8% | 1.5% |
| Biased | 5945 | 14.3% | 5.6% | 6.1% | ||
| 2 | RF | Unbiased | 200 | 0% | 2.1% | 0.8% |
| Biased | 172 | 8.4% | 5.8% | 6.8% | ||
| 3 | NB | Unbiased | 6513 | 0% | 0.4% | 0.1% |
| Biased | 5564 | 6.3% | 3.6% | 5.2% | ||
| 4 | XGB | Unbiased | 29,092 | 0% | 0.5% | 0.3% |
| Biased | 21,847 | 6.2% | 2.9% | 3.5% | ||
| 5 | LogReg | Unbiased | 5581 | 0% | 3% | 1.6% |
| Biased | 4383 | 13.2% | 7.3% | 7.6% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Markoulidakis, I.; Markoulidakis, G. The Point-of-Balanced Performance in Binary Classification: How to Simplify the Performance Metrics Ecosystem. Technologies 2026, 14, 139. https://doi.org/10.3390/technologies14030139
Markoulidakis I, Markoulidakis G. The Point-of-Balanced Performance in Binary Classification: How to Simplify the Performance Metrics Ecosystem. Technologies. 2026; 14(3):139. https://doi.org/10.3390/technologies14030139
Chicago/Turabian StyleMarkoulidakis, Ioannis, and Georgios Markoulidakis. 2026. "The Point-of-Balanced Performance in Binary Classification: How to Simplify the Performance Metrics Ecosystem" Technologies 14, no. 3: 139. https://doi.org/10.3390/technologies14030139
APA StyleMarkoulidakis, I., & Markoulidakis, G. (2026). The Point-of-Balanced Performance in Binary Classification: How to Simplify the Performance Metrics Ecosystem. Technologies, 14(3), 139. https://doi.org/10.3390/technologies14030139
