An Explainability-Driven SHAP-Weighted Ensemble Framework for Fraud Detection: Insights into Model Contribution Dynamics
Abstract
1. Introduction
2. Literature Review
2.1. The Scale and Adaptive Nature of Vehicle Insurance Fraud
2.2. The Machine Learning Paradigm: Capability and Persistent Limitations
2.2.1. The Demonstrated Value of Machine Learning
2.2.2. The Unresolved Imbalance Problem
2.2.3. The Evaluation Problem: Accuracy as a Misleading Metric
2.3. The Ensemble Learning Opportunity: Underexploited in This Domain
2.4. Explainable AI in Fraud Detection: Post Hoc and Architecturally Inert
2.5. Research Gaps and Motivation for an Explanation-Weighted Ensemble Framework
3. Methodology: A Framework for Explanation-Weighted Fraud Detection
3.1. Data Source and Feature Selection
3.2. Preprocessing Pipeline
3.3. Base Learner Selection and Imbalance-Aware Configuration
- Logistic Regression (LR): solver = “liblinear,” class_weight = “balanced,” C = 4.32. Provides a globally interpretable linear probabilistic baseline with well-calibrated output probabilities essential for soft voting;
- Robust Logistic Regression (LR_R): A variant of LR constructed by introducing a RobustScaler preprocessing stage based on the interquartile range (IQR). This design reduces sensitivity to extreme values in total_claim_amount, which frequently exhibits right-skewed distributions in insurance datasets, while retaining full LinearSHAP compatibility;
- Linear Support Vector Machine (SVM): A LinearSVC classifier with Platt scaling via CalibratedClassifierCV, introducing a complementary margin-based perspective. Configuration: C = 1, class_weight = “balanced.”.
3.4. The SHAP-Weighted Ensemble Architecture
3.4.1. Mathematical Formulation
3.4.2. Implementation Details
3.5. Ablation Study Design
3.6. Hyperparameter Optimization and Benchmarking
3.7. Evaluation Protocol
4. Results and Discussion
4.1. Cross-Validation Stability of Base Learners
4.2. Comparative Test-Set Performance with Statistical Inference
4.3. Ablation Study: Weighting Strategy Comparison
4.4. The Performance–Influence Paradox in the Equal-Weight Diagnostic Ensemble
4.5. SHAP Influence Analysis of the Three-Model SWE: Validating the Core Mechanism
4.6. Global Feature Importance: What the Ensemble Is Responding to
4.7. Prediction Correlation Analysis: The Mechanistic Explanation for Performance Parity
4.8. Dynamic Weight Analysis: Evidence of Instance-Adaptive Behavior
4.9. Threshold Optimization and Calibration Analysis
4.10. Cost-Sensitive Evaluation
4.11. Structured Case Study: The SWE Decision Trace in Practice
4.12. Synthesis: Positioning SWE as an Interpretability-Driven Architectural Contribution
5. Conclusions, Limitations, and Future Directions
5.1. Limitations
5.2. Future Directions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Harjai, S.; Khatri, S.K.; Singh, G. Detecting Fraudulent Insurance Claims Using Random Forests and Synthetic Minority Oversampling Technique. In Proceedings of the 2019 4th International Conference on Information Systems and Computer Networks (ISCON), Mathura, India, 21–22 November 2019; pp. 123–128. [Google Scholar]
- Aslam, F.; Hunjra, A.I.; Ftiti, Z.; Louhichi, W.; Shams, T. Insurance fraud detection: Evidence from artificial intelligence and machine learning. Res. Int. Bus. Financ. 2022, 62, 101744. [Google Scholar] [CrossRef]
- Benedek, B.; Ciumas, C.; Nagy, B.Z. Automobile insurance fraud detection in the age of big data—A systematic and comprehensive literature review. J. Financ. Regul. Compliance 2022, 30, 503–523. [Google Scholar] [CrossRef]
- Njeru, A.M. Detection of Fraudulent Vehicle Insurance Claims Using Machine Learning; University of Nairobi: Nairobi, Kenya, 2022. [Google Scholar]
- Mienye, I.D.; Sun, Y. A Survey of Ensemble Learning: Concepts, Algorithms, Applications, and Prospects|IEEE Journals & Magazine|IEEE Xplore. IEEE Access 2022, 10, 99129–99149. [Google Scholar] [CrossRef]
- Sagi, O.; Rokach, L. Ensemble learning: A survey. Wiley Interdiscip. Rev. Data Min. Knowl. Discov. 2018, 8, e1249. [Google Scholar] [CrossRef]
- Bourel, M.; Segura, A.M.; Crisci, C.; López, G.; Sampognaro, L.; Vidal, V.; Kruk, C.; Piccini, C.; Perera, G. Machine learning methods for imbalanced data set for prediction of faecal contamination in beach waters. Water Res. 2021, 202, 117450. [Google Scholar] [CrossRef] [PubMed]
- Khalil, A.A.; Liu, Z.; Fathalla, A.; Ali, A.; Salah, A. Machine Learning Based Method for Insurance Fraud Detection on Class Imbalance Datasets With Missing Values. IEEE Access 2024, 12, 155451–155468. [Google Scholar] [CrossRef]
- Dhieb, N.; Ghazzai, H.; Besbes, H.; Massoud, Y. A Secure AI-Driven Architecture for Automated Insurance Systems: Fraud Detection and Risk Measurement. IEEE Access 2020, 8, 58546–58558. [Google Scholar] [CrossRef]
- Liu, S.; Andrienko, G.; Wu, Y.; Cao, N.; Jiang, L.; Shi, C.; Wang, Y.-S.; Hong, S. Steering data quality with visual analytics: The complexity challenge. Vis. Inform. 2018, 2, 191–197. [Google Scholar] [CrossRef]
- Viaene, S.; Derrig, R.A.; Baesens, B.; Dedene, G. A Comparison of State-of-the-Art Classification Techniques for Expert Automobile Insurance Claim Fraud Detection. J. Risk Insur. 2002, 69, 373–421. [Google Scholar] [CrossRef]
- Dhieb, N.; Ghazzai, H.; Besbes, H.; Massoud, Y. Extreme Gradient Boosting Machine Learning Algorithm For Safe Auto Insurance Operations. In Proceedings of the 2019 IEEE International Conference on Vehicular Electronics and Safety (ICVES), Cairo, Egypt, 4–6 September 2019; pp. 1–5. [Google Scholar]
- Itri, B.; Mohamed, Y.; Mohammed, Q.; Omar, B. Performance comparative study of machine learning algorithms for automobile insurance fraud detection. In Proceedings of the 2019 Third International Conference on Intelligent Computing in Data Sciences (ICDS), Marrakech, Morocco, 28–30 October 2019; pp. 1–4. [Google Scholar]
- Hanafy, M.; Ming, R. Using machine learning models to compare various resampling methods in predicting insurance fraud. J. Theor. Appl. Inf. Technol. 2021, 99, 2819–2833. [Google Scholar]
- Komsrimorakot, P.; Siriborvornratanakul, T. Enhancing fraud detection in imbalanced motor insurance datasets using CP-SMOTE and Random Under-Sampling. J. Big Data 2025, 12, 172. [Google Scholar] [CrossRef]
- Nordin, S.-Z.S.; Wah, Y.B.; Haur, N.K.; Hashim, A.; Rambeli, N.; Jalil, N.A. Predicting automobile insurance fraud using classical and machine learning models. Int. J. Electr. Comput. Eng. (IJECE) 2024, 14, 911–921. [Google Scholar] [CrossRef]
- Kini, A.; Chelluru, R.; Naik, K.; Naik, D.; Aswale, S.; Shetgaonkar, P. Automobile insurance fraud detection: An overview. In Proceedings of the 2022 3rd International Conference on Intelligent Engineering and Management (ICIEM), London, UK, 27–29 April 2022; pp. 7–12. [Google Scholar]
- Khan, A.; Dhungana, K.; Gyawali, S.; Sharma, I. Forecasting Automobile Insurance Claims: Examining the Interplay of Policyholder Traits for Enhanced Predictive Insights. In Proceedings of the 2024 2nd International Conference on Intelligent Data Communication Technologies and Internet of Things (IDCIoT), Bengaluru, India, 4–6 January 2024; pp. 881–886. [Google Scholar]
- Davis, J.; Goadrich, M. The relationship between Precision-Recall and ROC curves. In Proceedings of the 23rd International Conference on Machine Learning, Pittsburgh, PA, USA, 25–29 June 2006; pp. 233–240. [Google Scholar]
- Saito, T.; Rehmsmeier, M. The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets. PLoS ONE 2015, 10, e0118432. [Google Scholar] [CrossRef] [PubMed]
- Vorobyev, I. Fraud risk assessment in car insurance using claims graph features in machine learning. Expert Syst. Appl. 2024, 251, 124109. [Google Scholar] [CrossRef]
- Mokgwatjane, K.; Paepae, T. An explainable ensemble machine learning approach for multi-domain, multiclass sentiment analysis in Amazon product reviews. Mach. Learn. Appl. 2026, 23, 100825. [Google Scholar] [CrossRef]
- Ngwenya, B.; Paepae, T.; Bokoro, P.N. Advancing SDG 6.3. 2 with machine learning-based virtual sensors for high-frequency nutrient monitoring. J. Water Process Eng. 2025, 79, 108831. [Google Scholar] [CrossRef]
- Teffo, N.; Bokoro, P.; Muremi, L.; Paepae, T. Performance evaluation of selected machine learning techniques in the detection of non-technical losses in the distribution system. In Proceedings of the 2024 32nd Southern African Universities Power Engineering Conference (SAUPEC), Stellenbosch, South Africa, 24–25 January 2024; pp. 1–5. [Google Scholar]
- Ding, Y.; Zhu, H.; Chen, R.; Li, R. An Efficient AdaBoost Algorithm with the Multiple Thresholds Classification. Appl. Sci. 2022, 12, 5872. [Google Scholar] [CrossRef]
- General Data Protection Regulation. Art. 22 GDPR. Automated Individual Decision-Making, Including Profiling. Intersoft Consulting 2020. Available online: https://gdpr-info.eu/art-22-gdpr (accessed on 19 March 2026).
- EUAI Act. The eu Artificial Intelligence Act; European Union: Brussels, Belgium, 2024; Available online: https://artificialintelligenceact.eu/ (accessed on 19 March 2026).
- Lundberg, S.M.; Lee, S.-I. A unified approach to interpreting model predictions. Adv. Neural Inf. Process. Syst. 2017, 30, 4768–4777. Available online: https://dl.acm.org/doi/10.5555/3295222.3295230 (accessed on 19 March 2026).
- Debener, J.; Heinke, V.; Kriebel, J. Detecting insurance fraud using supervised and unsupervised machine learning. J. Risk Insur. 2023, 90, 743–768. [Google Scholar] [CrossRef]
- ABDELRAHIM AQQAD. Insurance_Claims. Mendeley Data. 2023. Available online: https://data.mendeley.com/datasets/992mh7dk9y/2 (accessed on 19 March 2026).
- Akiba, T.; Sano, S.; Yanase, T.; Ohta, T.; Koyama, M. Optuna: A next-generation hyperparameter optimization framework. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, New York, NY, USA, 4–8 August 2019; pp. 2623–2631. [Google Scholar]
- Hamon, R.; Junklewitz, H.; Sanchez, I.; Malgieri, G.; De Hert, P. Bridging the gap between AI and explainability in the GDPR: Towards trustworthiness-by-design in automated decision-making. IEEE Comput. Intell. Mag. 2022, 17, 72–85. [Google Scholar] [CrossRef]
- Damen, V.; Wiersma, M.; Aydin, G.; van Haasteren, R. Explainable AI for EU AI Act compliance audits. Maandbl. Voor Account. En Bedrijfsecon. 2025, 99, 231–242. [Google Scholar] [CrossRef]




















| Feature | Type | Preprocessing and Domain Rationale |
|---|---|---|
| POLICYHOLDER CONTEXT | ||
| months_as_customer | Numerical | Standardized (z-score). Low tenure correlates with elevated fraud risk [8]. |
| age | Numerical | Discretized into quartile bins: [0–20], [21–40], [41–65], [65+]. Risk profiles differ systematically by age group. |
| insured_sex | Categorical | Binary encoded. Included as a demographic control variable. |
| INCIDENT CONTEXT | ||
| incident_type | Categorical | One-hot encoded. Primary fraud indicator; “Single Vehicle Collision” is widely associated with staged events. |
| collision_type | Categorical | One-hot encoded. Collision geometry provides discriminative contextual signal. |
| incident_severity | Categorical | One-hot encoded. Disproportionately severe damage relative to collision type flags inflation. |
| authorities_contacted | Categorical | One-hot encoded. Authority choice (police vs. none vs. ambulance) is a procedural fraud indicator. |
| number_of_vehicles_involved | Numerical | Standardized. Higher counts may indicate staged multi-vehicle collisions. |
| police_report_available | Binary | Label encoded. Absence of a police report correlates with fraudulent claims. |
| CLAIM CONTEXT | ||
| total_claim_amount | Numerical | Standardized. Unusually high amounts relative to incident severity are a primary fraud signal. |
| TARGET | ||
| fraud_reported | Binary | Target label: 1 (fraud), 0 (non-fraud). Class prevalence: 24.7% fraud. |
| Model | Mean ROC-AUC | Std. Dev. | Max ROC-AUC |
|---|---|---|---|
| Naive Bayes (NB) | 0.7642 | 0.0350 | 0.8171 |
| Random Forest (RF) | 0.7501 | 0.0362 | 0.8096 |
| Robust LR (LR_R) | 0.7437 | 0.0440 | 0.7923 |
| Logistic Regression (LR) | 0.7433 | 0.0439 | 0.7917 |
| Linear SVM | 0.7431 | 0.0417 | 0.7921 |
| XGBoost | 0.7410 | 0.0503 | 0.8069 |
| AdaBoost | 0.7372 | 0.0365 | 0.7768 |
| Gradient Boosting (GB) | 0.7252 | 0.0474 | 0.7925 |
| Model | ROC-AUC [95% CI] | PR-AUC [95% CI] | F1 [95% CI] | Recall | Precision |
|---|---|---|---|---|---|
| Random Forest (RF) | 0.778 | 0.607 | 0.679 | 0.735 | 0.632 |
| SWE (proposed) | 0.774 [0.681, 0.862] | 0.533 [0.405, 0.695] | 0.679 [0.569, 0.774] | 0.735 | 0.632 |
| SVE (ablation) | 0.776 [0.682, 0.863] | 0.537 [0.409, 0.700] | 0.679 [0.569, 0.774] | 0.735 | 0.632 |
| AWE (ablation) | 0.776 [0.682, 0.863] | 0.537 [0.409, 0.700] | 0.679 [0.569, 0.774] | 0.735 | 0.632 |
| Logistic Regression (LR) | 0.775 | 0.533 | 0.667 | 0.735 | 0.610 |
| Robust LR (LR_R) | 0.775 | 0.533 | 0.667 | 0.735 | 0.610 |
| Linear SVM | 0.777 | 0.540 | 0.667 | 0.673 | 0.660 |
| AdaBoost | 0.786 | 0.531 | 0.641 | 0.673 | 0.611 |
| XGBoost | 0.772 | 0.551 | 0.634 | 0.653 | 0.615 |
| Naive Bayes (NB) | 0.777 | 0.528 | 0.613 | 0.776 | 0.507 |
| Gradient Boosting (GB) | 0.777 | 0.533 | 0.568 | 0.551 | 0.587 |
| Model | Voting Weight | SHAP Influence | Ratio | ROC-AUC | Direction |
|---|---|---|---|---|---|
| Naive Bayes (NB) | 0.143 | 0.228 | 1.59× | 0.777 | Positive |
| Robust LR (LR_R) | 0.143 | 0.158 | 1.11× | 0.775 | Positive |
| Logistic Regression (LR) | 0.143 | 0.158 | 1.11× | 0.775 | Positive |
| Gradient Boosting (GB) | 0.143 | 0.150 | 1.05× | 0.777 | Positive |
| Random Forest (RF) | 0.143 | 0.124 | 0.87× | 0.778 | Negative |
| Linear SVM | 0.143 | 0.122 | 0.85× | 0.776 | Positive |
| AdaBoost | 0.143 | 0.060 | 0.42× | 0.786 | Positive |
| Model | Voting Weight | SHAP Influence | Ratio | ROC-AUC | Direction |
|---|---|---|---|---|---|
| Robust LR | 0.388 | 0.409 | 1.05× | 0.775 | Positive |
| Logistic Regression | 0.386 | 0.407 | 1.05× | 0.775 | Negative |
| Linear SVM | 0.226 | 0.184 | 0.81× | 0.776 | Positive |
| Instance | P(LR) | P(LR_R) | P(SVM) | w(LR) | w(LR_R) | w(SVM) | P_SWE | Outcome |
|---|---|---|---|---|---|---|---|---|
| True Positive (Fraud) | 0.886 | 0.885 | 0.637 | 0.404 | 0.408 | 0.188 | 0.839 | FRAUD (both) |
| False Positive (Legit) | 0.858 | 0.858 | 0.595 | 0.402 | 0.407 | 0.190 | 0.808 | FRAUD (both) |
| True Negative (Legit) | 0.277 | 0.277 | 0.130 | 0.378 | 0.381 | 0.240 | 0.242 | LEGIT (both) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Erasmus, N.C.; Paepae, T. An Explainability-Driven SHAP-Weighted Ensemble Framework for Fraud Detection: Insights into Model Contribution Dynamics. Information 2026, 17, 607. https://doi.org/10.3390/info17060607
Erasmus NC, Paepae T. An Explainability-Driven SHAP-Weighted Ensemble Framework for Fraud Detection: Insights into Model Contribution Dynamics. Information. 2026; 17(6):607. https://doi.org/10.3390/info17060607
Chicago/Turabian StyleErasmus, Nadia Charlene, and Thulane Paepae. 2026. "An Explainability-Driven SHAP-Weighted Ensemble Framework for Fraud Detection: Insights into Model Contribution Dynamics" Information 17, no. 6: 607. https://doi.org/10.3390/info17060607
APA StyleErasmus, N. C., & Paepae, T. (2026). An Explainability-Driven SHAP-Weighted Ensemble Framework for Fraud Detection: Insights into Model Contribution Dynamics. Information, 17(6), 607. https://doi.org/10.3390/info17060607

