Next Article in Journal
Evaluation of an Educational Campaign to Improve the Conscious Consumption of Recreationally Caught Fish
Next Article in Special Issue
Distribution-Free Stochastic Closed-Loop Supply Chain Design Problem with Financial Management
Previous Article in Journal
Optimal Order Policies for Dual-Sourcing Supply Chains under Random Supply Disruption
Previous Article in Special Issue
Does the Role of Media and Founder’s Past Success Mitigate the Problem of Information Asymmetry? Evidence from a UK Crowdfunding Platform
Open AccessArticle

An Empirical Comparison of Machine-Learning Methods on Bank Client Credit Assessments

Database/Bioinformatics Laboratory, College of Electrical and Computer Engineering, Chungbuk National University, Cheongju 28644, Korea
Microsoft Research, Montreal, QC H3A 3H3, Canada
Department of Information and Computer Sciences, National University of Mongolia, Sukhbaatar District, Building#3 Room#212, Ulaanbaatar 14201, Mongolia
Faculty of Information Technology, Ton Duc Thang University, Ho Chi Minh City 700000, Vietnam
Authors to whom correspondence should be addressed.
Sustainability 2019, 11(3), 699;
Received: 17 December 2018 / Revised: 17 January 2019 / Accepted: 21 January 2019 / Published: 29 January 2019
(This article belongs to the Special Issue Internet Finance, Green Finance and Sustainability)
Machine learning and artificial intelligence have achieved a human-level performance in many application domains, including image classification, speech recognition and machine translation. However, in the financial domain expert-based credit risk models have still been dominating. Establishing meaningful benchmark and comparisons on machine-learning approaches and human expert-based models is a prerequisite in further introducing novel methods. Therefore, our main goal in this study is to establish a new benchmark using real consumer data and to provide machine-learning approaches that can serve as a baseline on this benchmark. We performed an extensive comparison between the machine-learning approaches and a human expert-based model—FICO credit scoring system—by using a Survey of Consumer Finances (SCF) data. As the SCF data is non-synthetic and consists of a large number of real variables, we applied two variable-selection methods: the first method used hypothesis tests, correlation and random forest-based feature importance measures and the second method was only a random forest-based new approach (NAP), to select the best representative features for effective modelling and to compare them. We then built regression models based on various machine-learning algorithms ranging from logistic regression and support vector machines to an ensemble of gradient boosted trees and deep neural networks. Our results demonstrated that if lending institutions in the 2001s had used their own credit scoring model constructed by machine-learning methods explored in this study, their expected credit losses would have been lower, and they would be more sustainable. In addition, the deep neural networks and XGBoost algorithms trained on the subset selected by NAP achieve the highest area under the curve (AUC) and accuracy, respectively. View Full-Text
Keywords: automated credit scoring; decision making; machine learning; internet bank; sustainability automated credit scoring; decision making; machine learning; internet bank; sustainability
Show Figures

Figure 1

MDPI and ACS Style

Munkhdalai, L.; Munkhdalai, T.; Namsrai, O.-E.; Lee, J.Y.; Ryu, K.H. An Empirical Comparison of Machine-Learning Methods on Bank Client Credit Assessments. Sustainability 2019, 11, 699.

Show more citation formats Show less citations formats
Note that from the first issue of 2016, MDPI journals use article numbers instead of page numbers. See further details here.

Article Access Map by Country/Region

Back to TopTop