Comparative Study of Traditional and Machine Learning Models for SME Credit Risk Assessment
- DOI
- 10.2991/978-94-6239-774-3_34How to use a DOI?
- Keywords
- SME credit risk; small business lending; logistic regression; support vector machine; random forest; XGBoost; SBA 7(a) loans; ROC-AUC; precision-recall analysis; feature importance
- Abstract
Small and medium-sized enterprise (SME) credit risk assessment is an essential challenge for financial institutions because SME borrowers are characterized by high economic relevance, information opacity, diverse operating histories, and sensitivity to loan terms. This study presents an empirical comparison of traditional and machine learning models for SME credit risk assessment using official U.S. Small Business Administration 7(a) Freedom of Information Act (FOIA) open data. Loans approved in FY2010-FY2019 are used to develop the models, while loans approved from FY2020 through the current release are used as an out-of-time test sample. Charged-off loans are treated as default observations, and paid-in-full loans are treated as non-default observations. To minimize information leakage, variables that reveal post-origination outcomes are excluded. Logistic regression is used as the traditional benchmark, and its performance is compared with three challenger models: linear support vector machine, random forest, and XGBoost. The empirical results show that XGBoost performs best on the out-of-time test set in terms of accuracy, precision, recall, F1-score, ROC-AUC, PR-AUC, and Brier score (accuracy 0.8981, precision 0.4416, recall 0.8199, F1-score 0.5741, ROC-AUC 0.9193, PR-AUC 0.5838, and Brier score 0.0754). Feature analysis indicates that maturity, interest-rate structure, secondary-market sale status, SBA subprogram, lender geography, revolving status, guarantee amount, and business age are important risk signals. The findings confirm that machine learning can improve SME default identification under a temporal validation design while preserving financially meaningful model interpretation. Tree-based learning provides stronger discrimination and recall and captures nonlinear interaction effects more effectively, whereas logistic regression remains a useful transparent baseline.
- Copyright
- © 2026 The Author(s)
- Open Access
- Open Access This chapter is licensed under the terms of the Creative Commons Attribution-NonCommercial 4.0 International License (http://creativecommons.org/licenses/by-nc/4.0/), which permits any noncommercial use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons license and indicate if changes were made.
Cite this article
TY - CONF AU - Siqi Huang PY - 2026 DA - 2026/09/11 TI - Comparative Study of Traditional and Machine Learning Models for SME Credit Risk Assessment BT - Proceedings of the 2026 5th International Conference on Mathematical Statistics and Economic Analysis (MSEA 2026 PB - Atlantis Press SP - 358 EP - 371 SN - 2352-5428 UR - https://doi.org/10.2991/978-94-6239-774-3_34 DO - 10.2991/978-94-6239-774-3_34 ID - Huang2026 ER -