XGBoost versus LightGBM for Loan Approval Prediction: A Controlled, Like-for-Like Comparison
DOI:
https://doi.org/10.54536/ajdsai.v2i2.8054Keywords:
Credit Risk, Gradient Boosting, Lightgbm, Loan Approval Prediction, XGBoostAbstract
Deciding whether to approve a loan application is one of the highest-stakes decisions a lender makes: approving a risky applicant invites default, while turning away a sound one loses a good customer. Predicting the approval decision accurately is therefore central to comprehensive credit management. Gradient boosting has become the dominant approach for such tabular credit data, and two libraries, XGBoost and LightGBM, account for most of its practical use. Although both are widely applied to credit-risk and loan-approval prediction, they are seldom compared directly under matched conditions, which leaves little firm evidence for choosing between them; this study addresses that gap. A controlled, like-for-like comparison was carried out in which both models were trained on the same dataset of 45,000 loan applications using one pre-processing pipeline and a matched set of structural hyper-parameters, and were then evaluated with accuracy, precision, recall, F1-score and the area under the ROC curve. The comparison was re-examined with five-fold stratified cross-validation and repeated on a second, independent loan-approval dataset of a different character. On the primary dataset, XGBoost reached an accuracy of 93.5% and an ROC-AUC of 0.979, a small margin above LightGBM on every metric, and it retained the higher mean AUC under cross-validation; on the second, high-signal dataset both models exceeded 98% accuracy and XGBoost again held the higher cross-validated AUC, winning four of five folds. Across both dataset types, XGBoost therefore demonstrated a consistent, though modest, improvement in discrimination, an effect compatible with its level-wise tree growth and stronger regularization. The principal limitations are the use of a single matched configuration and two historical datasets. The findings indicate that both libraries are accurate and deployable, with XGBoost showing a marginal but reproducible edge in predictive discrimination.
References
Akinjole, A., Shobayo, O., Popoola, J., Okoyeigbo, O., & Ogunleye, B. (2024). Ensemble-based machine learning algorithm for loan default risk prediction. Mathematics, 12(21), Article 3423. https://doi.org/10.3390/math12213423
Alonso Robisco, A., & Carbo Martinez, J. M. (2022). Measuring the model risk-adjusted performance of machine learning algorithms in credit default prediction. Financial Innovation, 8(1), Article 70. https://doi.org/10.1186/s40854-022-00366-1
Alvi, J., Arif, I., & Nizam, K. (2024). Advancing financial resilience: A systematic review of default prediction models and future directions in credit risk management. Heliyon, 10(21), Article e39770. https://doi.org/10.1016/j.heliyon.2024.e39770
Anghel, A., Papandreou, N., Parnell, T., De Palma, A., & Pozidis, H. (2018). Benchmarking and optimization of gradient boosting decision tree algorithms. arXiv. https://arxiv.org/abs/1809.04559
Barbaglia, L., Manzan, S., & Tosetti, E. (2023). Forecasting loan default in Europe with machine learning. Journal of Financial Econometrics, 21(2), 569–596.
Bentejac, C., Csorgo, A., & Martinez-Munoz, G. (2021). A comparative analysis of gradient boosting algorithms. Artificial Intelligence Review, 54(3), 1937–1967. https://doi.org/10.1007/s10462-020-09896-5
Chang, V., Sivakulasingam, S., Wang, H., Wong, S. T., Ganatra, M. A., & Luo, J. (2024). Credit risk prediction using machine learning and deep learning: A study on credit card customers. Risks, 12(11), Article 174. https://doi.org/10.3390/risks12110174
Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). https://doi.org/10.1145/2939672.2939785
Emmanuel, I., Sun, Y., & Wang, Z. (2024). A machine learning-based credit risk prediction engine system using a stacked classifier and a filter-based feature selection method. Journal of Big Data, 11(1), Article 23. https://doi.org/10.1186/s40537-024-00882-0
Gu, Z., Lv, J., Wu, B., Hu, Z., & Yu, X. (2024). Credit risk assessment of small and micro enterprise based on machine learning. Heliyon, 10(5), Article e27096. https://doi.org/10.1016/j.heliyon.2024.e27096
Hancock, J. T., & Khoshgoftaar, T. M. (2020). CatBoost for big data: An interdisciplinary review. Journal of Big Data, 7(1), Article 94. https://doi.org/10.1186/s40537-020-00369-8
Haque, F. M. A., & Hassan, M. M. (2024). Bank loan prediction using machine learning techniques. American Journal of Industrial and Business Management, 14(12), 1690–1711. https://www.scirp.org/pdf/ajibm20241412_72123507.pdf
Hlongwane, R., Ramaboa, K. K. M., & Mongwe, W. (2024a). Enhancing credit scoring accuracy with a comprehensive evaluation of alternative data. PLOS ONE, 19(5), Article e0303566. https://doi.org/10.1371/journal.pone.0303566
Hlongwane, R., Ramaboa, K. K. M., & Mongwe, W. (2024b). A novel framework for enhancing transparency in credit scoring: Leveraging Shapley values for interpretable credit scorecards. PLOS ONE, 19(8), Article e0308718. https://doi.org/10.1371/journal.pone.0308718
Hossain, M. M., Mamun, M., Munir, A., Rahman, M. M., & Chowdhury, S. H. (2025). A secure bank loan prediction system by bridging differential privacy and explainable machine learning. Electronics, 14(8), Article 1691. https://doi.org/10.3390/electronics14081691
Hou, G., Tong, D. L., Liew, S. Y., & Choo, P. Y. (2025). Comparative analysis of resampling techniques for class imbalance in financial distress prediction using XGBoost. Mathematics, 13(13), Article 2186. https://doi.org/10.3390/math13132186
Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., & Liu, T.-Y. (2017). LightGBM: A highly efficient gradient boosting decision tree. In Advances in Neural Information Processing Systems (Vol. 30, pp. 3146–3154).
Ko, P.-C., Lin, P.-C., Do, H.-T., & Huang, Y.-F. (2022). P2P lending default prediction based on AI and statistical models. Entropy, 24(6), Article 801. https://doi.org/10.3390/e24060801
Kunwar, S., & Ghimire, B. R. (2026). Leakage-free early prediction of issue resolution time in agile software projects: A comparative study on dataset quality. American Journal of Data Science and Artificial Intelligence, 2(1), 58–68. https://doi.org/10.54536/ajdsai.v2i1.7792
Lawal, O., Okolie, A., Obunadike, C., Akwabeng, P. M., Ikhifa, M. O., & Alumona, P. (2026). Predicting food insecurity across U.S. census tracts: A machine learning analysis using the USDA Food Access Research Atlas. American Journal of Data Science and Artificial Intelligence, 2(1), 1–14. https://doi.org/10.54536/ajdsai.v2i1.6285
Markov, A., Seleznyova, Z., & Lapshin, V. (2022). Credit scoring methods: Latest trends and points to consider. The Journal of Finance and Data Science, 8, 180–201. https://doi.org/10.1016/j.jfds.2022.07.002
Miljkovic, T., & Wang, P. (2025). A dimension reduction assisted credit scoring method for big data with categorical features. Financial Innovation, 11(1), Article 29. https://doi.org/10.1186/s40854-024-00689-1
Nguyen, N., & Ngo, D. (2025). Comparative analysis of boosting algorithms for predicting personal default. Cogent Economics & Finance, 13(1), Article 2465971. https://doi.org/10.1080/23322039.2025.2465971
Ozkurt, C. (2024). Enhancing financial decision-making: Predictive modeling for personal loan eligibility with gradient boosting, XGBoost, and AdaBoost. Information Technology in Economics and Business, 1(1), 7–13. https://doi.org/10.69882/adba.iteb.2024072
Yang, C. (2024). Research on loan approval and credit risk based on the comparison of machine learning models. SHS Web of Conferences, 181, Article 02003. https://doi.org/10.1051/shsconf/202418102003
Ying, C., Shi, A., & Li, X. (2025). Hybrid boosted attention-based LightGBM framework for enhanced credit risk assessment in digital finance. Humanities and Social Sciences Communications, 12(1), Article 1036. https://doi.org/10.1057/s41599-025-05230-y
Zhu, M., Zhang, Y., Gong, Y., Xing, K., Yan, X., & Song, J. (2024). Ensemble methodology: Innovations in credit default prediction using LightGBM, XGBoost, and LocalEnsemble [Preprint]. arXiv. https://arxiv.org/abs/2402.17979
Zhu, X., Chu, Q., Song, X., Hu, P., & Peng, L. (2023). Explainable prediction of loan default based on machine learning models. Data Science and Management, 6(3), 123–133. https://doi.org/10.1016/j.dsm.2023.04.003
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Sagar Poudel, Dr. Bhoj Raj Ghimire

This work is licensed under a Creative Commons Attribution 4.0 International License.