XGBoost versus LightGBM for Loan Approval Prediction: A Controlled, Like-for-Like Comparison

Authors

  • Sagar Poudel Faculty of Science, Health and Technology, Nepal Open University, Lalitpur, Nepal
  • Bhoj Raj Ghimire Faculty of Science, Health and Technology, Nepal Open University, Lalitpur, Nepal

DOI:

https://doi.org/10.54536/ajdsai.v2i2.8054

Keywords:

Credit Risk, Gradient Boosting, Lightgbm, Loan Approval Prediction, XGBoost

Abstract

Deciding whether to approve a loan application is one of the highest-stakes decisions a lender makes: approving a risky applicant invites default, while turning away a sound one loses a good customer. Predicting the approval decision accurately is therefore central to comprehensive credit management. Gradient boosting has become the dominant approach for such tabular credit data, and two libraries, XGBoost and LightGBM, account for most of its practical use. Although both are widely applied to credit-risk and loan-approval prediction, they are seldom compared directly under matched conditions, which leaves little firm evidence for choosing between them; this study addresses that gap. A controlled, like-for-like comparison was carried out in which both models were trained on the same dataset of 45,000 loan applications using one pre-processing pipeline and a matched set of structural hyper-parameters, and were then evaluated with accuracy, precision, recall, F1-score and the area under the ROC curve. The comparison was re-examined with five-fold stratified cross-validation and repeated on a second, independent loan-approval dataset of a different character. On the primary dataset, XGBoost reached an accuracy of 93.5% and an ROC-AUC of 0.979, a small margin above LightGBM on every metric, and it retained the higher mean AUC under cross-validation; on the second, high-signal dataset both models exceeded 98% accuracy and XGBoost again held the higher cross-validated AUC, winning four of five folds. Across both dataset types, XGBoost therefore demonstrated a consistent, though modest, improvement in discrimination, an effect compatible with its level-wise tree growth and stronger regularization. The principal limitations are the use of a single matched configuration and two historical datasets. The findings indicate that both libraries are accurate and deployable, with XGBoost showing a marginal but reproducible edge in predictive discrimination.

References

Akinjole, A., Shobayo, O., Popoola, J., Okoyeigbo, O., & Ogunleye, B. (2024). Ensemble-based machine learning algorithm for loan default risk prediction. Mathematics, 12(21), Article 3423. https://doi.org/10.3390/math12213423

Alonso Robisco, A., & Carbo Martinez, J. M. (2022). Measuring the model risk-adjusted performance of machine learning algorithms in credit default prediction. Financial Innovation, 8(1), Article 70. https://doi.org/10.1186/s40854-022-00366-1

Alvi, J., Arif, I., & Nizam, K. (2024). Advancing financial resilience: A systematic review of default prediction models and future directions in credit risk management. Heliyon, 10(21), Article e39770. https://doi.org/10.1016/j.heliyon.2024.e39770

Anghel, A., Papandreou, N., Parnell, T., De Palma, A., & Pozidis, H. (2018). Benchmarking and optimization of gradient boosting decision tree algorithms. arXiv. https://arxiv.org/abs/1809.04559

Barbaglia, L., Manzan, S., & Tosetti, E. (2023). Forecasting loan default in Europe with machine learning. Journal of Financial Econometrics, 21(2), 569–596.

Bentejac, C., Csorgo, A., & Martinez-Munoz, G. (2021). A comparative analysis of gradient boosting algorithms. Artificial Intelligence Review, 54(3), 1937–1967. https://doi.org/10.1007/s10462-020-09896-5

Chang, V., Sivakulasingam, S., Wang, H., Wong, S. T., Ganatra, M. A., & Luo, J. (2024). Credit risk prediction using machine learning and deep learning: A study on credit card customers. Risks, 12(11), Article 174. https://doi.org/10.3390/risks12110174

Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). https://doi.org/10.1145/2939672.2939785

Emmanuel, I., Sun, Y., & Wang, Z. (2024). A machine learning-based credit risk prediction engine system using a stacked classifier and a filter-based feature selection method. Journal of Big Data, 11(1), Article 23. https://doi.org/10.1186/s40537-024-00882-0

Gu, Z., Lv, J., Wu, B., Hu, Z., & Yu, X. (2024). Credit risk assessment of small and micro enterprise based on machine learning. Heliyon, 10(5), Article e27096. https://doi.org/10.1016/j.heliyon.2024.e27096

Hancock, J. T., & Khoshgoftaar, T. M. (2020). CatBoost for big data: An interdisciplinary review. Journal of Big Data, 7(1), Article 94. https://doi.org/10.1186/s40537-020-00369-8

Haque, F. M. A., & Hassan, M. M. (2024). Bank loan prediction using machine learning techniques. American Journal of Industrial and Business Management, 14(12), 1690–1711. https://www.scirp.org/pdf/ajibm20241412_72123507.pdf

Hlongwane, R., Ramaboa, K. K. M., & Mongwe, W. (2024a). Enhancing credit scoring accuracy with a comprehensive evaluation of alternative data. PLOS ONE, 19(5), Article e0303566. https://doi.org/10.1371/journal.pone.0303566

Hlongwane, R., Ramaboa, K. K. M., & Mongwe, W. (2024b). A novel framework for enhancing transparency in credit scoring: Leveraging Shapley values for interpretable credit scorecards. PLOS ONE, 19(8), Article e0308718. https://doi.org/10.1371/journal.pone.0308718

Hossain, M. M., Mamun, M., Munir, A., Rahman, M. M., & Chowdhury, S. H. (2025). A secure bank loan prediction system by bridging differential privacy and explainable machine learning. Electronics, 14(8), Article 1691. https://doi.org/10.3390/electronics14081691

Hou, G., Tong, D. L., Liew, S. Y., & Choo, P. Y. (2025). Comparative analysis of resampling techniques for class imbalance in financial distress prediction using XGBoost. Mathematics, 13(13), Article 2186. https://doi.org/10.3390/math13132186

Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., & Liu, T.-Y. (2017). LightGBM: A highly efficient gradient boosting decision tree. In Advances in Neural Information Processing Systems (Vol. 30, pp. 3146–3154).

Ko, P.-C., Lin, P.-C., Do, H.-T., & Huang, Y.-F. (2022). P2P lending default prediction based on AI and statistical models. Entropy, 24(6), Article 801. https://doi.org/10.3390/e24060801

Kunwar, S., & Ghimire, B. R. (2026). Leakage-free early prediction of issue resolution time in agile software projects: A comparative study on dataset quality. American Journal of Data Science and Artificial Intelligence, 2(1), 58–68. https://doi.org/10.54536/ajdsai.v2i1.7792

Lawal, O., Okolie, A., Obunadike, C., Akwabeng, P. M., Ikhifa, M. O., & Alumona, P. (2026). Predicting food insecurity across U.S. census tracts: A machine learning analysis using the USDA Food Access Research Atlas. American Journal of Data Science and Artificial Intelligence, 2(1), 1–14. https://doi.org/10.54536/ajdsai.v2i1.6285

Markov, A., Seleznyova, Z., & Lapshin, V. (2022). Credit scoring methods: Latest trends and points to consider. The Journal of Finance and Data Science, 8, 180–201. https://doi.org/10.1016/j.jfds.2022.07.002

Miljkovic, T., & Wang, P. (2025). A dimension reduction assisted credit scoring method for big data with categorical features. Financial Innovation, 11(1), Article 29. https://doi.org/10.1186/s40854-024-00689-1

Nguyen, N., & Ngo, D. (2025). Comparative analysis of boosting algorithms for predicting personal default. Cogent Economics & Finance, 13(1), Article 2465971. https://doi.org/10.1080/23322039.2025.2465971

Ozkurt, C. (2024). Enhancing financial decision-making: Predictive modeling for personal loan eligibility with gradient boosting, XGBoost, and AdaBoost. Information Technology in Economics and Business, 1(1), 7–13. https://doi.org/10.69882/adba.iteb.2024072

Yang, C. (2024). Research on loan approval and credit risk based on the comparison of machine learning models. SHS Web of Conferences, 181, Article 02003. https://doi.org/10.1051/shsconf/202418102003

Ying, C., Shi, A., & Li, X. (2025). Hybrid boosted attention-based LightGBM framework for enhanced credit risk assessment in digital finance. Humanities and Social Sciences Communications, 12(1), Article 1036. https://doi.org/10.1057/s41599-025-05230-y

Zhu, M., Zhang, Y., Gong, Y., Xing, K., Yan, X., & Song, J. (2024). Ensemble methodology: Innovations in credit default prediction using LightGBM, XGBoost, and LocalEnsemble [Preprint]. arXiv. https://arxiv.org/abs/2402.17979

Zhu, X., Chu, Q., Song, X., Hu, P., & Peng, L. (2023). Explainable prediction of loan default based on machine learning models. Data Science and Management, 6(3), 123–133. https://doi.org/10.1016/j.dsm.2023.04.003

Downloads

Published

2026-09-04

How to Cite

Poudel, S., & Ghimire, B. R. (2026). XGBoost versus LightGBM for Loan Approval Prediction: A Controlled, Like-for-Like Comparison. American Journal of Data Science and Artificial Intelligence, 2(2), 20-29. https://doi.org/10.54536/ajdsai.v2i2.8054

Similar Articles

You may also start an advanced similarity search for this article.