A Comparison of Classical and Machine Learning Techniques for Binary Logistic Regression Estimation

Document Type : Original Article

Authors
1 Department of Mathematics, College of Basic Education, Salahaddin University, Erbil, Kurdistan Region, Iraq.
2 Department of Mathematics, College of Science, Salahaddin University, Erbil, Kurdistan Region, Iraq.
10.24271/psr.2025.530530.2236
Abstract
Binary Logistic Regression (BLR) is widely used for classification, but parameter estimation can be unstable due to reliance on numerical optimization, class separation, and high-dimensional data. This research aims to address these challenges by evaluating and comparing four distinct approaches to estimating BLR parameters: the classical Maximum Likelihood Estimation (MLE) based on the Newton-Raphson Method (NRM), regularized approaches via Gradient Descent (GD) with Lasso Regression (LR, L1) and Ridge Regression (RR, L2) penalties, and the metaheuristic Particle Swarm Optimization (PSO). To evaluate these methods, real-world clinical data from 1,511 breast cancer patients treated at Hiwa Hospital in Sulaymaniyah between 2019 and 2023 were utilized. The main objective was to evaluate each method’s predictive accuracy in identifying the need for chemotherapy, using evaluation metrics such as Area Under the Curve (AUC), Log-Loss, and Mean Squared Error (MSE). Using N-split Cross-Validation, advanced statistical tests were applied to confirm the significance and reliability of the observed differences, including the Friedman test, the Wilcoxon test, and the Holm correction procedure. Therefore, this study's primary goal is to determine the most reliable and effective technique for estimating BLR parameters. Consequently, statistical models in real-world applications are more accurate. According to the results, MLE and PSO performed noticeably better, obtaining the highest AUC values and the lowest levels of loss (Log-Loss and MSE). Their performance was consistent across all N-split values, with no significant differences between them in most statistical tests. LR performed mediocrely, falling short of MLE and PSO, especially in the AUC metric, though the differences were not always statistically significant. However, RR did not perform well. Its Log-Loss and MSE values were extremely high and markedly different from the other methods, even though its AUC was similar to the others. This illustrates the model's instability and poor probability estimation, likely caused by over-regularization.
Keywords
Crossmark
Subjects