Enhancing Phishing Detection: A Comparative Analysis of Machine Learning Techniques

Document Type : Original Article

Authors
1 Department of Computer Science, Darbandikhan Technical Institute, Sulaimani Polytechnic University, Sulaimani, Iraq
2 Department of Computer Science, Darbandikhan Technical Institute, Sulaimani Polytechnic University, Sulaimaniya, Iraq
3 Computer Science Department, Kurdistan Technical Institute, Sulaimani, Iraq
4 Department of Computer Science, Cihan University Sulaimaniya, Sulaymaniyah City 46001, Iraq
10.24271/psr.2025.464499.1640
Abstract
Phishing attacks trick people into giving away personal information on fake websites. These attacks are still a big problem for cybersecurity because traditional methods often can't keep up with new tricks. This study evaluated various supervised machine learning (ML) classifiers to identify phishing websites and examined how feature selection can enhance their performance. Seven ML classifiers were tested: Naive Bayes, Multilayer Perceptron (MLP), AdaBoost, Bagging, Support Vector Machine (SVM), JRip, and REP Tree. These classifiers were trained on a dataset consisting of 2,500 instances, comprising 1,220 legitimate websites and 1,280 phishing sites, using 10-fold cross-validation. Feature selection techniques, specifically Principal Component Analysis (PCA) and Incremental Wrapper Subset Selection (IWSS), were employed to improve the input features. The results showed that PCA significantly improved detection rates through dimensionality reduction while retaining key discriminative features, achieving an accuracy of 99.8%. Among the classifiers, MLP outperformed the others with an accuracy of 92.94%, and when PCA was applied to MLP, the accuracy increased to 97.5%. These findings highlight the importance of feature optimization in identifying phishing websites and demonstrate that MLP is a suitable model for real-world applications.
Keywords
Crossmark
Subjects