An ML Ensemble Model for Evaluating Student Performance and Recommending Academic Specializations in Science or Literature

Document Type : Original Article

Authors
1 Department of Artificial Intelligence and Data Science, College of Science and Technology, University of Human Development, Sulaymaniyah, Kurdistan Region, Iraq.
2 Department of Computer Science, College of Science, University of Halabja, Halabja, Kurdistan Region, Iraq.
3 Department of Software and Informatics Engineering, University of Raparin, Ranya, Kurdistan Region, Iraq.
10.24271/psr.2025.545273.2347
Abstract
Evaluating and predicting student performance is an essential part of modern education systems, helping to assess teaching approaches and improve student outcomes. Students’ academic performance analysis at the final stage of primary school is one of the most significant challenges in Educational Data Mining (EDM). Students often face various difficulties when deciding on specializations in the science or literature fields upon completing primary school and starting secondary school. With the increasing availability of student data, Machine Learning (ML) has emerged as a valuable tool for predicting academic performance, identifying at-risk students, and designed to enhancing organizational efficiency and overall student success. This study proposed an ensemble model that combines three ML algorithms Random Forest (RF), K-nearest Neighbors (KNN), and ADABOOST. Proposed ensemble mode prediction integrated through a voting mechanism. The significance of ensemble learning models lies in their ability to overcome the limitations of individual models while amplifying their strengths. Collected data from the E-Parwarda contains information from 26 schools in the Halabja province. During the preprocessing steps, authors utilized the K-Means algorithm to generate new class labels based on the relevant subjects for each department, thereby enhancing labelling accuracy by reflecting students’ academic strengths. After obtaining the findings, evaluated the proposed ensemble model against commonly used algorithms such as KNN, RF, ADABOOST, Decision Tree (DT), and Neural Network (NN). Proposed ensemble model's accuracy is 97.24% for random hold-out and cross-validation techniques, which is higher than the accuracies of the compared models.
Keywords
Crossmark
Subjects