Understanding students’ mathematical achievement through feature selection and machine learning: Evidence from PISA 2018 Comprendiendo el rendimiento en matemáticas de los estudiantes a través de selección de variables y aprendizaje automático: evidencia de PISA 2018


Yildirim U. H., KAYHAN ATILGAN Y., Turfan D.

Revista de Psicodidactica, cilt.31, sa.2, 2026 (SSCI, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 31 Sayı: 2
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1016/j.psicod.2026.500187
  • Dergi Adı: Revista de Psicodidactica
  • Derginin Tarandığı İndeksler: Social Sciences Citation Index (SSCI), Scopus, Psicodoc, DIALNET
  • Anahtar Kelimeler: Classification, Educational data mining, Explainable artificial intelligence, Feature selection, Machine learning, Mathematical achievement
  • Hacettepe Üniversitesi Adresli: Evet

Özet

The Programme for International Student Assessment (PISA) provides a framework for examining academic achievement. Using the Turkish sample of PISA 2018, this study investigates mathematics achievement, focusing on students performing below and above the baseline proficiency level (Level 2). To address this objective, a machine learning classification approach is adopted, integrating feature selection, supervised algorithms, and model interpretation techniques. The analysis includes 6,863 students from 186 schools and 62 student- and school-level variables derived from the PISA questionnaires. Three feature selection methods – Boruta, Mutual Information, and ReliefF – are used to identify subsets of variables that serve as inputs for five classification models: Decision Tree, Bagging, Random Forest, AdaBoost, and Support Vector Machine. Model performance is evaluated using repeated stratified cross-validation across plausible values and multiple classification metrics. The results indicate that models built on substantially reduced feature sets can achieve predictive performance comparable to models using all variables. Among the evaluated configurations, the combination of Mutual Information feature selection with the top 20 variable subset and the Random Forest classifier provides a stable and practical useful balance between prediction performance and model parsimony. Model predictions are examined using conditional SHAP values, partial dependence plots, individual conditional expectation plots, and accumulated local effects plots. Results show that socio-economic background, perceived test difficulty, occupational expectations, school stratum type and behaviors hindering learning are among the variables most strongly associated with model predictions. These analyses capture predominantly non-linear associations between predictors and model predictions, and these relationships vary across students, indicating heterogeneous model responses.