Sucipto, Didik Dwi Prasetya, Triyanna Widiyaningtyas
The dataset of questions in the Cognitive Bloom's Taxonomy (BT) classification case uses a lot of private datasets. The lack of public datasets leads to a lack of performance comparison in the classification model. This study aims to determine the impact of the dataset performance model on the performance of the classification model. The classification model uses a classification algorithm model with optimization in feature extraction and feature selection on cognitive BT, including Naïve Bayes (NB), Support Vector Machine (SVM), and Random Forest (RF) on three public datasets. Three groups of datasets were used: small, medium, and large. The evaluation technique uses precision, accuracy, recall, specificity, and f1-score metrics and tests performance stability with k-fold cross-validation. The result of the study is that the overall classifier's performance depends on The degree of accuracy with which the dataset represents the original distribution relative to its size. The accuracy can be improved by adding extraction and selection features. The statistically tested datasets have significant differences between the models. The Tukey HSD test shows that the differences between the models have significant values. © 2024 IEEE.
Universitas Negeri Malang, Department of Electrical Engineering and Informatics, Malang, Indonesia