An Evaluation of the Impact of Dataset Size on Classification Performance in the Cognitive Bloom's Taxonomy

Closed

Sucipto, Didik Dwi Prasetya, Triyanna Widiyaningtyas

2024 2024 Beyond Technology Summit on Informatics International Conference, BTS-I2C 2024 Conference paper Cited by 0 Quartile

Abstract

The dataset of questions in the Cognitive Bloom's Taxonomy (BT) classification case uses a lot of private datasets. The lack of public datasets leads to a lack of performance comparison in the classification model. This study aims to determine the impact of the dataset performance model on the performance of the classification model. The classification model uses a classification algorithm model with optimization in feature extraction and feature selection on cognitive BT, including Naïve Bayes (NB), Support Vector Machine (SVM), and Random Forest (RF) on three public datasets. Three groups of datasets were used: small, medium, and large. The evaluation technique uses precision, accuracy, recall, specificity, and f1-score metrics and tests performance stability with k-fold cross-validation. The result of the study is that the overall classifier's performance depends on The degree of accuracy with which the dataset represents the original distribution relative to its size. The accuracy can be improved by adding extraction and selection features. The statistically tested datasets have significant differences between the models. The Tukey HSD test shows that the differences between the models have significant values. © 2024 IEEE.

Affiliations

Universitas Negeri Malang, Department of Electrical Engineering and Informatics, Malang, Indonesia