K-Means Algorithm in Illiteracy Clustering

Closed

Triyanna Widiyaningtyas, Anwar Ludfianto, Denny Widhiyanuriyawan, Muhammad Anandha Fritama, Wahyu Styo Pratama

2023 ICEEIE 2023 - International Conference on Electrical, Electronics and Information Engineering Conference paper Cited by 1 Quartile

Abstract

Illiteracy is a condition where a person cannot read and write to communicate and express opinions in social life. Illiteracy is a problem that exists in almost all countries, including Indonesia. Even though the illiteracy rate continues to decrease annually, several regions still have a high quantity of illiteracy. Additionally, there are fewer illiterates in any area. One way to find out the distribution of illiterate population data is to map the population based on the level of illiteracy so that it can be prioritized for handling in areas with high illiteracy rates. In data mining, the technique for mapping the illiterate population can be done using the clustering method, namely by grouping the population of each region by giving labels based on their illiterate class. This study aims to cluster illiterate population data based on their illiteracy level and find the optimal number of classes in the illiteracy level clustering process. The class clustering algorithm uses k-means clustering. This study applies three scenarios with different k values, namely k=2, k=3, and k=5. The clustering results are validated using the cluster validation process to find the optimal number of clusters with a silhouette coefficient (SC). The results showed that the SC values at k=2, k=3, and k=5 were 0.24, 0.59, and 0.35. The results of this validation show that when k=3 produces the most excellent SC value, the study concludes that k=3 is the most ideal number of clusters. © 2023 IEEE.

Affiliations

Universitas Negeri Malang, Department of Electrical Engineering and Informatics, Malang, Indonesia; Universitas Brawijaya, Mechanical Engineering Department, Malang, Indonesia