Rahmawati Febrifyaning Tias, Triyanna Widiyaningtyas, Wahyu Sakti Gunawan Irianto, Mas Nurul Hamidah, Salahudin Robo
Recommendation systems are essential in guiding users to relevant items in large-scale platforms. One of the simplest yet widely adopted approaches is K-Nearest Neighbor (KNN), in both user-based and item-based forms. Although considered as a baseline, the performance of KNN remains highly dependent on similarity measures and data conditions such as scarcity and cold-start. This study presents a comprehensive evaluation with the Collaborative Filtering approach on the MovieLens-100K dataset, using three similarity functions namely Cosine, Pearson Correlation Coefficient (PCC), and Jaccard, and two evaluation scenarios, namely Random Split (5-Fold Cross Validation), Cold-Start (User-Item) The results show that in the Random Split scenario, both User-KNN and Item-KNN produce competitive performance on Cosine and Jaccard with MAE ranging from 0.82-0.84 while RMSE ranges from 1.04-1.07. In contrast, PCC produces higher errors, especially on User-KNN. Under cold-start conditions, Item-KNN is superior in handling new users and User-KNN shows better accuracy for new items. And in terms of computational efficiency, Item-KNN proved to be computationally faster. These findings confirm the limitations of traditional baselines and emphasize the importance of scenario-dependent evaluation. Furthermore, this study shows that a hybrid strategy combining User-KNN and Item-KNN can provide a more robust solution for recommendation systems. Overall, this study contributes to a systematic baseline analysis that not only measures accuracy but also takes into account computational efficiency, while offering further development of recommendation research. © 2025 IEEE.
Malang State University, Department of Electronics and Informatics Engineering, Malang, Indonesia; University of Bhayangkara, Informatic Engineering, Surabaya, Indonesia; Universitas Yapis Papua, Department of Information Systems, Papua, Indonesia