The Effect of Word Removal Method on The Performance of SVM and Naïve Bayes for Sentiment Analysis

Closed

Aisyah Larasati, Rudi Nurdiansyah, Abdul Muid, Rafida Salwa Nida, Ilham Firmansyah, Agus Rachmad Purnama

2025 2025 9th International Conference on Electrical, Electronics and Information Engineering, ICEEIE 2025 Conference paper Cited by 0 Quartile

Abstract

This study aims to analyze the effect of word removal methods (Classical and Zipf's) on the performance of SVM and Naïve Bayes for a sentiment analysis. The source of data is customer reviews on a certain online platform at the Google Play. This study takes 10,000 data samples to experiment. Data is split into training and testing data sets with a ratio of 7 5% : 2 5%, in which the training data set in the form of a balanced data set (having a positive and a negative class with a ratio of 50:50). The stemming algorithm in the text preprocessing is the Porter algorithm. The performance of the classifier model is measured based on the accuracy, computation time, and the Receiver Operating Characteristics (ROC) curve. The SVM model parameters to be set are C, gamma, and kernel type. The Naïve Bayes model parameter is the alpha value. This study finds that the best model to perform sentiment analysis regarding model accuracy and ROC curve is the SVM model using the classical word removal. However, the best computation time is achieved by Naïve Bayes and Zipf's removal word method. Thus, the selection of word removal method and the modelling algorithm in a sentiment analysis should consider the trade-off between model accuracy, ROC curve, and computation time © 2025 IEEE.

Affiliations

Universitas Negeri Malang, Dept. of Mechanical and Industrial Engineering, Malang, Indonesia; Universitas Nahdlatul Ulama Sidoarjo, Dept. of Industrial Engineering, Sidoarjo, Indonesia