Aisyah Larasati, Rudi Nurdiansyah, Vertic E.B. Darmawan, Nikmatus Sholikha, Yuh-Wen Chen, Agus Rachmad Purnama
This study aims to analyze the effect of stemming methods, which are Nazief-Adriani and Porter, on the performance of the SVM Model for sentiment analysis. This study retrieves data by scraping customer reviews on Google Play and takes 10,000 data samples to do the experiment. There are 144 experimental combinations. This study uses the classical method as its stop-word removal method. The performance of the classifier model is measured based on the accuracy, computation time, and Receiver Operating Characteristics (ROC) curve. The SVM model parameters to be set are C, gamma, and kernel type. This study performs a grid search optimization to find the best parameter of the SVM model. This study finds that SVM with RBF kernel function, Porter stemming, and classical removal method results in the best accuracy and the highest area under the ROC curve. This study also finds that a 50%:50% data distribution of the class label (a balanced data set) does not always result in a better performance compared to a 70%:30 data distribution (an imbalanced data set). © 2025 IEEE.
Universitas Negeri Malang, Dept. of Mechanical and Industrial Engineering, Malang, Indonesia; Da-Yeh University, Dept. of Accounting and Information Management, Dacun, Taiwan; Universitas Nahdlatul Ulama Sidoarjo, Dept. of Industrial Engineering, Sidoarjo, Indonesia