Muhammad Fatkhur Rizal, Triyanna Widiyaningtyas, Hary Suswanto, Wiwik Suharso, Salahudin Robo, Rahmawati Febrifyaning Tias, Erfan Ainul Yakin
The robustness of lexicon-based sentiment analysis for rating prediction remains insufficiently examined, particularly under realistic cross-domain and temporally ordered conditions. This study evaluates three lexicon methods TextBlob, VADER, and AFINN combined with threshold and calibrated mapping strategies for converting polarity scores into 1-5-star ratings. Experiments were conducted on 50,000 Amazon Books reviews using leakage-safe temporal splits and evaluated with exact accuracy, within one accuracy, and RMSE. Results show that threshold-only mapping yields strong ordinal proximity but weak exact accuracy, with TextBlob achieving 85.87% within-one accuracy yet only 32.70% exact match. Calibration substantially improves performance: calibrated TextBlob attains 62.70% exact accuracy, surpassing the majority baseline (62.56%) while reducing RMSE relative to threshold mapping. VADER exhibits similar improvements after calibration. These findings demonstrate that lexicon-based models remain practical, interpretable, and computationally efficient for rating prediction, but only when paired with appropriate calibration procedures. The study provides explicit task definitions, reproducible evaluation protocols, and insights relevant to interpretable sentiment components in recommendation systems. © 2025 IEEE.
Universitas Negeri Malang, Department of Electrical Engineering and Informatics, Malang, Indonesia; Universitas Muhammadiyah Jember, Department of Information Systems, Jember, Indonesia; Universitas Yapis Papua, Department of Information Systems, Jayapura, Indonesia; Universitas Bhayangkara Surabaya, Department of Informatics, Malang, Indonesia