Muhamad Fahrurozi
STMIK IKMI CIREBON

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

NAÏVE BAYES SENTIMENT ACCURACY WITH CHI-SQUARE AND INFORMATION GAIN Muhamad Fahrurozi; Dian Ade Kurnia; Yudhistira Arie Wijaya; Puji Pramudya Marta; Khaerul Anam
Antivirus : Jurnal Ilmiah Teknik Informatika Vol 20 No 1 (2026): Mei 2026
Publisher : Universitas Islam Balitar

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.35457/1nmksc45

Abstract

This study aims to evaluate the effectiveness of the Chi-Square and Information Gain feature selection techniques in improving the accuracy of sentiment classification on application reviews using the Naive Bayes algorithm. The main issue addressed in this research is the high dimensionality of features in Indonesian-language text data, which may potentially affect model performance. To address this, the study applies a text preprocessing pipeline consisting of sentiment labeling, normalization, and feature extraction using TF-IDF with an initial 5000 features. The dataset contains 225,043 Gojek application reviews, which were reduced to 220,860 valid entries after cleaning, and subsequently divided into training and testing sets using a stratified split. Experimental results show that Chi-Square achieved the highest accuracy of 0.877864 with 3000 features, while Information Gain reached an accuracy of 0.877751 with 4000 features. Both values are slightly lower than the model without feature selection, which achieved an accuracy of 0.878045. These findings indicate that feature reduction does not improve the performance of Naive Bayes, as the algorithm performs more effectively when retaining a broader distribution of words and more complete contextual representations. In conclusion, feature selection for probabilistic models should be applied cautiously, especially on Indonesian-language review data that are informal and exhibit high variation in expression.