Claim Missing Document
Check
Articles

Found 25 Documents
Search

Analisis Sentimen Ulasan Mobile JKN pada Playstore dengan Perbandingan Akurasi Algoritma Naïve Bayes dan SVM Pranata, Eka Arya; Budiman, Fikri; Kurniawan, Defri
Building of Informatics, Technology and Science (BITS) Vol 7 No 1 (2025): June (2025)
Publisher : Forum Kerjasama Pendidikan Tinggi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bits.v7i1.7334

Abstract

The facilities provided by BPJS Health by releasing the Mobile JKN application, with this application the administrative process that previously had to be done directly can be done online and more flexibly. This research aims to see the sentiment of the community towards the JKN Mobile application review by comparing the SVM and Naïve Bayes algorithms. As well as optimizing the Naïve Bayes algorithm by using grid search. Reviews are taken from Google play with the help of Google Play Scraper API, the dataset taken amounted to 7,000 reviews. The results of using Naïve Bayes with an accuracy value of 86%, after tuning optimization using Grid Search significantly increases the accuracy value of the Naïve Bayes algorithm to 91% and for the SVM algorithm has an accuracy value of 92%. From the trial, it was found that the SVM algorithm is still better than the Naïve Bayes algorithm even though it has been optimized, but by optimizing the accuracy value Naïve Bayes is closer to SVM performance. This research can provide insight into the comparison of the two algorithms in identifying JKN Mobile reviews and the need for optimization to improve the performance of algorithms in sentiment analysis, besides that this research also contributes to the improvement and development of the JKN Mobile application so that it is useful for the community.
Analisis Sentimen Pengguna X terhadap Kasus Korupsi Gula Tom Lembong Menggunakan Naïve Bayes, SVM, dan Random Forest Kuncoro, Aneira Vicentiya; Budiman, Fikri; Kurniawan, Defri
Building of Informatics, Technology and Science (BITS) Vol 7 No 3 (2025): December 2025
Publisher : Forum Kerjasama Pendidikan Tinggi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bits.v7i3.8577

Abstract

The alleged sugar import corruption case involving Tom Lembong has become one of the most widely discussed public issues on social media, generating diverse reactions. This phenomenon illustrates how public opinion on legal issues is often influenced by perceptions of the public figures involved. This study aims to analyze public sentiment regarding the case on the social media platform X (formerly Twitter). The dataset consists of 1,802 tweets collected through a crawling process using the X API with the keyword “Tom Lembong.” The research stages include data cleaning, case folding, text normalization, tokenizing, stopword removal, stemming, sentiment labeling using a lexicon-based approach, and feature extraction with the Term Frequency–Inverse Document Frequency (TF-IDF) method. The prepared dataset was then tested using three classification algorithms: Naïve Bayes, Support Vector Machine (SVM), and Random Forest. The results show that the SVM algorithm achieved the highest accuracy (84%), followed by Random Forest (80%) and Naïve Bayes (76%). Based on the sentiment labeling results, positive sentiment dominated with 61%, while negative sentiment accounted for 39%. Although the analyzed issue concerns an alleged corruption case, the dominance of positive sentiment indicates that public opinion tends to focus on Tom Lembong’s personal image or public track record, which is viewed positively rather than on the substance of the legal allegations. These findings demonstrate the effectiveness of the SVM algorithm in analyzing high-dimensional text and provide insights into how public perception of legal issues can be influenced by image factors and the socio-political context on social media.
Optimizing Feature Extraction for Naïve Bayes Sentiment Analysis Achmad, Achmad; Budiman, Fikri
Journal of Applied Informatics and Computing Vol. 10 No. 1 (2026): February 2026
Publisher : Politeknik Negeri Batam

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30871/jaic.v10i1.12041

Abstract

The rapid growth of e-commerce platforms such as Tokopedia has generated a large volume of user reviews containing diverse opinions about products and services. These reviews reflect consumer perceptions and provide valuable insights for business decision-making. This study aims to enhance sentiment analysis performance by optimizing the Naïve Bayes algorithm through a comparison of two feature extraction techniques, namely Bag of Words (BoW) and Term Frequency–Inverse Document Frequency (TF-IDF). The dataset consists of 5,400 Tokopedia product reviews obtained from the Kaggle platform, which are categorized into positive and negative sentiments. The research process includes text preprocessing consisting of text cleaning, case folding, tokenization, stopword removal, and stemming, feature extraction using Bag of Words (BoW) and Term Frequency–Inverse Document Frequency (TF-IDF), handling data imbalance using the Synthetic Minority Over-sampling Technique (SMOTE), and model training using the Naïve Bayes. The dataset is divided into 80% training data and 20% testing data, and model performance is evaluated using accuracy, precision, recall, and F1-score. The results show that BoW achieved the highest accuracy of 93%, while TF-IDF reached 83%, indicating that BoW provides more effective feature representation and more stable performance for Naïve Bayes-based sentiment analysis on this dataset.
Comparison of Random Forest and LSTM for Tokopedia Sentiment Analysis Saputra, Fahrizal Denta; Budiman, Fikri
Journal of Applied Informatics and Computing Vol. 10 No. 1 (2026): February 2026
Publisher : Politeknik Negeri Batam

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30871/jaic.v10i1.12042

Abstract

Tokopedia is one of the largest e-commerce platforms in Indonesia, where every transaction generates user reviews containing opinions about the products or services received. These reviews provide important information about product quality, but the very large quantity makes manual analysis inefficient. This study aims to automatically classify Tokopedia review sentiment and compare the performance of machine learning and deep learning methods. The dataset used was obtained from Kaggle and has undergone an initial cleaning stage, including removing irrelevant columns and manually labeling into two sentiment classes, positive and negative. The research methodology includes several stages, namely data preprocessing (cleaning, case-folding, stopword removal, tokenization, normalization, and stemming), feature extraction using TF-IDF for Random Forest and word embedding for LSTM, implementation of Random Forest and Long Short-Term Memory (LSTM) models, and model evaluation using confusion matrix. Experimental results show that LSTM provides the best performance with 94% accuracy, while Random Forest achieves 92% accuracy. These findings indicate that LSTM is more effective in understanding language context, resulting in more accurate sentiment classification and is useful for decision making in the e-commerce field.
Analisis Sentimen Ulasan Film pada IMDb Menggunakan Support Vector Machine dengan Penerapan Chi-Square dan Word Embedding Janto, Dwi; Marjuni, Aris; Budiman, Fikri
Jurnal Teknologi Informasi dan Ilmu Komputer Vol 13 No 3: Juni 2026
Publisher : Fakultas Ilmu Komputer, Universitas Brawijaya

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.25126/jtiik.2026133

Abstract

Analisis sentimen adalah metode untuk mendeteksi sentimen pada opini dalam teks. Tujuan dari penelitian ini adalah untuk menggabungkan pemilihan fitur chi-square dan teknik word embedding untuk meningkatkan akurasi analisis sentimen pada 50.000 data ulasan film dari IMDb. Chi-square digunakan untuk memilih fitur yang paling relevan dalam data teks, sementara word embedding seperti Word2Vec, FastText, dan GloVe digunakan untuk merepresentasikan kata dalam bentuk vektor numerik.Proses analisis diawali dengan pre-processing data, dilanjutkan dengan pemilihan fitur menggunakan uji chi-square, representasi fitur menggunakan word embedding, dan klasifikasi menggunakan Support Vector Machine (SVM). Hasil penelitian menunjukkan bahwa menggabungkan metode chi-square dengan word embedding meningkatkan kinerja model SVM dibandingkan tanpa feature selection dan word embedding. Hasil kinerja terbaik diperoleh dengan menggunakan gabungan SVM, Word2Vec, dan chi-square, dengan akurasi sebesar 88,63%, precision sebesar 87,64%, recall sebesar 89,61%, dan F1-score sebesar 88,61%. Penelitian ini juga menunjukkan bahwa pemilihan fitur chi-square dapat secara efektif mengurangi dimensionalitas data tanpa mengurangi kualitas informasi, dan word embedding meningkatkan akurasi dengan meningkatkan representasi kata-kata. Hasil penelitian ini menegaskan bahwa kombinasi metode-metode ini dapat digunakan secara efektif untuk analisis sentimen, khususnya pada kumpulan data besar seperti IMDb.   Abstract This study discusses sentiment analysis of movie reviews taken from the public IMDb dataset. The main objective of this research is to enhance the performance of the classification model in detecting positive and negative sentiments by utilizing the Chi-Square feature selection method and word embedding techniques. The classification model used is the Support Vector Machine (SVM). Three word embedding techniques tested in this study include Word2Vec, FastText, and GloVe. The study also examines the effectiveness of Chi-Square as a feature selection method to improve model accuracy. The experimental results show that the combination of SVM with Word2Vec supported by Chi-Square feature selection provides the best performance with an accuracy of 88.63%, precision of 87.64%, recall of 89.61%, and an F1-Score of 88.61%. Conversely, the use of GloVe resulted in the lowest performance with an accuracy of 71.92% before feature selection. After Chi-Square feature selection, the performance improved to 79.87%. This study reinforces the conclusion that the use of word embedding techniques together with the Chi-Square feature selection method can significantly enhance the performance of the SVM model in sentiment analysis. This research contributes to developing a more effective approach for sentiment analysis using SVM-based classification methods, by combining feature selection and word embedding as a performance improvement strategy.