p-Index From 2021 - 2026
7.518
P-Index
This Author published in this journals
All Journal Teknika Jurnal Sains dan Teknologi Jurnal Simetris JSI: Jurnal Sistem Informasi (E-Journal) International Journal of Advances in Intelligent Informatics IJCIT (Indonesian Journal on Computer and Information Technology) Jurnal Pilar Nusa Mandiri SINTECH (Science and Information Technology) Journal Jurnal Informatika Universitas Pamulang Jurnal Nasional Komputasi dan Teknologi Informasi WIDYA LAKSANA Jurnal Informatika Kaputama (JIK) EVOLUSI : Jurnal Sains dan Manajemen JTIK (Jurnal Teknik Informatika Kaputama) Jurnal Ilmu Teknik dan Komputer Jurnal Tekinkom (Teknik Informasi dan Komputer) Infotek : Jurnal Informatika dan Teknologi Jurnal Media Informatika JUSTIAN - Jurnal Sistem Informasi Akuntansi J-Intech (Journal of Information and Technology) Ilmu Komputer untuk Masyarakat Jurnal Pengabdian Masyarakat Bidang Sains dan Teknologi Jurnal Ilmu Komputer Dan Informatika Journal of Artificial Intelligence and Engineering Applications (JAIEA) Jurnal Komputer Teknologi Informasi Sistem Komputer (JUKTISI) Jurnal Abdimas Le Mujtamak Mestaka: Jurnal Pengabdian Kepada Masyarakat TAMIKA: Jurnal Tugas Akhir Manajemen Informatika & Komputerisasi Akuntansi Dedikasi Saintek Jurnal Pengabdian Masyarakat SOROT: Jurnal Pengabdian Kepada Masyarakat Madani: Jurnal Pengabdian Masyarakat dan Kewirausahaan Jurnal Komputer dan Teknologi (JUKOMTEK) DEDIKASI SAINTEK Jurnal Pengabdian Masyarakat Indonesian Community Service Journal of Computer Science (IndoComs) Jurnal Nasional Komputasi dan Teknologi Informasi
Claim Missing Document
Check
Articles

Sistem Prediksi Risiko Penyakit Jantung Berbasis Machine Learning dan Framework Streamlit Reymond Syahputra Hidayana; Fransiska Regina; Rendi Rendi; Riski Annisa

Publisher :

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.32672/jnkti.v8i6.10158

Abstract

Abstrak - Penelitian ini menggunakan algoritma pembelajaran mesin untuk membangun sistem yang dapat memprediksi risiko penyakit jantung. Dalam dataset Cleveland Heart Disease, tiga algoritma Logistic Regression, XGBoost, dan Naive Bayes digunakan dengan pembagian data uji dan latih sebesar 80:20. Pembersihan data, pemisahan fitur dan target, pelatihan model, dan evaluasi menggunakan metrik akurasi, presisi, recall, f1-score, dan AUC dilakukan. Hasil pengujian menunjukkan bahwa Logistic Regression adalah yang terbaik dengan skor akurasi, presisi, recall, dan f1-score sebesar 0,90, dan AUC sebesar 0,94. Selanjutnya, model terbaik diterapkan pada sistem prediksi berbasis web yang menggunakan framework Streamlit. Selain data pengguna, sistem dapat menampilkan risiko penyakit jantung secara informatif. Berdasarkan hasil penelitian, model Logistic Regression dapat digunakan sebagai alat bantu awal dalam mendeteksi risiko penyakit jantung secara efektif.Kata kunci : Prediksi Penyakit Jantung; Machine Learning; Logistic Regression; Klasifikasi; Streamlit; Abstract - This study employs machine learning algorithms to develop a system capable of predicting the risk of heart disease. Using the Cleveland Heart Disease dataset, three algorithms—Logistic Regression, XGBoost, and Naive Bayes—were applied with an 80:20 train-test split. Data cleaning, feature–target separation, model training, and evaluation using accuracy, precision, recall, f1-score, and AUC metrics were conducted. The results indicate that Logistic Regression performs the best, achieving accuracy, precision, recall, and f1-score values of 0.90, and an AUC of 0.94. The best-performing model was then deployed in a web-based prediction system using the Streamlit framework. In addition to user input, the system provides an informative display of heart disease risk. Based on the findings, the Logistic Regression model can serve as an effective preliminary tool for detecting heart disease risk.Keywords: Heart Disease Prediction; Machine Learning; Logistic Regression; Classification; Streamlit;
Performance Evaluation of the BERT Model in Sentiment Analysis of DANA Application User Reviews Hazael Susanto; Weiskhy Steven Dharmawan; Riski Annisa; Lady Agustin Fitriana
Journal of Artificial Intelligence and Engineering Applications (JAIEA) Vol. 5 No. 3 (2026): June 2026
Publisher : Yayasan Kita Menulis

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.59934/jaiea.v5i3.2359

Abstract

The rapid growth of digital wallets in Indonesia generates a large volume of user reviews on platforms such as the Google Play Store that cannot be efficiently analyzed manually. This study aims to evaluate the performance of the BERT (Bidirectional Encoder Representations from Transformers) model in sentiment classification tasks on a dataset of DANA application user reviews collected from the Google Play Store. The BERT model is fine-tuned using labeled Indonesian-language data with three sentiment classes: positive, negative, and neutral. Specialized preprocessing strategies are applied to handle the characteristics of informal text, abbreviations, and code-switching phenomena prevalent in Indonesian user reviews. Evaluation is conducted using accuracy, precision, recall, and F1-score metrics. Experimental results indicate that the fine-tuned IndoBERT model achieves an accuracy of 91.24% with a weighted F1-score of 0.91 on a test dataset of 6,106 samples. The Negative class achieves the highest performance with an F1-score of 0.95, followed by the Positive class (0.88) and Neutral class (0.84). This study provides empirical evidence of the effectiveness of the IndoBERT Transformer architecture for sentiment analysis in the Indonesian-language fintech domain and can serve as a reference for developing deep learning-based NLP systems in similar contexts.
Performance Evaluation of Machine Learning Algorithms in Sentiment Analysis of Spotify Reviews Frizi Olivian; Sahrul Bariyah; Grant Christo Budiyanto; Riski Annisa; Lady Agustin Fitriana; Weiskhy Steven Dharmawan
Journal of Artificial Intelligence and Engineering Applications (JAIEA) Vol. 5 No. 3 (2026): June 2026
Publisher : Yayasan Kita Menulis

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.59934/jaiea.v5i3.2362

Abstract

The rapid growth of digital music streaming platforms has generated a massive volume of user reviews on the Google Play Store, making manual analysis practically infeasible. This study evaluates and compares the performance of three machine learning algorithms Support Vector Machine (SVM), Neural Network (Multilayer Perceptron), and Random Forest in classifying sentiments from Spotify user reviews written in Indonesian. A total of 10,000 reviews were collected from the Google Play Store using the google-play-scraper library and processed through a text preprocessing pipeline comprising cleaning, case folding, word normalization, tokenization, stopword removal, and stemming using the Sastrawi library. Sentiment labeling was performed automatically using the InSet lexicon, categorizing reviews into three classes: Positive (56.63%), Neutral (30.60%), and Negative (12.76%). Feature extraction was conducted using the TF-IDF method, with an 80:20 train-test split strategy and stratified sampling to maintain class distribution. Model performance was evaluated based on accuracy, precision, recall, and F1-score metrics. The results demonstrate that SVM and Neural Network achieved equivalent and superior accuracy of 0.937, with macro F1-scores of 0.908 and 0.907, respectively, outperforming Random Forest which recorded an accuracy of 0.853 and a macro F1-score of 0.777. These findings indicate that SVM and Neural Network are more optimal and reliable for sentiment classification of Indonesian-language Spotify reviews, while Random Forest requires further improvement, particularly in recognizing minority classes.
Topic Modeling of Clash of Clans Player Reviews Using NLP-Based Latent Dirichlet Allocation (LDA) Machine Learning Method Rai Markus Panamuan; Debi Handika; Muhamad Rizki Pratama; Weiskhy Steven Dharmawan; Lady Agustin Fitriana; Riski Annisa
Journal of Artificial Intelligence and Engineering Applications (JAIEA) Vol. 5 No. 3 (2026): June 2026
Publisher : Yayasan Kita Menulis

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.59934/jaiea.v5i3.2364

Abstract

The rapid growth of the mobile gaming industry has generated millions of player reviews on platforms like the Google Play Store. Clash of Clans, developed by Supercell, is one of the world's most popular mobile strategy games, generating a vast volume of user reviews that are difficult to analyze manually. This study applies Latent Dirichlet Allocation (LDA), a generative probabilistic machine learning model based on Natural Language Processing (NLP), to identify and cluster key topics discussed in player reviews on the Google Play Store. A total of 10,000 player reviews were collected through web scraping, followed by NLP-based text preprocessing including tokenization, stopword removal, and lemmatization. The LDA model was optimized using a coherence score evaluation of 0.512, resulting in the identification of five dominant discussion topics: technical issues and bugs, game updates and balance, gameplay and strategy, monetization and in-app purchases, and social interactions and clan systems. The results show that LDA-based topic modeling provides structured and actionable insights for game developers to understand player feedback and improve game quality. This research contributes to the field of NLP-based mobile game review analysis.
Analisis Sentimen Ulasan Aplikasi Detik.Com di Google Play Store Menggunakan Pendekatan Lexicon-Based dan Machine Learning Talcha Ilham Putri; Riski Annisa; Muhammad Fahmi Julianto
Jurnal Komputer Teknologi Informasi Sistem Komputer (JUKTISI) Vol. 5 No. 1 (2026): Juni 2026
Publisher : LKP KARYA PRIMA KURSUS

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.62712/juktisi.v5i1.1195

Abstract

Ulasan pengguna pada Google Play Store merupakan sumber informasi yang dapat digunakan untuk mengetahui tingkat kepuasan pengguna terhadap suatu aplikasi. Namun, jumlah ulasan yang terus bertambah menyebabkan proses analisis secara manual menjadi kurang efektif. Oleh karena itu, diperlukan metode analisis sentimen untuk mengidentifikasi kecenderungan opini pengguna secara otomatis. Penelitian ini bertujuan untuk menganalisis sentimen ulasan pengguna aplikasi Detik.com menggunakan pendekatan Lexicon-Based serta membandingkan kinerja algoritma Naive Bayes, Decision Tree, Random Forest, dan Support Vector Machine (SVM). Data penelitian diperoleh dari Google Play Store melalui proses web scraping. Tahapan penelitian meliputi preprocessing data, pelabelan sentimen menggunakan pendekatan Lexicon-Based, ekstraksi fitur menggunakan TF-IDF, pembagian data latih dan data uji, proses klasifikasi menggunakan algoritma machine learning, serta evaluasi model menggunakan metrik accuracy, precision, recall, dan F1-score. Hasil pelabelan sentimen menunjukkan bahwa dari 1.000 ulasan yang dianalisis, sebanyak 591 ulasan (59,10%) termasuk sentimen positif dan 409 ulasan (40,90%) termasuk sentimen negatif. Berdasarkan hasil pengujian, algoritma Decision Tree memperoleh performa terbaik dengan nilai akurasi sebesar 79,6%, precision sebesar 80,1%, recall sebesar 79,6%, dan F1-score sebesar 79,7%. Sementara itu, Random Forest memperoleh akurasi sebesar 77,6%, SVM sebesar 74,5%, dan Naive Bayes sebesar 71,9%. Hasil penelitian menunjukkan bahwa pendekatan Lexicon-Based yang dikombinasikan dengan algoritma machine learning mampu digunakan untuk menganalisis sentimen ulasan pengguna aplikasi Detik.com secara efektif, dengan Decision Tree sebagai algoritma yang memberikan kinerja terbaik pada dataset penelitian.
Analisis Sentimen Program Bantuan Sosial Menggunakan Metode Machine Learning Anna; Riski Annisa; Panny Agustia Rahayuningsih
Jurnal Nasional Komputasi dan Teknologi Informasi Vol. 9 No. 2 (2026): April, 2026
Publisher : Program Studi Teknik Komputer, Fakultas Teknik. Universitas Serambi Mekkah

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.32672/thpxny55

Abstract

Abstrak - Program bantuan sosial merupakan instrumen penting yang diimplementasikan oleh pemerintah untuk meningkatkan kesejahteraan masyarakat dan mengurangi kesenjangan ekonomi. Studi ini bertujuan untuk menganalisis sentimen publik terhadap program bantuan sosial di Indonesia dengan memanfaatkan metode machine learning. Jaringan media sosial menyediakan data teks berupa jawaban, pandangan, dan komentar publik. Data teks mentah diproses terlebih dahulu menggunakan pembersihan teks, case folding, tokenisasi, penghapusan stop word, dan stemming sebelum analisis sentimen dilakukan. Selain itu, dua annotator menggunakan aturan anotasi yang disediakan untuk mengkategorikan data teks secara manual dengan sentimen (positif, negatif, atau netral). Tiga model machine learning—Bidirectional Encoder Representations from Transformers (BERT), Long Short-Term Memory (LSTM), dan Regresi Logistik—digunakan untuk menilai sentimen. Kinerja model diuji menggunakan metrik presisi, recall, dan F1-score untuk menentukan akurasi dan efektivitasnya. Dengan F1-score 0,93, temuan menunjukkan bahwa model BERT berkinerja terbaik dalam analisis sentimen. Analisis sentimen mengungkapkan bahwa sentimen netral mendominasi tanggapan publik terhadap program bantuan sosial, yang menunjukkan bahwa publik belum memiliki opini yang kuat, baik positif maupun negatif, terhadap program bantuan sosial. Temuan ini memberikan informasi berharga bagi para pembuat kebijakan dan pelaksana program untuk mengevaluasi program bantuan sosial secara komprehensif, mengidentifikasi area yang perlu ditingkatkan, dan meningkatkan kualitas layanan untuk memaksimalkan manfaat bagi masyarakat. Kata kunci: Logistic Regression; Machine Learning; Analisis Sentimen; Program Bantuan Sosial; Twitter; Abstract - Social assistance programs are important instruments implemented by the government to improve public welfare and reduce economic disparities. This study aims to analyze public sentiment towards social assistance programs in Indonesia by utilizing machine learning methods. Social media networks provided text data in the form of public answers, views, and comments. Raw text data was preprocessed using text cleaning, case folding, tokenization, stop word removal, and stemming before sentiment analysis was performed. Additionally, two annotators used the supplied annotation rules to manually categorize the text data with sentiment (positive, negative, or neutral). Three machine learning models—Bidirectional Encoder Representations from Transformers (BERT), Long Short-Term Memory (LSTM), and Logistic Regression—were used to assess sentiment. Model performance was tested using precision, recall, and F1-score metrics to determine their accuracy and efficacy. With an F1-score of 0.93, the findings demonstrated that the BERT model performed the best in sentiment analysis. Sentiment analysis revealed that neutral sentiment dominates public responses to social assistance programs, indicating that the public does not yet have a strong opinion, positive or negative, towards social assistance programs. This finding provides valuable information for policy makers and program implementers to comprehensively evaluate social assistance programs, identify areas that need improvement, and improve service quality to maximize benefits to the community. Keywords: Logistic Regression; Machine Learning; Sentiment Analysis; Social Assistance Program; Twitter;
Analisis Sentimen Komentar Trailer Film Menggunakan Pendekatan Lexicon-Based dan Machine Learning Fariska Adela Nurhidayah; Riski Annisa; Muhammad Fahmi Julianto
Jurnal Nasional Komputasi dan Teknologi Informasi Vol. 9 No. 3 (2026): Juni, 2026
Publisher : Program Studi Teknik Komputer, Fakultas Teknik. Universitas Serambi Mekkah

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.32672/fjfa1e80

Abstract

Abstrak – Seiring berkembangnya media sosial, YouTube telah menjadi salah satu tempat utama di mana orang dapat menggunakan kolom komentar untuk berbagi pemikiran mereka tentang film. Penelitian ini menggunakan kombinasi teknik berbasis leksikon dan machine learning untuk memeriksa sentimen penonton mengenai trailer film Andai Ibu Tidak Menikah dengan Ayah. Sejumlah langkah preprocessing, termasuk cleaning, case folding, normalisasi, tokenisasi, stopword removal, dan stemming, diterapkan pada data setelah dikumpulkan melalui scraping komentar YouTube. Leksikon Sentimen Indonesia (InSet Lexicon) digunakan untuk pelabelan sentimen, dan pendekatan TF-IDF digunakan untuk ekstraksi fitur. Synthetic Minority Oversampling Technique (SMOTE) digunakan untuk mengoreksi ketidakseimbangan data. Metode Naïve Bayes, Logistic Regression, dan Support Vector Machine (SVM) kemudian digunakan untuk mengklasifikasikan sentimen. Metrik akurasi, presisi, recall, dan F1-score digunakan untuk menilai kinerja model. Dengan akurasi 81.90%, presisi 79.95%, recall 81.90%, dan F1-score 80.87%, hasil ini menunjukkan bahwa algoritma Naïve Bayes berkinerja terbaik. Sementara itu, akurasi SVM dan Logistic Regression masing-masing adalah 66.67% dan 62.86%. Hasil ini menunjukkan bahwa Naïve Bayes mengungguli algoritma lain dalam klasifikasi sentimen dari komentar trailer film. Kata kunci : Analisis Sentimen; Lexicon-Based; Machine Learning; TF-IDF; YouTube;   Abstract - As social media has grown, YouTube has become one of the main places where people may use comment sections to share their thoughts about movies. This study uses a combination of lexicon-based and machine learning techniques to examine viewer sentiment regarding the trailer for the film Andai Ibu Tidak Menikah dengan Ayah. A number of preprocessing steps, including as cleaning, case folding, normalization, tokenization, stopword removal, and stemming, were applied to the data after it was gathered via YouTube comment scraping. The Indonesian Sentiment Lexicon (InSet Lexicon) was used for sentiment labeling, and the TF-IDF approach was used for feature extraction. The Synthetic Minority Oversampling Technique (SMOTE) was used to correct data imbalance. The Naïve Bayes, Logistic Regression, and Support Vector Machine (SVM) methods were then used to classify sentiment. Accuracy, precision, recall, and F1-score metrics were used to assess the model's performance. With an accuracy of 81.90%, precision of 79.95%, recall of 81.90%, and F1-score of 80.87%, the findings show that the Naïve Bayes algorithm performed the best. In the meantime, the accuracy of SVM and Logistic Regression was 66.67% and 62.86%, respectively. These results show that Naïve Bayes outperforms the other algorithms in sentiment classification from movie trailer comments. Keywords: Sentiment Analysis; Lexicon-Based; Machine Learning; TF-IDF; YouTube;
Prediksi Produksi Kelapa Sawit Menggunakan Algoritma Machine Learning Berdasarkan Data Operasional 2025 Cici; Riski Annisa; Muhammad Fahmi Julianto
Jurnal Nasional Komputasi dan Teknologi Informasi Vol. 9 No. 3 (2026): Juni, 2026
Publisher : Program Studi Teknik Komputer, Fakultas Teknik. Universitas Serambi Mekkah

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.32672/h4zecc92

Abstract

Abstrak - Produksi kelapa sawit adalah salah satu parameter penting untuk melihat seberapa efisien operasi di perkebunan. Penelitian ini bertujuan untuk meramalkan produksi kelapa sawit (Janjang per Pokok) dengan menggunakan algoritma Machine Learning yang didasarkan pada data operasional tahun 2025. Dataset ini terdiri dari 1.347 catatan operasional yang mencakup 13 variabel fitur. Variabel-variabel ini termasuk luas lahan, produktivitas tenaga kerja, serta fitur turunan seperti Total_Hari kerja pemanen dan Hari Kerja per Hektar. Metode yang digunakan meliputi Regresi Linier, Regresi Pohon Keputusan, dan Regresi Hutan Acak dengan pembagian data sebesar 80% untuk latihan dan 20% untuk pengujian. Hasil evaluasi menunjukkan bahwa Random Forest Regressor memberikan kinerja terbaik dengan nilai R² mencapai 0,9393, Root Mean Square Error (RMSE) sebesar 0,0728, dan Mean Absolute Error (MAE) sebesar 0,0345. Kinerja ini jauh lebih baik dibandingkan dengan Decision Tree (R² 0,8372) dan Linear Regression (R² 0,7966). Analisis pentingnya fitur menunjukkan bahwa variabel Jumlah janjang/tandan dan Jumlah pohon sawit memberikan kontribusi paling besar untuk prediksi. kebaruan penelitian ini ada pada pengembangan fitur operasional khusus untuk perkebunan dan penggunaan pembelajaran ensemble pada data produksi nyata tahun 2025. Model ini diharapkan bisa menjadi alat yang membantu dalam pengambilan keputusan untuk mengoptimalkan penggunaan sumber daya saat panen. Kata kunci : Prediksi; Produksi Minyak Sawit; Pembelajaran Mesin; Hutan Acak; Data Operasional;   Abstract - Palm oil production is one of the important parameters to see how efficient the operation in the plantation. This study aims to forecast palm oil production (Fruits per Plant) using Machine Learning algorithms based on operational data in 2025. This dataset consists of 1,347 operational records that include 13 variable features. These variables include land area, labor productivity, as well as derived features such as Total_Harvester_Working_Days and Working_Days per Hectare. The methods used include Linear Regression, Decision Tree Regression, and Random Forest Regression with a data division of 80% for training and 20% for testing. The evaluation results show that Random Forest Regressor provides the best performance with an R² value reaching 0.9393, Root Mean Square Error (RMSE) of 0.0728, and Mean Absolute Error (MAE) of 0.0345. This performance is significantly better than Decision Tree (R² 0.8372) and Linear Regression (R² 0.7966). Feature importance analysis shows that the variables Number of bunches/stems and Number of oil palm trees provide the greatest contribution to the prediction. The novelty of this research lies in the development of operational features specifically for plantations and the use of ensemble learning on real production data for 2025. This model is expected to be a tool that helps in decision-making to optimize resource use during harvest. Keywords: Prediction; Palm Oil Production; Machine Learning; Random Forest; Operational Data;
Analisis Sentimen Ulasan Berbasis Lexicon Menggunakan Pembobotan Tf-IDF dan Algoritma Machine Learning Meltiana; Riski Annisa; Muhammad Fahmi Julianto
Jurnal Nasional Komputasi dan Teknologi Informasi Vol. 9 No. 3 (2026): Juni, 2026
Publisher : Program Studi Teknik Komputer, Fakultas Teknik. Universitas Serambi Mekkah

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.32672/gzkd8504

Abstract

Abstrak - Warung Kopi Asiang merupakan salah satu objek dalam penelitian ini, di mana ulasan pengguna yang terdapat di Google Maps dengan menggunakan kombinasi pendekatan berbasis Lexicon, pembobotan TF-IDF, dan algoritma machine learning. Sebanyak 2.358 ulasan dikumpulkan melalui teknik web scraping menggunakan Instant Data Scraper versi 1.4.1.Proses preprocessing diterapkan pada seluruh data,  mencakup tahapan case folding, cleaning, normalisasi, tokenization, stopword removal, dan stemming.Pelabelan sentimen dilakukan menggunakan metode berbasis lexicon, menghasilkan tiga kategori positif sebanyak 1. 289 ulasan (58,62%), negatif sebanyak 632 ulasan (28,74%), dan netral sebanyak 278 ulasan (12,64%). Data yang telah dilabeli kemudian dikonversi ke dalam bentuk numerik menggunakan TF-IDF dan diklasifikasikan dengan tiga algoritma yakni Support Vector Machine (SVM), Decision Tree, dan Logistic Regression. Optimasi  model dilakukan melalui hyperparameter tuning menggunakan GridSearchCV. Dari hasil evaluasi  Support Vector Machine (SVM) dan Logistic Regression sama-sama mencapai akurasi tertinggi sebesar 80%, di mana pada SVM akurasi optimal tersebut sudah tercapai sejak model dasar (default).Decision Tree menghasilkan akurasi sebesar 71%. Temuan ini membuktikan bahwa integrasi metode Lexicon, TF-IDF, dan machine learning bekerja secara efektif dalam menganalisis sentimen ulasan di Google Maps. Kata kunci : Analisis sentimen; Google Maps; Lexicon; Machine Learning;   Abstract - Warung Kopi Asiang is one of the subjects of this study, in which user reviews found on Google Maps were analysed using a combination of a lexicon-based approach, TF-IDF weighting, and machine learning algorithms. A total of 2,358 reviews were collected via web scraping using Instant Data Scraper version 1.4.1. Preprocessing was applied to all data, covering the stages of case folding, cleaning, normalisation, tokenisation, stopword removal, and stemming; sentiment labelling was performed using a lexicon-based method, resulting in three categories: 1, 289 reviews (58.62%), 632 negative reviews (28.74%), and 278 neutral reviews (12.64%). The labelled data was then converted into numerical form using TF-IDF and classified using three algorithms: Support Vector Machine (SVM), Decision Tree, and Logistic Regression. Model optimisation was carried out via hyperparameter tuning using GridSearchCV. The evaluation results showed that both Support Vector Machine (SVM) and Logistic Regression achieved the highest accuracy of 80%, with SVM having reached this optimal accuracy from the base (default) model. Decision Tree achieved an accuracy of 71%. These findings demonstrate that the integration of the Lexicon method, TF-IDF, and machine learning works effectively in analysing sentiment in Google Maps reviews Keywords: Sentiment analysis; Google Maps; Lexicon; Machine Learning;
Perbandingan Algoritma Machine Learning Berbasis TF-IDF dan SMOTE untuk Analisis Sentimen Ulasan Aming Coffee Emi Wulandari; Riski Annisa; Muhammad Fahmi Julianto
Jurnal Nasional Komputasi dan Teknologi Informasi Vol. 9 No. 3 (2026): Juni, 2026
Publisher : Program Studi Teknik Komputer, Fakultas Teknik. Universitas Serambi Mekkah

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.32672/hk03ka40

Abstract

Abstrak - Ulasan pelanggan pada Google Maps dapat dimanfaatkan untuk mengetahui persepsi pelanggan terhadap kualitas produk dan layanan suatu perusahaan. Namun, banyaknya jumlah ulasan dan data yang tidak terstruktur menyebabkan proses analisis secara manual menjadi kurang efektif. Penelitian ini bertujuan untuk melakukan analisis sentimen pada ulasan pelanggan Aming Coffee menggunakan metode pembobotan Term Frequency–Inverse Document Frequency (TF-IDF) dan algoritma machine learning. Data penelitian diperoleh melalui proses scraping sebanyak 4.000 ulasan pelanggan dari Google Maps Aming Coffee. Tahapan penelitian meliputi preprocessing data yang terdiri atas cleaning, case folding, tokenisasi, normalisasi, stopword removal, dan stemming, kemudian dilakukan pembobotan TF-IDF, pembagian data, klasifikasi menggunakan algoritma Naïve Bayes, Support Vector Machine (SVM), dan Decision Tree C4.5, serta penerapan Synthetic Minority Over-sampling Technique (SMOTE) pada data latih untuk menangani ketidakseimbangan kelas. Evaluasi model dilakukan menggunakan accuracy, precision, recall, F1-score, dan confusion matrix. Hasil penelitian menunjukkan bahwa model tanpa SMOTE memperoleh akurasi tertinggi, yaitu Naïve Bayes sebesar 90,63%, diikuti SVM sebesar 90,50%, dan C4.5 sebesar 88,25%. Setelah penerapan SMOTE, akurasi model mengalami penurunan, namun nilai precision meningkat sehingga model menjadi lebih mampu memperhatikan kelas sentimen minoritas. Hasil Exploratory Data Analysis (EDA) menunjukkan bahwa sentimen positif mendominasi ulasan pelanggan Aming Coffee dengan kata yang paling sering muncul antara lain “kopi”, “coffee”, “enak”, “mantap”, dan “ramai”. Penelitian ini menunjukkan bahwa kombinasi TF-IDF, machine learning, dan SMOTE dapat digunakan untuk analisis sentimen ulasan pelanggan serta memberikan informasi yang bermanfaat dalam evaluasi kualitas layanan dan pengambilan keputusan bisnis. Kata kunci : Analisis Sentimen; TF-IDF; Naïve Bayes; Support Vector Machine; Decision Tree C4.5; SMOTE;   Abstract - Customer reviews on Google Maps can be utilized to understand customer perceptions of a company's products and services. However, the large volume of unstructured reviews makes manual analysis less effective. This study aims to perform sentiment analysis on Aming Coffee customer reviews using the Term Frequency–Inverse Document Frequency (TF-IDF) weighting method and machine learning algorithms. Research data were collected through scraping 4,000 customer reviews from Google Maps. The research stages included data preprocessing consisting of cleaning, case folding, tokenization, normalization, stopword removal, and stemming, followed by TF-IDF weighting, data splitting, sentiment classification using Naïve Bayes, Support Vector Machine (SVM), and Decision Tree C4.5 algorithms, as well as the implementation of Synthetic Minority Over-sampling Technique (SMOTE) on training data to address class imbalance. Model evaluation was conducted using accuracy, precision, recall, F1-score, and confusion matrix. The results showed that models without SMOTE achieved the highest accuracy, with Naïve Bayes reaching 90.63%, followed by SVM at 90.50% and C4.5 at 88.25%. After applying SMOTE, model accuracy decreased, while precision increased, indicating improved attention to minority sentiment classes. Exploratory Data Analysis (EDA) results revealed that positive sentiment dominated Aming Coffee customer reviews, with frequently occurring words including “kopi,” “coffee,” “enak,” “mantap,” and “ramai.” This study demonstrates that the combination of TF-IDF, machine learning algorithms, and SMOTE can be effectively applied to sentiment analysis and provide useful insights for service quality evaluation and business decision-making. Keywords: Sentiment Analysis; TF-IDF; Naïve Bayes; Support Vector Machine; Decision Tree C4.5; SMOTE;