Claim Missing Document
Check
Articles

Found 1 Documents
Search

Identifikasi Pesan Penipuan Berkedok Hadiah Digital Menggunakan Machine Learning Muhammad Shafari Rahmat; Syahrizal Andhika
Jurnal Teknik dan Science Vol. 5 No. 2 (2026): Juni : Jurnal Teknik dan Science
Publisher : Asosiasi Dosen Muda Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.56127/jts.v5i2.2948

Abstract

Digital prize scams are increasingly distributed through text messages, social media, and instant messaging platforms using persuasive expressions, suspicious links, and urgent instructions. This study aims to develop a machine learning-based approach for identifying Indonesian-language scam messages disguised as digital prize notifications. The dataset consisted of 6,000 text messages, comprising 3,000 scam messages and 3,000 non-scam messages. The data were manually labeled by two independent annotators and processed through case folding, text cleaning, tokenization, stopword removal, and stemming. Text features were transformed into numerical representations using Term Frequency–Inverse Document Frequency (TF-IDF). Five classification algorithms were evaluated, including Multinomial Naive Bayes, Support Vector Machine, Logistic Regression, Random Forest, and XGBoost. Model performance was assessed using stratified five-fold cross-validation based on accuracy, precision, recall, F1-score, and ROC-AUC. The results showed that XGBoost achieved the best performance, with an accuracy of 97.0%, precision of 97.2%, recall of 96.8%, F1-score of 97.0%, and ROC-AUC of 99.3%. Scam messages were commonly characterized by words related to prizes, winning notifications, free offers, claims, and urgency. These findings indicate that TF-IDF combined with XGBoost can effectively support the automatic detection of digital prize scam messages. Future studies should incorporate URL characteristics, sender metadata, and contextual language models to improve detection performance.