Claim Missing Document
Check
Articles

Found 13 Documents
Search

Classification of Online Gambling Spam Comments on YouTube Using Support Vector Machine Umbu Anaagung Pariamalinya; Josua Josen A. Limbong; Julius Panda Putra Naibaho
Indonesian Journal of Artificial Intelligence and Data Mining Vol. 9 No. 1 (2026): March 2026
Publisher : Universitas Islam Negeri Sultan Syarif Kasim Riau

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

While digital transformation has established YouTube as a major communication platform, the site has also become vulnerable to online gambling spam in Indonesia. This study investigates the effectiveness of the Support Vector Machine (SVM) algorithm for automated spam detection as an alternative to manual moderation. A total of 9,169 comments were collected from gaming, education, and entertainment channels using the YouTube Data API v3 and were used to train and evaluate the model with an 80:20 data split. The experimental results show that SVM achieved an accuracy of 99.62% and an F1-score of 0.996, demonstrating strong capability in identifying spam comments written in informal and modified promotional language. The main contribution of this study is the development of a highly accurate and practical spam detection approach for Indonesian YouTube comments, which can support more efficient moderation systems. However, the model still has limitations in detecting sarcastic content. Therefore, future research should explore deep learning models such as BERT to improve contextual understanding and strengthen automated moderation in digital environments.
Clustering HIV Screening Data in Teluk Bintuni Using K-Means Yuliana Manobi; Alex De Kweldju; Julius Panda Putra Naibaho
J-INTECH ( Journal of Information and Technology) Vol 14 No 02 (2026): Journal of Information and Technology
Publisher : LPPM Universitas Bhinneka Nusantara

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.32664/j-intech.v14i02.2368

Abstract

Human Immunodeficiency Virus (HIV) remains a major public health challenge in Papua, Indonesia, where geographical barriers and limited resources complicate service delivery. This study applies clustering methods to HIV screening data from 22 Health Service Units (UPK) in Teluk Bintuni Regency during 2024–2025. The final dataset consisted of 29 valid unit-year observations containing variables related to total tests, HIV-positive cases, and sex-disaggregated distributions. Data preprocessing followed the Knowledge Discovery in Database (KDD) framework, including data selection, cleaning, log transformation, normalization, and the construction of derived variables such as positivity rate and gender ratios. Four clustering algorithms were compared, namely K-Means, Hierarchical Clustering, Fuzzy C-Means, and DBSCAN, using Silhouette Score, Davies-Bouldin Index, and Calinski-Harabasz Index. The results indicate that K-Means produced the most stable and interpretable clustering structure, forming three groups of UPKs: intermediate screening units with moderate coverage and low positivity, priority units with limited testing but high positivity rates, and active screening units with the highest testing volume and case detection. These findings reveal heterogeneity in HIV screening performance across UPKs and support differentiated intervention strategies. Units with high apparent positivity but low testing coverage require expanded outreach, field verification, and improved access to HIV screening services, while active screening units should be strengthened through counseling, referral, and follow-up services. This study demonstrates the usefulness of clustering analysis in identifying service gaps and supporting evidence-based HIV intervention planning at the local level.
PENERAPAN ALGORITMA BERT UNTUK DETEKSI KOMENTAR SPAM JUDI ONLINE DI YOUTUBE Fahri Akbar Rosid Asro; Julius Panda Putra Naibaho; Alex De Kweldju
JIPI (Jurnal Ilmiah Penelitian dan Pembelajaran Informatika) Vol 11, No 2 (2026)
Publisher : STKIP PGRI Tulungagung

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.29100/jipi.v11i2.8122

Abstract

Penyebaran komentar spam yang mengandung promosi perjudian daring di YouTube semakin meningkat dan mengganggu kualitas interaksi digital. Deteksi manual tidak lagi efektif karena volume komentar yang tinggi dan kompleksitas gaya bahasa yang digunakan. Penelitian ini bertujuan untuk menerapkan model In-doBERT yang telah di‑fine‑tune guna mendeteksi komentar spam bertema perjudian dalam bahasa Indonesia. Data mentah diperoleh melalui scraping sebanyak 9.056 komentar dari 14 video kategori game. Setelah dilakukan pra‑pemrosesan yang meliputi pembersihan teks, normalisasi, tokenisasi, penghapusan stopword, dan stemming, 263 komentar dihapus sehingga tersisa 8.793 komentar bersih untuk pelabelan. Komentar dikategorikan secara semi‑otomatis ke dalam dua kelas, spam dan non‑spam, dengan bantuan pencocokan kata kunci dan validasi manual. Model IndoBERT dilatih menggunakan pendekatan transfer learning dan dievaluasi dengan metrik akurasi, presisi, recall, dan F1‑score. Hasil eksperimen menunjukkan akurasi sekitar 99,54% (atau dibulatkan menjadi 99,6%) dan F1‑score 1,00, dengan hanya delapan kesalahan klasifikasi dari 1.759 data uji. Vis-ualisasi melalui wordcloud dan histogram mengungkap pola domi-nan istilah promosi dalam komentar spam. Temuan ini menunjukkan bahwa pendekatan berbasis transformer efektif dalam memahami konteks semantik, meskipun disamarkan melalui eufemisme atau kode, serta dapat diandalkan untuk meningkatkan akurasi sistem moderasi otomatis pada platform digital.