Claim Missing Document
Check
Articles

Found 3 Documents
Search

Hate Speech Analysis of YouTube Comments on the 2024 Indonesian Presidential Debate Using IndoBERT Agus Sasmito Aribowo; Yuli Fauziah; Yusna Bantulu; Shoffan Saifullah; Azfa Mutiara Ahmad Fubalo
Kinetik: Game Technology, Information System, Computer Network, Computing, Electronics, and Control Vol. 11, No. 3, August 2026
Publisher : Universitas Muhammadiyah Malang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.22219/kinetik.v11i3.2604

Abstract

The rapid digitization of political campaigns has intensified the spread of hate speech on social media, threatening democratic discourse and social cohesion. In Indonesia, YouTube comments on the 2024 presidential election debates have emerged as a critical yet underexplored source of polarizing content. However, existing detection systems struggle with Indonesian-specific linguistic features, code-mixing, and implicit political sarcasm, while high annotation costs and class imbalance further limit the scalability of supervised approaches. To address these challenges, this study introduces a large-scale dataset of 38,742 YouTube comments collected from the five official debate stages and labeled using a cost-effective semi-supervised framework (20% expert-annotated, 80% pseudo-labeled). We systematically evaluate four classification models —IndoBERT, mBERT, SVM, and Random Forest—under identical experimental conditions using evaluation metrics optimized for imbalanced data. Experimental results demonstrate that IndoBERT consistently outperforms all baseline models, achieving an average accuracy of 89.7% and a macro F1-score of 0.89 across all debate stages. Notably, IndoBERT maintains high recall for the minority hate speech class (0.81–0.90), confirming its superior ability to capture localized political rhetoric and contextual nuances that multilingual and classical models frequently miss. This study contributes a publicly available Indonesian political hate speech dataset, validates a scalable semi-supervised annotation pipeline, and provides empirical evidence that domain-specific Transformer models are essential for reliable content moderation in politically charged, low-resource environments.
Development of a Web-Based Smart Ecosystem Platform for Sustainable Public Service Automation Shoffan Saifullah; Muhammad Iqbal; Lisnawanty Lisnawanty; Weiskhy Steven Dharmawan; Fahmi Raditya
Jurnal Infortech Vol. 8 No. 1 (2026): June 2026
Publisher : LPPM Universitas Bina Sarana Informatika

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.31294/infortech.v8i1.12830

Abstract

Public service delivery in Indonesia continues to face fundamental challenges including inefficient manual administrative processes, error-prone document validation, and the absence of real-time tracking systems. This research aims to develop the PANDU (Pelayanan Publik Digital Terpadu) platform as a web-based smart ecosystem that automates public services sustainably. The platform is built using the Waterfall development method with Model-View-Controller architecture based on Laravel 12 framework, Filament 4.0 administration panel, and Tailwind CSS 4.0 responsive interface. Four main smart features are integrated: automatic document validation, duplicate request detection within a 30-day window, category-based related service recommendations, and automatic priority calculation using multi-criteria scoring algorithm. The platform produces three separate panels for citizens, officers, and administrators, equipped with real-time tracking system through public API and configurable multi-step approval workflows. Black box testing results using equivalence partitioning technique demonstrate one hundred percent functional success rate, while usability evaluation using System Usability Scale yields an average score in the Excellent category with Acceptable acceptability level. The PANDU platform successfully bridges the gap between smart government theoretical frameworks and operational implementation, providing significant contribution to accelerating sustainable digital transformation of public services in Indonesia.
Comparative Analysis of Email Spam Detection Using SVM with TF-IDF and Word2Vec on Multilingual Datasets Kaifa Ahlal Katamsyi; Ahmad Taufiq Akbar; Andi Nurkholis; Hari Prapcoyo; Bagus Muhammad Akbar; Shoffan Saifullah
Paradigma - Jurnal Komputer dan Informatika Vol. 28 No. 1 (2026): March 2026 Period
Publisher : LPPM Universitas Bina Sarana Informatika

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.31294/p.v28i1.12339

Abstract

The rapid growth of email communication has increased the prevalence of spam emails, which can disrupt productivity and compromise information security. This study presents a comparative analysis of two text representation methods—TF-IDF and Word2Vec—for spam email classification using a Support Vector Machine (SVM) with a Radial Basis Function kernel. The experiments utilized Indonesian and English email datasets totaling 5,421 emails, split into 75% training and 25% testing sets. Two scenarios were evaluated: baseline with default parameters and after hyperparameter optimization using Grid Search combined with K-Fold Cross Validation. The results indicate that TF-IDF consistently outperformed Word2Vec across both languages, achieving the highest accuracy of 0.9562 on the English dataset after tuning. Word2Vec showed substantial improvement following parameter adjustment, reducing the performance gap with TF-IDF. The findings highlight the importance of hyperparameter optimization for enhancing the quality of feature representations and improving classification performance. This study also demonstrates that TF-IDF provides more stable results across different linguistic contexts, while Word2Vec benefits significantly from careful tuning. The results provide practical insights for implementing efficient spam email detection systems in multilingual environments. Future research could explore additional classifiers, deep learning approaches, and contextual embeddings to further improve classification accuracy and robustness.