Claim Missing Document
Check
Articles

KLASIFIKASI SENTIMEN MASYARAKAT TERHADAP KASUS TUNTUTAN 17+8 MENGGUNAKAN NAÏVE BAYES CLASSIFIER Arif Kurniawan; Muhammad Fikry; Novi Yanti; Surya Agustian
EXPERT: Jurnal Manajemen Sistem Informasi dan Teknologi Vol 16, No 1 (2026): Juni
Publisher : Universitas Bandar Lampung (UBL)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.36448/expert.v16i1.4947

Abstract

Perkembangan media sosial telah mendorong munculnya berbagai opini masyarakat terhadap isu-isu publik, termasuk kasus Tuntutan 17+8 yang menjadi perhatian luas di Indonesia. Analisis sentimen menjadi pendekatan yang penting untuk mengidentifikasi kecenderungan opini masyarakat secara sistematis. Namun, data teks pada media sosial umumnya bersifat tidak terstruktur dan mengandung berbagai noise sehingga memerlukan tahapan preprocessing yang tepat sebelum dilakukan klasifikasi. Penelitian ini bertujuan untuk menganalisis sentimen masyarakat terhadap kasus Tuntutan 17+8 menggunakan metode Naïve Bayes Classifier serta mengevaluasi pengaruh tahapan preprocessing terhadap performa model melalui pendekatan ablation study. Data penelitian berupa komentar TikTok yang diproses melalui tahapan preprocessing meliputi case folding, normalisasi, stopword removal, dan stemming. Selanjutnya, fitur teks diekstraksi menggunakan TF-IDF dan diklasifikasikan menggunakan algoritma Multinomial Naïve Bayes. Hasil penelitian menunjukkan bahwa setiap kombinasi preprocessing memberikan pengaruh yang berbeda terhadap performa model. Pada dataset tambahan sebanyak 8.370 komentar, performa terbaik diperoleh pada kombinasi case folding dan normalisasi dengan akurasi 89,87% dan F1-score 91,89%. Sementara itu, pada dataset utama sebanyak 1.525 komentar, performa terbaik diperoleh pada kombinasi normalisasi dan stopword removal dengan akurasi 81,20% dan F1-score 80,31%. Hasil evaluasi menggunakan confusion matrix menunjukkan bahwa model memiliki kemampuan klasifikasi yang baik, terutama pada kelas sentimen positif. Penelitian ini membuktikan bahwa pemilihan tahapan preprocessing yang tepat berperan penting dalam meningkatkan performa klasifikasi sentimen menggunakan Naïve Bayes Classifier.
Analisis Komparatif Konfigurasi Multilayer Perceptron pada Classifier Head RoBERTa untuk Klasifikasi Ujaran Kebencian Ikhwan Habibi; Surya Agustian; Jasril Jasril; Muhammad Affandes
TIN: Terapan Informatika Nusantara Vol 7 No 2 (2026): July 2026
Publisher : Forum Kerjasama Pendidikan Tinggi (FKPT)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/tin.v7i2.10512

Abstract

The widespread dissemination of hate speech and offensive language on social media has increased the demand for accurate automated text classification systems. Although RoBERTaForSequenceClassification has been widely used for various for text classification task, the effect of its default classifier head configuration on classification performance has not yet been systematically evaluated. As the main contribution, this study conducts a controlled evaluation of 32 multilayer perceptron (MLP)-based classifier head configurations, varying the number of hidden layers, activation functions, and dropout rates, against the default classifier head on the English HASOC 2021 dataset for two subtasks: binary and multiclass classification. Each configuration was evaluated using Stratified 5-Fold Cross-Validation with Macro-F1 as the evaluation metric, after which the best-performing configuration was further evaluated on an independent test set. For the binary task, the best configuration achieved a test Macro-F1 of 80.90%, about 0.3 percentage points higher than the baseline's 80.59%. For the multiclass task, the configuration with the highest validation performance instead achieved a test Macro-F1 of 65.68%, about 0.4 percentage points lower than the baseline's 66.11%, showing that an advantage observed during cross-validation does not always hold on the test set. Further analysis revealed that excessively deep hidden layers combined with aggressive dimensional compression can sharply degrade performance on the multiclass task. These findings indicate that the effect of classifier head configuration is small and task-dependent, so systematic evaluation remains necessary before adopting a given configuration in place of the default classifier head when fine-tuning RoBERTa-based models.
Klasifikasi Sentimen Pada Dataset yang Terbatas Menggunakan Algoritma Convolutional Neural Network M Ridho Saputra; Surya Agustian; Jasril Jasril; Novriyanto Novriyanto
Bulletin of Computer Science Research Vol. 5 No. 4 (2025): June 2025
Publisher : Forum Kerjasama Pendidikan Tinggi (FKPT)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bulletincsr.v5i4.613

Abstract

This study aims to analyze public responses to the appointment of Kaesang Pangarep as the Chairman of the Indonesian Solidarity Party (PSI) using a sentiment classification approach based on the Convolutional Neural Network (CNN) algorithm. The primary dataset consists of 300 Indonesian-language tweets categorized into three sentiment classes: positive, negative, and neutral. The limited size of the training data presents a major challenge, as it can hinder the model's ability to generalize. To address this issue, data augmentation was carried out by incorporating external datasets with Covid-19 and Open Topic themes. The preprocessing stages include text cleaning, normalization, and tokenization. The developed CNN model utilizes a layered architecture and applies regularization techniques such as L2 and dropout to reduce the risk of overfitting. Accuracy, F1-score, precision, and recall were used as performance evaluation metrics. Experimental results show that the best performance was achieved when the Kaesang and Covid-19 datasets were combined, yielding an F1-score of 0.62 on the validation set and 0.51 on the test set. These findings indicate that adding external data can improve classification accuracy, even under limited data conditions. This study contributes to the development of deep learning-based sentiment classification methods for Indonesian-language texts.
Analisis Sentimen Ulasan Aplikasi Indodax Pada Google Play Store Dengan Algoritma Random Forest Muhammad Iqbal Maulana; Yusra Yusra; Muhammad Fikry; Surya Agustian; Siti Ramadhani
Bulletin of Computer Science Research Vol. 5 No. 4 (2025): June 2025
Publisher : Forum Kerjasama Pendidikan Tinggi (FKPT)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bulletincsr.v5i4.626

Abstract

Crypto assets have become a global phenomenon with a significant increase in the number of investors in Indonesia. Indodax, as the largest crypto asset trading platform in Indonesia, has contributed to the growth of this ecosystem and received many user reviews through the Google Play Store. With more than 5 million downloads and 100 thousand reviews, sentiment analysis is an important tool to understand user perceptions of Indodax services. The results of manual labeling show that the majority of reviews are positive (3989 reviews), while neutral and negative sentiments are 477 and 534 reviews respectively. From the research and testing that has been carried out using the Random Forest method and optimizing with Hyperparameter Tuning GridSearchCV on 4 test scenarios. The best results were obtained in Scenario 4 (3 Preprocessing Stages (Cleaning, Case Folding, and Tokenization) + Random Forest & Hyperparameter Tuning) producing the best value, with Precision 81%, Recall 64%, F1-Score 70% and Accuracy 89%. With the best parameter values ??{'criterion': 'entropy', 'max_depth': None, 'max_features': 'sqrt', 'min_samples_leaf': 1, 'min_samples_split': 2, 'n_estimators': 100}. This study shows that every experimental model that is optimized produces a higher value than experimental model that is not optimized.
Optimasi Klasifikasi Hate Speech dan Offensive Language melalui Frozen RoBERTa Feature Extraction dan Random Forest Marsha Cahyani Dwisyakilla; Surya Agustian; Novriyanto Novriyanto; Muhammad Affandes
Bulletin of Computer Science Research Vol. 6 No. 4 (2026): June 2026
Publisher : Forum Kerjasama Pendidikan Tinggi (FKPT)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bulletincsr.v6i4.1157

Abstract

Hate speech and offensive content detection on social media remains a significant challenge in Natural Language Processing (NLP) due to the characteristics of Twitter data, which are typically short, informal, and contain various elements such as mentions, URLs, hashtags, and emotional expressions that complicate the classification process. End-to-end Transformer fine-tuning approaches generally require substantial computational resources; therefore, this study explores a more computationally efficient approach by utilizing RoBERTa as a frozen feature extractor combined with Random Forest as the classifier. This approach enables the exploitation of contextual representations generated by Transformer models without requiring full model retraining.The study employs the HASOC 2021 English Track dataset, which consists of two classification tasks: Task A for binary classification (HOF and NOT) and Task B for multi-class classification (HATE, OFFN, PRFN, and NONE). The classification pipeline is optimized through the incorporation of handcrafted features, oversampling, Random Forest hyperparameter tuning, and threshold tuning in specific scenarios. Model performance is evaluated using accuracy, precision, recall, and F1-macro, with F1-macro serving as the primary metric due to class imbalance. The best-performing model achieved an F1-macro score of 0.80 on Task A and 0.64 on Task B. These results indicate that the combination of frozen RoBERTa representations and Random Forest provides strong performance for binary hate speech and offensive content classification. However, the performance on Task B highlights the difficulty of distinguishing linguistically similar categories, such as HATE, OFFN, and PRFN, suggesting that fine-grained multi-class classification remains a challenging task. Overall, the findings indicate that RoBERTa-based frozen feature extraction constitutes a computationally efficient alternative for hate speech detection on English Twitter data, although further improvements are required to enhance performance in multi-class classification settings.
Analisis Efektivitas IndoBERT untuk Klasifikasi Multilabel Terjemahan Hadis Bukhari Menggunakan Logistic Regression Achmad Yamin Harahap; Nazruddin Safaat H; Surya Agustian; Suwanto Sanjaya; Teddie D
Bulletin of Computer Science Research Vol. 6 No. 4 (2026): June 2026
Publisher : Forum Kerjasama Pendidikan Tinggi (FKPT)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bulletincsr.v6i4.1219

Abstract

Hadith serves as the second source of guidance after the Quran, directing Muslims in various aspects of life; the *Sahih al-Bukhari* collection is among the most renowned. The complex nature of their meanings often encompassing multiple categories of messages poses a significant challenge for manual text classification, particularly as data volume grows. In this study, the content of the hadith often includes multiple message types, such as recommendations, prohibitions, and general information. This research aims to evaluate an automated classification system for Indonesian translations of *Sahih al-Bukhari* hadith, categorizing them into three classes: Information, Recommendation, and Prohibition. The study is motivated by the vast number of hadith, which requires significant time and deep understanding for people to grasp the core message of each one. This classification system is intended to facilitate the identification of primary messages, thereby making the processes of searching, studying, and understanding hadith more effective and efficient. IndoBERT is employed to generate contextual vector representations capable of capturing deeper semantic meaning, while Logistic Regression is selected for its efficiency and stability with high-dimensional data. Evaluation is conducted using a train-validation-test split approach, alongside accuracy and macro F1-score metrics. The study achieved an average F1-score of 67.43%, demonstrating that the combination of IndoBERT and Logistic Regression yields strong, consistent classification performance for this multi-label task.
Enhancing Hate Speech and Offensive Language Detection using CatBoost with RoBERTa-based Contextual Embeddings Muhammad Elfarizi; Surya Agustian; Fitra Kurnia; Suwanto Sanjaya; Fitri Insani
SISTEMASI Vol 15, No 7 (2026): Sistemasi: Jurnal Sistem Informasi
Publisher : Universitas Islam Indragiri

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.32520/stmsi.v15i7.6637

Abstract

The widespread dissemination of hate speech and offensive content on social media platforms has become a critical societal issue, highlighting the need for reliable automated detection systems. This study proposes a hybrid approach that leverages frozen embeddings from the pre-trained language model cardiffnlp/twitter-roberta-base-offensive as a high-level semantic feature extractor, combined with the CatBoost gradient boosting algorithm as the final classifier. The proposed method was evaluated on the HASOC 2021 English dataset through six experimental scenarios and compared with a TF-IDF baseline using CatBoost's default hyperparameters. Experimental results demonstrate that the proposed approach achieved a Macro F1-score of 0.7924 for the binary classification task (Task 1A) and 0.6113 for the multiclass classification task (Task 1B), outperforming the TF-IDF baseline, which achieved scores of 0.7724 and 0.5798, respectively. The proposed system demonstrated a clear performance improvement and achieved results comparable to those of the top-ranked teams on the official HASOC 2021 leaderboard, while avoiding the computational cost associated with fine-tuning large pre-trained language models.
Klasifikasi Hate Speech dan Offensive Language Menggunakan BERT dan Support Vector Machine Muhammad Tirta Syakban; Surya Agustian; Muhammad Fikry; Muhammad Affandes
Bulletin of Computer Science Research Vol. 6 No. 3 (2026): April 2026
Publisher : Forum Kerjasama Pendidikan Tinggi (FKPT)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bulletincsr.v6i3.1061

Abstract

Hate speech and offensive language have become increasingly complex problems on social media, requiring classification approaches that can effectively capture linguistic context. While transformer-based models with end-to-end fine-tuning have become the dominant approach, the use of transformers as fixed feature extractors combined with classical machine learning algorithms remains relatively underexplored, particularly in benchmark settings such as HASOC 2021. This study aims to investigate the effectiveness of a feature-based transformer approach by combining embeddings from BERT and RoBERTa with Support Vector Machine (SVM) classifiers using multiple kernel configurations, including Linear, RBF, Polynomial, and LinearSVC. Experiments were conducted on Sub-task A and Sub-task B by comparing traditional feature-based methods (TF-IDF) with transformer-based embeddings. The experimental results show that RoBERTa embeddings consistently outperform other feature extraction methods. On the test dataset, the combination of RoBERTa and SVM achieves competitive performance compared to other systems in HASOC 2021. In Sub-task B, the optimal model achieves a Macro F1-score of 0.61, outperforming several BERT-based and classical baseline systems.These findings demonstrate that using transformer embeddings as fixed feature representations combined with optimized SVM classifiers can serve as an effective alternative to fine-tuning approaches, particularly in achieving more stable performance under class imbalance conditions. This study contributes by highlighting the potential of feature-based transformer methods as a flexible and competitive strategy for hate speech and offensive language detection.
Co-Authors .Safrizal, Safrizal Abdillah, Rahmad Achmad Yamin Harahap Afdhal Zikri Afriyanti, Liza Aftari, Dhea Putri AGUNG SUCIPTO Ahmad Kurniawan Ahmad, Rizmah Zakiah Nur Alfitra Salam Arasy, Abdurrahman Arif Kurniawan Ash Shiddicky Aulia Ramadhani Ayu Fransiska Baehaqi Dermawan, Jozu Dzaky Abdillah Salafy Eka Pandu Cynthia El Saputra, Yoga Elin Haerani Elvia Budianita Fadhilah Syafria Faizah Husniah Fauzan Ray T Fauzi Ihsan Febi Yanto Febrian Rizki Adi Sutiyo Fitra Kurnia Fitri Insani Fitri Insani Fitri, Dina Deswara Fuji Astuti Gusti, Siska Kurnia Habib Hakim Sinaga Hadi, Mukhlis Halimah Heru Wibowo Idhafi, Zaky Iffa, Marwika Rifattul Ihsan, Miftahul Iis Afrianty Iis Afrianty Ikhwan Habibi Ilham Habibi Hasibuan Illahi, Ridho Iman Fauzi Aditya Sayogo Indri Pangestuti Iwan Iskandar Jasril Jasril Jasril Jasril Jasril Jasril Jauhari, Najwa Lestari Handayani Lubis, Anggun Tri Utami BR. M Ridho Saputra Marsha Cahyani Dwisyakilla Melyana Hasibuan Miftah Farid Muhammad Affandes Muhammad Affandes Muhammad Elfarizi Muhammad Fikry Muhammad Fikry Muhammad Iqbal Maulana Muhammad Irsyad Muhammad Irsyad Muhammad Ravil Muhammad Tirta Syakban Muktar Sahbuddin Mukti M Kusairi Mulyadi, Syahrul Nadila Handayani Putri naldi, Afri Nazir, Alwis Nazruddin Safaat Nazruddin Safaat H Nazruddin Safaat H Nazruddin Safaat H Negara, Benny Sukma Novi Yanti Novriyanto Novriyanto Novriyanto Novriyanto Nurika Dwi Wahyuni Nurul Fatiara Okfalisa Okfalisa Oktavia, Lola Pangestu, Yoga Pizaini Pizaini Pranata, Joni Prima Yohana Putri Zahwa Putri, Adilah Atikah Putri, Atika Rahmad Abdillah Rahmad Kurniawan Ramadhani, Siti Reski Mai Candra Reski Mai Candra Rizqa Raaiqa Bintana Safrizal, Afri Naldi Salam Kurniawan Saputra, Ikhsan Dwi Saputra, Nugroho Wahyu Satira, Husna Sinaga, Habib Hakim Siti Ramadhani Siti Ramadhani Sri Puji Utami A. Subhi, Yazid Abdullah Suci Rahayu Sulistia Ningsih, Sulistia Suwanto Sanjaya Syaiful Azhar Tarmizi, Veci Cahyono Teddie D Trya Ayu Pratiwi Utari, Roid Fitrah Wahyu Reinaldy Wan Sobri Amin Yusra Yusra Yusra Yusra Yusra, Yusra