Claim Missing Document
Check
Articles

Klasifikasi Sentimen Terhadap Topik Pindah Ibu Kota Negara Pada Twitter Menggunakan Metode Naïve Bayes Classifier Dermawan, Jozu; Yusra, Yusra; Fikry, Muhammad; Agustian, Surya; Oktavia, Lola
Jurnal Sistem Komputer dan Informatika (JSON) Vol. 5 No. 3 (2024): Maret 2024
Publisher : Universitas Budi Darma

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30865/json.v5i3.7475

Abstract

Towards the middle of 2019, President Joko Widodo announced plans to relocate Indonesia's capital city. This caused pros and cons in the community, which were widely observed in various social media. To quickly measure the level of public sentiment towards the policy of moving the National Capital City (IKN), whose construction is already underway, a classification system that has good performance is needed. This research proposes a classification of public sentiment on the topic using the Naïve Bayes Classifier method. The data used in this study amounted to 4000 tweets that have been classified into two classes, namely 2000 positive class data and 2000 negative class data. The purpose of this research is how to apply the Naïve Bayes Classifier method in classifying sentiment on the topic of moving the nation's capital and determine the accuracy level of the method. The application of the Naïve Bayes classification method using TF-IDF features to classify 10% of the data as testing data resulted in an accuracy of 77.00%, for a precision value of 77.06%, recall 77.08% and f1-score of 77.00%. Based on the results achieved, the Naïve Bayes Classifier method is good at text classification tasks, with a fairly good accuracy rate.
Klasifikasi Sentimen Masyarakat Terhadap Kaesang Pangarep pada Media Sosial Twitter/X Menggunakan MLP Classifier dengan Fitur FastText Tarmizi, Veci Cahyono; Agustian, Surya; Okfalisa, Okfalisa; Pizaini, Pizaini
TIN: Terapan Informatika Nusantara Vol 6 No 7 (2025): December 2025
Publisher : Forum Kerjasama Pendidikan Tinggi (FKPT)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/tin.v6i7.8815

Abstract

Social media has become a primary channel for the public to express their opinions and reactions toward various political developments in Indonesia. One of the prominent discussions revolves around Kaesang Pangarep’s appointment as the Chairman of the Indonesian Solidarity Party (PSI). This study aims to analyze and classify public sentiment regarding this issue by employing the Multi-Layer Perceptron (MLP) algorithm integrated with FastText-based text representation. The dataset was collected from Twitter using keywords such as “Kaesang PSI”, and was further expanded with additional data from general topics including Covid-19 and Open Topic, ensuring a balanced distribution across positive, neutral, and negative sentiment categories for a more comprehensive representation of public opinion. The model’s performance was evaluated through four metrics: accuracy, precision, recall, and F1 Score. The experimental results demonstrate that the MLP–FastText model achieved consecutive scores of 0. 5129 for F1 Score, 0. 6035 for accuracy, 0. 5319 for precision, and 0. 5996 for recall. These findings indicate that the combination of MLP and FastText effectively captures sentiment patterns within textual data, particularly in the context of unstructured and dynamic social media content, and performs well when enhanced with relevant external data augmentation strategies.
Classification of Phishing URL Attacks Using Random Forest Algorithm Based on Feature Importance Melyana Hasibuan; Rahmad Abdillah; Surya Agustian; Reski Mai Candra
bit-Tech Vol. 8 No. 2 (2025): bit-Tech
Publisher : Komunitas Dosen Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.32877/bt.v8i2.3511

Abstract

The development of information technology and increasing digital activities have made URL-based phishing threats more complex and difficult to detect. Phishing attacks target not only individuals but also organizations, requiring detection systems that are accurate, efficient, and capable of handling high-dimensional data. Machine learning approaches, particularly Random Forest, have been widely applied for phishing detection; however, further evaluation is needed regarding the role of feature selection in improving efficiency without reducing performance. This study aims to evaluate the performance of the Random Forest algorithm for phishing URL detection and to analyze the impact of feature selection based on feature importance. This research adopts the Knowledge Discovery in Databases (KDD) framework, including data selection, preprocessing, feature selection, modeling, and evaluation stages. The PhiUSIIL-2024 dataset is used, with two modeling scenarios: Random Forest using all features (RF Full) and Random Forest using the top 30 features selected through feature importance (RF Top-30). Model performance is evaluated using accuracy, precision, recall, and F1-score metrics under different data split ratios. The experimental results show that both models achieve very high and stable classification performance, with evaluation metrics close to or reaching 100%. The RF Top-30 model maintains performance comparable to the RF Full model despite using fewer features. This study concludes that feature importance-based feature selection effectively simplifies the Random Forest model without sacrificing performance, making it suitable for efficient URL phishing detection systems.
Klasifikasi Sentimen Komentar Youtube Tentang Pembatalan Indonesia Sebagai Tuan Rumah Piala Dunia U-20 Menggunakan Algoritma Naïve Bayes Classifer Ilham Habibi Hasibuan; Elvia Budianita; Surya Agustian; Pizaini Pizaini
Jurnal Sistem Komputer dan Informatika (JSON) Vol. 5 No. 2 (2023): Desember 2023
Publisher : Universitas Budi Darma

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30865/json.v5i2.7096

Abstract

Text mining is a method used to perform tasks such as document classification, clustering, information extraction, sentiment analysis, and information retrieval. The Federation Internationale Football Association (FIFA), the international football governing body, has designated Indonesia as the host country for the U-20 World Cup starting in 2019. Indonesia is expected to be the choice venue for the U-20 World Cup in 2021. However, due to the Covid outbreak -19, the World Cup was rescheduled and is now scheduled to take place in 2023. Indonesia officially relinquished its position as host on March 31 2023. One of the reasons is the many factions that oppose the presence of the Israeli national team in Indonesia. As a result, various public reactions responded to Indonesia's decision to cancel holding the U-20 World Cup, especially on the Narasi tv YouTube channel video entitled "The U-20 World Cup Failed to Be Held in Indonesia, Let's Look at it from Two Perspectives | Discussion". Since the video was uploaded until August 16 2023, the total comments generated were 4,629 comments. This research uses a Naïve Bayes classifier approach. Naïve Bayes Classifier (NBC) is a direct probabilistic classifier that exploits Bayes' Theorem under strong independence conditions. The tests carried out show that the model performance when using stopword removal and stemming techniques is superior in classifying classes in the dataset. The F1-Score is 59.70% and the Accuracy value is 63.43%. Furthermore, after identifying the most efficient model for applying naïve Bayes classification, evaluation was carried out on validation data resulting in an F1-Score of 58.72% and an accuracy rate of 61.65%. Classification analysis shows that Indonesian people have a negative view or are disappointed with the cancellation
SENTIMENT CLASSIFICATION OF PUBLIC PERCEPTIONS ON RP200 TRILLION HIMBARA STIMULUS USING NAÏVE BAYES Wan Sobri Amin; Muhammad Fikry; Rahmad Abdillah; Surya Agustian
Jurnal Riset Informatika Vol. 8 No. 2 (2026): Maret 2026
Publisher : Kresnamedia Publisher

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.34288/jri.v8i2.500

Abstract

The government's policy in the form of a fund stimulus of Rp200 trillion to the Himpunan Bank Milik Negara (HIMBARA) is a strategic step to maintain national economic stability and encourage real sector recovery. However, the implementation of public policy is inseparable from the response and public perception that develops on social media. This study aims to classify public sentiment towards the Rp200 trillion fund stimulus policy to Bank HIMBARA based on Instagram user comments and test the performance of the Naïve Bayes Classifier method in analyzing public policy sentiment. This study uses a quantitative approach with text mining and machine learning methods. Data in the form of 1.309 Instagram comments was collected through web scraping techniques from several online media accounts, then processed through text preprocessing and manual labeling stages into positive, neutral, and negative sentiments. Feature weighting was carried out using TF-IDF, then the data were classified using Multinomial Naïve Bayes and Complement Naïve Bayes. The results show that the Complement Naïve Bayes model achieved the best performance with an accuracy of 84%, an F1-score of 81%, and a high ROC-AUC value. These findings indicate that the majority of public sentiment toward the stimulus policy tends to be positive, and that the Naïve Bayes method is effective for social media–based sentiment analysis.
Penerapan Saliency Maps dalam Explainable AI Untuk Deteksi Penyakit Paru-Paru pada Citra X-Ray Dada dengan Deep Learning Wahyu Reinaldy; Benny Sukma Negara; Muhammad Irsyad; Muhammad Affandes; Surya Agustian
TIN: Terapan Informatika Nusantara Vol 7 No 1 (2026): June 2026
Publisher : Forum Kerjasama Pendidikan Tinggi (FKPT)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/tin.v7i1.9962

Abstract

Early identification of lung diseases is very important so that medical personnel can quickly provide first aid and further study the patient's condition. In this study, a model was developed to classify chest X-ray images of the lungs using the VGG16 architecture. These chest X-ray images were categorized into three groups: COVID-19, normal lungs, and pneumonia. A combination of hyperparameters, including a learning rate of 0.001, 50 epochs, and a batch size of 16, was used to train the model, achieved an accuracy of 96%. Several evaluation metrics, including precision, recall, f1-score, and confusion matrix, were used to assess the model. In addition, saliency map methods were used to visually interpret the model's prediction output and display the areas of the chest X-ray images that most influenced the model's decision-making. The saliency map visualization findings show that the model focuses its predictions on regions of the lungs associated with the disease, which helps in understanding the algorithm's decision-making process.
Integrasi Efficient Channel Attention (ECA) pada DenseNet169 untuk Klasifikasi Multi-Kelas Citra X-ray Dada: Integration of Efficient Channel Attention (ECA) in DenseNet169 for Multi-Class Chest X-ray Classification Satira, Husna; Negara, Benny Sukma; Irsyad, Muhammad; Jasril, Jasril; Agustian, Surya
MALCOM: Indonesian Journal of Machine Learning and Computer Science Vol. 6 No. 3 (2026): MALCOM July 2026
Publisher : Institut Riset dan Publikasi Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.57152/malcom.v6i3.2799

Abstract

Penyakit paru seperti pneumonia dan COVID-19 masih menjadi masalah kesehatan global yang memerlukan deteksi dini dan akurat. X-ray dada merupakan modalitas pencitraan yang umum digunakan, namun interpretasinya masih bergantung pada keahlian serta ketelitian radiolog. Penelitian ini bertujuan untuk mengevaluasi efektivitas integrasi Efficient Channel Attention (ECA) pada arsitektur DenseNet169 untuk klasifikasi multi-kelas pada citra X-ray dada. ECA dipilih karena merupakan mekanisme attention yang ringan serta mampu menangkap hubungan antarkanal fitur dengan kompleksitas parameter yang rendah. Dataset yang digunakan terdiri dari 5.228 citra yang terbagi ke dalam tiga kelas, yaitu COVID-19, pneumonia, dan normal. Penelitian ini membandingkan model DenseNet169 baseline dan DenseNet169 dengan ECA menggunakan parameter pelatihan yang sama untuk memastikan perbandingan yang adil. Evaluasi dilakukan menggunakan metrik accuracy, precision, recall, dan F1-score. Hasil menunjukkan bahwa model baseline memperoleh accuracy sebesar 96,75%, sedangkan model dengan ECA memperoleh 96,27%. Dengan demikian, integrasi ECA belum memberikan peningkatan performa yang signifikan dibandingkan dengan DenseNet169 pada dataset yang digunakan. Temuan ini menunjukkan bahwa efektivitas ECA dipengaruhi oleh karakteristik arsitektur model, strategi integrasi, serta karakteristik dataset, sehingga penerapan attention mechanism tidak selalu menghasilkan peningkatan performa secara langsung.
Klasifikasi Hate Speech dan Offensive Language Menggunakan Hybrid RoBERTa dan XGBoost dengan Optimasi Hyperparameter: Hate Speech and Offensive Language Classification Using Hybrid RoBERTa and XGBoost with Hyperparameter Optimization Jauhari, Najwa; Agustian, Surya; Syafria, Fadhilah; Affandes, Muhammad
MALCOM: Indonesian Journal of Machine Learning and Computer Science Vol. 6 No. 3 (2026): MALCOM July 2026
Publisher : Institut Riset dan Publikasi Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.57152/malcom.v6i3.2834

Abstract

Ujaran kebencian dan bahasa ofensif di media sosial merupakan masalah serius yang membutuhkan deteksi otomatis yang akurat. Penelitian ini mengusulkan pendekatan hybrid yang menggabungkan model Robustly Optimised BERT Approach (RoBERTa) berbasis Twitter (cardiffnlp/twitter?roberta?base?offensive) sebagai ekstraktor fitur dengan algoritma XGBoost untuk klasifikasi pada dataset Hate Speech and Offensive Content Identification (HASOC) 2021 berbahasa Inggris. Kebaruan penelitian ini terletak pada integrasi optimasi hyperparameter Optuna, Multi?Seed Ensemble, Division Calibration, dan oversampling pada ruang fitur embedding yang belum pernah diterapkan secara bersamaan dalam satu pipeline pada HASOC 2021. Dua tugas diselesaikan : Task 1A adalah klasifikasi biner untuk membedakan konten Hate and Offensive (HOF) dari yang tidak (NOT). Task 1B adalah klasifikasi multi-kelas yang membagi tweet menjadi empat kategori: (ujaran kebencian terhadap kelompok) HATE, bahasa ofensif terhadap individu (OFFN), kata kasar tanpa target spesifik (PRFN), dan konten aman (NONE).  Ketidakseimbangan kelas ditangani dengan oversampling pada ruang fitur, dan optimasi hyperparameter dilakukan untuk meningkatkan performa. Hasil evaluasi pada data uji menunjukkan Macro F1?Score sebesar 0,8083 untuk Task 1A dan 0,6541 untuk Task 1B. Perbandingan dengan papan peringkat HASOC 2021 menunjukkan bahwa skor tersebut sebanding dengan tim peringkat 5 (Task 1A) dan peringkat 3 (Task 1B) yang berbasis BERT murni.
Analisis Sensitivitas Confidence Threshold pada Semi-Supervised FixMatch untuk Klasifikasi Multi-Kelas Citra Chest X-Ray Ahmad Kurniawan; Muhammad Irsyad; Benny Sukma Negara; Surya Agustian; Nazruddin Safaat H
TIN: Terapan Informatika Nusantara Vol 7 No 1 (2026): June 2026
Publisher : Forum Kerjasama Pendidikan Tinggi (FKPT)

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/tin.v7i1.10028

Abstract

Optimizing the confidence threshold in pseudo-labeling is a critical technical challenge in Semi-Supervised Learning (SSL) for multi-class medical image classification, since an overly strict threshold limits the utilization of unlabeled data, while an overly loose threshold introduces low-quality pseudo-labels into the training process. This study applies the FixMatch method with a DenseNet-169 architecture as the backbone to classify three lung disease classes COVID-19, Pneumonia, and Normal under conditions of highly limited labeled data. The dataset used is the Covid19, Pneumonia, and Normal Chest X-Ray Images dataset from Mendeley Data, comprising 5,218 images, split into 70% training, 10% validation, and 20% testing. The experiment was systematically designed using three proportions of labeled data (5%, 10%, 15%) and three confidence threshold values (τ = 0.90; 0.95; 0.99), resulting in nine experimental scenarios. The results show that τ = 0.95 with 15% labeled data achieved the best performance (accuracy 97.41%, F1-Score 97.49%, AUC 0.9963) because it balances pseudo-label selectivity with a sufficient effective data volume: at a low label ratio (5%), the limited volume of labeled data means the lower mask rate at τ = 0.95 is not adequately compensated, so τ = 0.99 performs slightly better; whereas at a high label ratio (15%), the selectivity of τ = 0.95 produces high-quality pseudo-labels with adequate volume, driving improved generalization. This study contributes an empirical analysis of confidence threshold sensitivity in FixMatch for multi-class CXR classification with limited labeled data. These findings reveal that the effectiveness of the confidence threshold is contextual to label availability, and determining the optimal threshold cannot be separated from the available labeled data ratio.
Perbandingan Kinerja Random forest dan SVM Pada Klasifikasi Tingkat Kekumuhan Permukiman Menggunakan SMOTE Nurika Dwi Wahyuni; Fadhilah Syafria; Novi Yanti; Surya Agustian
Building of Informatics, Technology and Science (BITS) Vol 8 No 1 (2026): June 2026
Publisher : Forum Kerjasama Pendidikan Tinggi

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47065/bits.v8i1.10101

Abstract

Classifying slum levels is essential for a structured, data-driven analysis of settlement conditions. This study compares the performance of Random forest and Support vector machine (SVM) in classifying slum levels in Pekanbaru City across two scenarios with and without SMOTE using slum indicator scoring data. Its contributions include analyzing SMOTE's impact on model performance and evaluating the top 10 features against the full feature set. The dataset comprises 992 RT-level records from Disperkim Pekanbaru City (2020, 2021, and 2023) featuring 16 slum indicator scores based on PUPR Ministerial Regulation No. 14/2018, categorized into three classes: Non-Slum, Low Slum, and Moderate Slum. Following the KDD process (selection, preprocessing, transformation, data mining, evaluation, and analysis), the data was split 80:20 using stratified sampling and evaluated based on accuracy, precision, recall, F1-score, and confusion matrix. Results show that the Linear SVM without SMOTE achieved perfect evaluation metrics (1.0000); however, this is interpreted cautiously as the class labels derive from strict regulatory scoring rules, making class boundaries inherently linear. Random forest saw its F1-score rise from 0.9660 to 0.9700 after SMOTE, while the most significant improvement occurred in SVM RBF, jumping from 0.9214 to 0.9779. Testing the top 10 features led to a decreased F1-score across models, indicating that utilizing all 16 features remains optimal for this dataset.
Co-Authors .Safrizal, Safrizal Abdillah, Rahmad Achmad Yamin Harahap Afdhal Zikri Afriyanti, Liza Aftari, Dhea Putri AGUNG SUCIPTO Ahmad Kurniawan Ahmad, Rizmah Zakiah Nur Alfitra Salam Arasy, Abdurrahman Arif Kurniawan Ash Shiddicky Aulia Ramadhani Ayu Fransiska Baehaqi Dermawan, Jozu Dzaky Abdillah Salafy Eka Pandu Cynthia El Saputra, Yoga Elin Haerani Elvia Budianita Fadhilah Syafria Faizah Husniah Fauzan Ray T Fauzi Ihsan Febi Yanto Febrian Rizki Adi Sutiyo Fitra Kurnia Fitri Insani Fitri Insani Fitri, Dina Deswara Fuji Astuti Gusti, Siska Kurnia Habib Hakim Sinaga Hadi, Mukhlis Halimah Heru Wibowo Idhafi, Zaky Iffa, Marwika Rifattul Ihsan, Miftahul Iis Afrianty Iis Afrianty Ikhwan Habibi Ilham Habibi Hasibuan Illahi, Ridho Iman Fauzi Aditya Sayogo Indri Pangestuti Iwan Iskandar Jasril Jasril Jasril Jasril Jasril Jasril Jauhari, Najwa Lestari Handayani Lubis, Anggun Tri Utami BR. M Ridho Saputra Marsha Cahyani Dwisyakilla Melyana Hasibuan Miftah Farid Muhammad Affandes Muhammad Affandes Muhammad Elfarizi Muhammad Fikry Muhammad Fikry Muhammad Iqbal Maulana Muhammad Irsyad Muhammad Irsyad Muhammad Ravil Muhammad Tirta Syakban Muktar Sahbuddin Mukti M Kusairi Mulyadi, Syahrul Nadila Handayani Putri naldi, Afri Nazir, Alwis Nazruddin Safaat Nazruddin Safaat H Nazruddin Safaat H Nazruddin Safaat H Negara, Benny Sukma Novi Yanti Novriyanto Novriyanto Novriyanto Novriyanto Nurika Dwi Wahyuni Nurul Fatiara Okfalisa Okfalisa Oktavia, Lola Pangestu, Yoga Pizaini Pizaini Pranata, Joni Prima Yohana Putri Zahwa Putri, Adilah Atikah Putri, Atika Rahmad Abdillah Rahmad Kurniawan Ramadhani, Siti Reski Mai Candra Reski Mai Candra Rizqa Raaiqa Bintana Safrizal, Afri Naldi Salam Kurniawan Saputra, Ikhsan Dwi Saputra, Nugroho Wahyu Satira, Husna Sinaga, Habib Hakim Siti Ramadhani Siti Ramadhani Sri Puji Utami A. Subhi, Yazid Abdullah Suci Rahayu Sulistia Ningsih, Sulistia Suwanto Sanjaya Syaiful Azhar Tarmizi, Veci Cahyono Teddie D Trya Ayu Pratiwi Utari, Roid Fitrah Wahyu Reinaldy Wan Sobri Amin Yusra Yusra Yusra Yusra Yusra, Yusra