Claim Missing Document
Check
Articles

Addressing Extreme Class Imbalance in Multilingual Complaint Classification Using XLM-RoBERTa Ariyanto, Muhammad; Alzami, Farrikh; Sani, Ramadhan Rakhmat; Gamayanto, Indra; Naufal, Muhammad; Winarno, Sri; Iswahyudi
Journal of Applied Informatics and Computing Vol. 10 No. 1 (2026): February 2026
Publisher : Politeknik Negeri Batam

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30871/jaic.v10i1.11606

Abstract

Government complaint management systems often suffer from extreme class imbalance, where a few public service categories accumulate most reports while many others remain under-represented. This research examines whether simple class weighting can improve fairness in multilingual transformer models for automatic routing of Indonesian citizen complaints on the LaporGub Central Java e-governance platform. The dataset comprises 53,877 Indonesian-language complaints spanning 18 service categories with an imbalance ratio of about 227:1 between the largest and smallest classes. After cleaning and deduplication, we stratify the data into training, validation, and test sets. We compare three approaches: (i) a linear support vector machine (SVM) with term frequency inverse document frequency (TF-IDF) unigram and bigram and class-balanced weights, (ii) a cross-lingual RoBERTa (XLM-RoBERTa-base) model without class weighting, and (iii) an XLM-RoBERTa-base model with a class-weighted cross-entropy loss. Fairness is operationalised as equal importance for categories and quantified primarily using the macro-averaged F1-score (Macro-F1), complemented by per-class F1, weighted F1, and accuracy. The unweighted XLM-RoBERTa model outperforms the SVM baseline in Macro-F1 (0.610 vs 0.561). The class-weighted variant attains similar Macro-F1 (0.608) while redistributing performance towards minority categories. Analysis shows that class weighting is most beneficial for categories with a few hundred to several thousand samples, whereas extremely rare categories with fewer than 200 complaints remain difficult for all models and require additional data-centric interventions. These findings demonstrate that multilingual transformer architectures combined with simple class weighting can provide a more balanced backbone for automated complaint routing in Indonesian e-government, particularly for low- and medium-frequency service categories.
Exploring Public Opinion on the 'Makan Bergizi Gratis' Program on X: A Comparative Analysis of IndoBERT-Large and NusaBERT-Large Models Arunia, Aurelya Prameswari; Sani, Ramadhan Rakhmat; Dewi, Ika Novita; Sulistyono, MY Teguh
Journal of Applied Informatics and Computing Vol. 10 No. 1 (2026): February 2026
Publisher : Politeknik Negeri Batam

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30871/jaic.v10i1.11757

Abstract

Program Makan Bergizi Gratis (MBG) has triggered extensive discourse on social media platform X, which serves as a primary space for public expression of opinions toward government policies. This study aims to analyze public sentiment toward the MBG program while simultaneously comparing the performance of two prominent Transformer-based models, namely IndoBERT-Large and NusaBERT-Large. This research adopts a quantitative approach employing supervised learning on 10,201 Indonesian-language posts (tweets) collected through web scraping from February 2024 to September 2025. A total of 2,000 samples were manually annotated as ground truth, achieving a high level of inter-annotator reliability (Cohen’s Kappa, κ = 0.81). The experimental results indicate that IndoBERT-Large outperforms NusaBERT-Large, achieving an accuracy of 83.00%, while NusaBERT-Large demonstrates competitive performance with an accuracy of 80.50%. Substantively, public discourse is dominated by negative sentiment, accounting for nearly 50% of the total data, reflecting public concerns regarding budgetary constraints and technical implementation issues. Positive sentiment ranges between 33% and 36%, indicating sustained and substantial public support for the program. These findings confirm the effectiveness of Transformer-based models in accurately capturing the dynamics of public opinion toward government policies using social media data.
IMPLEMENTASI METODE DESIGN THINKING DAN SYSTEM USABILITY SCALE PADA USER EXPERIENCE APLIKASI BELAJAR BAHASA INGGRIS TALKTALES MELALUI CERITA RAKYAT Muhammad Nabhan Rifa’i; Ramadhan Rakhmat Sani; Suharnawi Suharnawi; Resha Meiranadi Caturkusuma
Jurnal Sistem Informasi dan Informatika (Simika) Vol. 8 No. 1 (2025): Jurnal Sistem Informasi dan Informatika (Simika)
Publisher : Program Studi Sistem Informasi, Universitas Banten Jaya

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.47080/simika.v8i1.3742

Abstract

Indonesia faces significant challenges in improving English proficiency among its population. According to the EF Education First English Proficiency Index 2023, Indonesia ranks 79th out of 113 countries. On the other hand, the current generation begins to forget cultural elements such as folklore or myths that have been passed down from the nation's ancestors. This study aims to design an English learning application for children and teenagers using the design thinking method. The stages of design thinking that are used are empathize, define, ideate, prototyping, and testing. In the prototyping stage, low-fidelity and high-fidelity prototypes were created to visualize the application's design and functionality. Testing was conducted by using task scenarios and the System Usability Scale (SUS). The task scenario testing results revealed that the effectiveness and efficiency rate of 85.71%, indicated that most tasks could be completed successfully by users. The SUS testing results showed an average score of 86.5%, indicated that the application has a high level of usability and well-received by users. Thus, the application's interface is considered easy to use and effective in supporting the English learning process for the target users. This research provides a positive contribution to the development of educational applications using a design thinking approach.
PENDEKATAN EXPLAINABLE MACHINE LEARNING UNTUK ANALISIS FAKTOR DROP OUT MAHASISWA MENGGUNAKAN XGBOOST Agnes Putri Istiwana; Ramadhan Rakhmat Sani; Yuventius Tyas Catur Pramudi
Rabit : Jurnal Teknologi dan Sistem Informasi Univrab Vol 11 No 1 (2026): Januari
Publisher : LPPM Universitas Abdurrab

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.36341/rabit.v11i1.7218

Abstract

The problem of student dropout is a strategic issue in higher education because it has a direct impact on academic quality, institutional efficiency, and university accreditation. This study aims to statistically analyze the factors that contribute to the variation in GPA of students who have dropped out using the Explainable Machine Learning approach. The predictive model was built using the Extreme Gradient Boosting (XGBoost) algorithm to obtain optimal prediction performance, while the Shapley Additive Explanations (SHAP) method was used to provide interpretation of the contribution of each feature in the model. The research dataset includes academic, administrative, and demographic data of students who have dropped out in the last two academic years. The evaluation results show that the XGBoost model shows excellent predictive performance with an R² value of 0.820 indicating that most of the GPA variation can be explained by the model, and is supported by a Root Mean Squared Error (RMSE) value of 0.344 and a Mean Absolute Error (MAE) of 0.172 indicating that the prediction error rate is relatively low. SHAP analysis revealed that the number of credits taken and tuition payment status were the two factors that statistically significantly contributed to GPA changes in the predictive model. This study provides more comprehensive insights by combining high predictive performance and model interpretability, enabling educational institutions to identify student academic risk earlier and based on data.
Penerapan Arsitektur Enterprise pada Pelayanan Pendaftaran Anggota pada Fitnation Premiere Gym Indra Gamayanto; Galih Mentari Pangesti; Ramadhan Rakhmat Sani
JOINS (Journal of Information System) Vol 7 No 2 (2022): Edisi November 2022
Publisher : Fakultas Ilmu Komputer, Universitas Dian Nuswantoro

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.33633/joins.v7i2.6396

Abstract

Fitnation Premiere Gym merupakan usaha yang bergerak di bidang jasa layanan fitness center kepada masyarakat yang ingin berolahraga dalam menjaga kebugaran tubuh dan pembentukan tubuh yang ideal. Untuk meningkatkan kualitas manajemen dan pelayanan, perlu adanya sistem informasi yang membantu dalam proses manajemen dan kinerja. Aktivitas bisnis utamanya adalah pendaftaran anggota, sistem pembayaran dan paket program mereka masih menggunakan sistem, database, aplikasi dan teknologi yang belum saling terintegrasi. Pengelolaan data, verifikasi pembayaran, pelaporan dan persetujuan kegiatan masih dilakukan secara manual. Dalam hal ini memiliki banyak resiko seperti kesalahan dalam pencatatan, waktu yang dibutuhkan untuk pendaftaran anggota relatif menjadi lebih lama, dan informasi yang dibutuhkan tidak terintegrasi. Dalam perencanaan strategis IS/TI sangat diperlukan suatu enterprise architecture agar dapat tercapai keselarasan strategi IS/TI dengan strategi bisnis dari organisasi untuk mengatasi permasalahan guna mewujudkan visi, misi, dan tujuan organisasi. Penelitian ini menggunakan metodologi TOGAF ADM dimulai dari preliminary phase, requirement management, architecture vision, business architecture, information system architecture, technology architecture, opportunities and solution, hingga migration planning. Hasil dari penelitian ini, yaitu suatu usulan model TOGAF ADM yang disesuaikan dengan proses dan kebutuhan bisnis dari fitness center dalam merancang Enterprise Architecture untuk perencanaan strategis IS/TI.
Deep Learning Factor Investing in the Indonesian Stock Market Atha Rohmatullah, Fawwaz; Alzami, Farrikh; Rakhmat Sani, Ramadhan; Novita Dewi, Ika; Winarno, Sri; Sulistyono, Teguh
IJCCS (Indonesian Journal of Computing and Cybernetics Systems) Vol 19, No 4 (2025): October
Publisher : IndoCEISS in colaboration with Universitas Gadjah Mada, Indonesia.

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.22146/ijccs.112549

Abstract

Traditional linear factor models often fail to capture the complex, non-linear dynamics of emerging stock markets. This research designs and validates a novel Recurrence Plot (RP) matrices with β-VAE deep learning methodology to discover non-linear investment factors within the Indonesian context. We demonstrate that this framework is a systematically superior "factor factory" compared to a linear RP with PCA baseline, discovering twice as many high-quality factors (Sharpe > 0.3) and generating 7-fold more alpha on average. A key finding is the model's ability to disentangle high-frequency predictive signals (identified by SHAP) from more valuable, low-frequency profitable trends (validated by backtesting). The champion factor from this process yields a robust annualized alpha of 6.65% with a minimal max drawdown of -7.73% from 2018 to 2025. This study concludes that the RP -> β -VAE approach is a robust and resilient framework for discovering safer, non-linear sources of return unexplained by conventional models.
Perbandingan Metode Seleksi Fitur Chi-Square dan Information Gain untuk Peningkatan Interpretabilitas dan Optimasi Kinerja Model TabNet Annisa Ratna Salsabilla; Ramadhan Rakhmat Sani; Ika Novita Dewi
Jurnal Nasional Teknologi dan Sistem Informasi Vol 11 No 3 (2025): Desember 2025
Publisher : Departemen Sistem Informasi, Fakultas Teknologi Informasi, Universitas Andalas

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.25077/TEKNOSI.v11i3.2025.253-262

Abstract

Breast cancer is one of the most significant global health issues. Machine learning approaches offer the potential to accurately analyze clinical data and aid in early diagnosis. However, conventional machine learning models are often limited in their ability to model complex nonlinear relationships in medical data, which can reduce predictive accuracy. This study employs a deep learning architecture because of its ability to model such relationships. Specifically, the TabNet model was chosen because it is designed for tabular data and offers better interpretability. The public Wisconsin Diagnostic Breast Cancer (WDBC) dataset, which has 30 features and an imbalanced class distribution, was used in this study. Feature selection was necessary to handle the high-dimensional data, and SMOTE-ENN was used for class balancing. Two feature selection methods, Chi-Square and Information Gain, were compared to determine the most effective approach. Hyperparameter optimization was performed using Optuna and validated with stratified k-fold cross-validation to ensure optimal performance. The results of the experiment demonstrate that feature selection and optimization significantly improve performance. The base model with Chi-Square feature selection achieved an accuracy rate of 64.91%. Meanwhile, the Chi-Square model with Optuna optimization increased accuracy to 98.25%. This is 3.51% higher than the accuracy of 94.74% achieved by the optimized model without feature selection. In the final comparison, both methods demonstrated distinct advantages: Chi-Square (75% features) excelled in achieving 100% precision and more efficient computation time. Information Gain (75% features), on the other hand, was the only method to achieve 100% recall, which is crucial for minimizing false negatives. These results demonstrate that the optimal method depends on the context. Information Gain is best for maximum diagnostic sensitivity, and Chi-Square is best for performance balance and efficiency.
Spatiotemporal Analysis of Peatland Fire Hotspots and Fire Intensity in Riau Province Using MODIS–VIIRS Multisensor Satellite Data Najwa Ratu Afi; Ramadhan Rakhmat Sani; Ricardus Anggi Pramunendar; Nurul Anisa Sri Winarsih; Ika Novita Dewi
Journal of Applied Informatics and Computing Vol. 10 No. 3 (2026): June 2026
Publisher : Politeknik Negeri Batam

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30871/jaic.v10i3.12686

Abstract

Peatland fires in Riau Province frequently occur during the dry season and contribute significantly to regional haze, environmental degradation and carbon emissions. Effective monitoring of these fires remains challenging due to their widespread distribution and varying intensity across peatland areas. This research aims to analyze the spatiotemporal characteristics of peatland fire hotspots in Riau Province using multisensor satellite observations from the NASA Fire Information for Resource Management System (FIRMS). The dataset integrates Moderate Resolution Imaging Spectroradiometer (MODIS) and Visible Infrared Imaging Radiometer Suite (VIIRS) data from the Suomi-NPP, NOAA-20 and NOAA-21 satellites. After applying filtering criteria of confidence ≥70% and Fire Radiative Power (FRP) ≥5 megawatts (MW), a total of 7,297 significant hotspots were identified during the July–October 2025 dry season. The results show that fire activity peaked in July with a maximum daily FRP of 25,611 MW and a monthly total of 65,120 MW, followed by a decline in September and a slight increase in October. The FRP distribution was highly right-skewed, with an average value of13.2 MW, while the most intense hotspots reached 189.4 MW. Estimated carbon dioxide (CO₂) emissions reached approximately 122,472 tons, indicating substantial environmental impacts. Spatial clustering and persistence analysis revealed several high-risk peatland zones with repeated fire occurrences. These findings demonstrate the importance of multisensor satellite monitoring for improving early fire detection, emission assessment and disaster mitigation strategies in peatland regions.
Model Hybrid Random Forest dan Information Gain untuk meningkatkan Performa Algoritma Machine Learning pada Deteksi Malicious Software Fauzi Adi Rafrastara; Wildanil Ghozi; Ramadhan Rakhmat Sani; L. Budi Handoko
Jurnal Informatika dan Rekayasa Perangkat Lunak Vol. 6 No. 2 (2024): September
Publisher : Universitas Wahid Hasyim

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

The evolution of malware, or malicious software, has raised increasing concerns, targeting not only computers but also other devices like smartphones. Malware is no longer just monomorphic but has evolved into polymorphic, metamorphic, and oligomorphic forms. With this massive development, conventional antivirus software is becoming less effective at countering it. This is due to malware's ability to propagate itself using different fingerprint and behavioral patterns. Therefore, an intelligent machine learning-based antivirus is needed, capable of detecting malware based on behavior rather than fingerprints. This research focuses on the implementation of a machine learning model for malware detection using ensemble algorithms and feature selection to achieve optimal performance. The ensemble algorithm used is Random Forest, evaluated and compared with k-Nearest Neighbor and Decision Tree as state-of-the-art methods. To enhance classification performance in terms of processing speed, the feature selection method applied is Information Gain, with 22 features. The highest results were achieved using the Random Forest algorithm and Information Gain feature selection method, reaching a score of 99.0% for accuracy and F1-Score. By reducing the number of features, processing speed can be increased by almost fivefold.
Deteksi Serangan Denial of Service (DoS) dan Spoofing pada Internet of Vehicles menggunakan Algoritma K-Nearest Neighbor (KNN) Wildanil Ghozi; Fauzi Adi Rafrastara; Ramadhan Rakhmat Sani; Abdussalam Abdussalam
Jurnal Informatika dan Rekayasa Perangkat Lunak Vol. 6 No. 2 (2024): September
Publisher : Universitas Wahid Hasyim

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

The implementation of Internet of Things (IoT) technology in motor vehicles has been increasing over time and is known as the Internet of Vehicles (IoV). IoV is becoming more essential to society as it provides comfort, safety, and efficiency in driving. Unfortunately, the use of internet technology in IoV brings the potential for cyber-attacks, such as Denial of Service (DoS) and Spoofing. Intrusion Detection Systems in IoV have not yet fully matured, as this technology is still relatively new. Therefore, the potential threats and their significant impact make research on this topic urgently needed. This study aims to evaluate the performance of the k-Nearest Neighbor (kNN) classification algorithm in detecting cyber-attacks on IoV. The predicted classes in this study consist of six categories: Benign, DoS, Gas-Spoofing, Steering Wheel-Spoofing, Speed-Spoofing, and RPM-Spoofing. These two types of attacks on IoV (DoS and Spoofing) pose risks to the operational safety of vehicles, which can endanger drivers and other road users. The dataset used is a public dataset called CIC IoV2024. The performance of the kNN algorithm is also compared to three other state-of-the-art algorithms, including Naïve Bayes, Deep Neural Network, and Random Forest. The results show that k-Nearest Neighbor (kNN) achieved the best performance with a score of 98.7% for both accuracy and F1-Score metrics. kNN outperformed Naïve Bayes, which ranked second with a score of 98.1% accuracy and 98.0% F1-Score. Thus, the kNN algorithm can be recommended as a classifier in the development of an intrusion detection system for IoV
Co-Authors ., Junta Zeniarza ., Junta Zeniarza Abdussalam Abdussalam Abdussalam Abdussalam, Abdussalam Abu Salam Ade Nurul Aisyah Agnes Putri Istiwana Agung Priyo Utomo, Rino Ahmad Khotibul Umam, Ahmad Khotibul Al zami, Farrikh Alzami, Farrikh Annisa Ratna Salsabilla Ardytha Luthfiarta ARIYANTO, MUHAMMAD Arta Moro Sundjaja, Arta Moro Arunia, Aurelya Prameswari Asih Rohmani Asih Rohmani Asih Rohmani, Asih Atha Rohmatullah, Fawwaz Bernadette Chayeenee Norman , Maria Budi Harjo Budi, Setyo Candra Irawan Catur Supriyanto Christy Atika Sari Darnell Ignasius Defri Kurniawan Defri Kurniawan Diana Aqmala Doheir, Mohamed Dwi Puji Prabowo, Dwi Puji Eko Hari Rachmawanto Elkaf Rahmawan Pramudya Erika Devi Udayanti Fahmi Amiq Farah Syadza Mufidah Farrikh Al Zami Farrikh Al Zami Fauzi Adi Rafrastara Fauzi Adi Rafrastara Florentina Esti Nilasari Florentina Esti Nilawati Galih Mentari Pangesti Guruh Fajar Shidik Hanny Haryanto Harun Al Azies Hercio Venceslau Silla Heru Lestiawan Hussein, Jasim Nadheer Hussein, Jassim Nadheer Ifan Rizqa Ignasius, Darnell Ika Novita Dewi Ika Novita Dewi Ikhwansyah Kurniawan Indra Gamayanto Iswahyudi ISWAHYUDI ISWAHYUDI Ivan Bayu Fachreza Junta Zeniarja Karin, Tan Regina Kiki Widia Kurniawan, Defri L. Budi Handoko Lekso Budi Handoko Maszuda, Akbar Alvian Megantara, Rama Aria Melati Anggreni Sitorus Muhammad Fais Ramadhani Muhammad Nabhan Rifa’i Muhammad Naufal MY. Teguh Sulistyono Nadya Azizah Najwa Ratu Afi Nida Aulia Karima Novita Dewi , Ika Nugraha, Purwa Esti Paramita, Cinantya Pergiwati, Dewi Priyo Utomo, Rino Agung Pulung Nurtantio Andono Purwanto Purwanto Ramadhani, Dwi Arya Resha Meiranadi Caturkusuma Rhyan David Levandra Ricardus Anggi Pramunendar Richard Emmerig S. Sukamto, Titien Sarker, Md. Kamruzzaman Sasono Wibowo Sendi Novianto Sendi Novianto Sendi Novianto Setyo Budi Setyo Budi Sirait, Tamsir Hasudungan Soares, Gilardinho Javiere Oscoraldo Pedrosa Sri Winarno Sri Winarno Suharnawi Suharnawi Suharnawi Suharnawi Suharnawi Sukamto, Titien S. Sukamto, Titien Suhartini Sulistyono, Teguh Syahrizal, Muhammad Iqbal Titien Suhartini Sukamto Titien Suhartini Sukamto Utomo, Danang Wahyu Wibowo, Isro' Rizky Wildanil Ghozi Winarsih, Nurul Anisa Sri Wulan Puspita Loka Yani Parti Astuti Yanuaresta, Dianna Yunita Ayu Pratiwi Yupie Kusumawati Yuventius Tyas Catur Pramudi Zahro, Azzula Cerliana Zami, Farrikh Al