Claim Missing Document
Check
Articles

COMPARATIVE ANALYSIS OF BAGGING AND BOOSTING MODELS IN ENSEMBLE LEARNING FOR GRADUATION PREDICTION Sartika Lina Mulani Sitio; Darmawati; Yuda Samudra
JITK (Jurnal Ilmu Pengetahuan dan Teknologi Komputer) Vol. 11 No. 3 (2026): JITK Issue February 2026
Publisher : LPPM Nusa Mandiri

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.33480/jitk.v11i3.7579

Abstract

Student graduation prediction is an important aspect in supporting academic decision-making in higher education. However, conventional evaluation approaches have not been able to identify the risk of early graduation delays. This study aims to compare the performance of two ensemble learning approaches, namely Bagging using Random Forest and Boosting using XGBoost, in predicting student graduation. The study used  the Predict Students' Dropout and Academic Success dataset  consisting of 4,424 student data. Both models were trained on the same data and evaluated using the Accuracy, Precision, Recall, F1-Score, and ROC-AUC metrics. The results of the experiment showed that both models had almost equal accuracy, i.e. 82.6% for Random Forest and 82.5% for XGBoost. However, XGBoost showed better performance on Recall (0.878) and F1-Score (0.834), which indicated a higher ability to detect students who actually graduated. Based on these results, this study concludes that XGBoost is more effective than Random Forest in the context of predicting student graduation and is more suitable to be applied to  the Academic Early Warning System in universities
PENERAPAN ALGORITMA K-MEANS CLUSTERING UNTUK ANALISIS POLA DATA EKONOMI HISTORIS Abed Neco; Firman Aziz Saputra; Nazar Fadhil Abdullah; Rizky Ramadhani; Testarina Tatiana Hermansyah; Sartika Lina Mulani Sitio
JRIS : Jurnal Rekayasa Informasi Swadharma Vol 5, No 2 (2025): JURNAL JRIS EDISI JULI 2025
Publisher : Institut Teknologi dan Bisnis (ITB) Swadharma

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.56486/jris.vol5no2.879

Abstract

Historical economic and financial data are available in vast volumes, yet extracting non-trivial insights hidden within them remains a significant challenge, primarily due to the reliance on traditional, hypothesis-driven analysis methods. In the Indonesian context, the comprehensive application of clustering techniques to uncover objective data narratives remains unexplored, mainly raising the urgency of developing a data-driven approach. This study aims to address this gap by demonstrating the capabilities and flexibility of the K-Means algorithm as a robust exploratory analysis method. The study employs a comparative case study approach on five independent datasets purposefully selected to cover diverse domains and periods: bank merger trends (1971–1988), critical macroeconomic indicators (1992–2003), state-owned bank financial performance (2004–2014), bird’s nest exports (2017–2021), and comparable economic data from the United States (1930–1955). Methodologically, each dataset was rigorously pre-processed before being clustered using the K-Means algorithm, with the quality of the results quantitatively evaluated using the Silhouette Score, Davies-Bouldin Index, and Inertia metrics. The results demonstrate powerful clustering performance, with three of the five case studies achieving Silhouette Scores above 0.70, indicating dense and well-defined data segmentation. Key findings demonstrate that the formed clusters successfully map historical periods objectively; for example, the algorithm automatically isolates the extreme anomaly of the 1998 monetary crisis as a unique cluster, identifies the peak of the banking merger era as a phase of intense consolidation, and groups state-owned banks into distinct strategic segments based on their capital and profitability profiles. This study confirms that K-Means is an effective exploratory analysis method, capable of transforming complex historical data into structured insights to support more informed and evidence-based policy formulation.Meskipun data ekonomi dan keuangan historis tersedia dalam volume yang sangat besar, upaya untuk mengekstrak wawasan non-trivial yang tersembunyi di dalamnya tetap menjadi tantangan signifikan, terutama karena ketergantungan pada metode analisis tradisional yang bersifat hypothesis-driven. Dalam konteks Indonesia, aplikasi teknik Clustering secara komprehensif untuk mengungkap narasi data yang objektif masih belum banyak dieksplorasi, sehingga memunculkan urgensi untuk mengembangkan pendekatan berbasis data. Penelitian ini bertujuan untuk menjawab kesenjangan tersebut dengan mendemonstrasikan kapabilitas dan fleksibilitas algoritma K-Means sebagai metode analisis eksplorasi yang tangguh. Untuk mencapai tujuan ini, penelitian menerapkan pendekatan studi kasus komparatif pada lima dataset independen yang sengaja dipilih guna mencakup domain dan periode waktu yang beragam: tren merger bank (1971-1988), indikator makroekonomi kritis (1992-2003), kinerja keuangan bank BUMN (2004-2014), ekspor komoditas sarang burung walet (2017-2021), dan data ekonomi pembanding dari Amerika Serikat (1930-1955). Secara metodologis, setiap dataset diproses secara ketat melalui pra-pemrosesan sebelum dikelompokkan menggunakan K-Means, dengan kualitas hasil dievaluasi secara kuantitatif melalui metrik Silhouette score, Davies-Bouldin Index, dan Inertia. Hasil penelitian menunjukkan kinerja klasterisasi yang sangat kuat, di mana tiga dari lima studi kasus mencapai Silhouette score di atas 0.70, yang mengindikasikan segmentasi data yang padat dan terdefinisi dengan baik. Temuan utama menunjukkan bahwa klaster yang terbentuk berhasil memetakan periode-periode historis secara objektif; sebagai contoh, algoritma ini secara otomatis mengisolasi anomali ekstrem krisis moneter 1998 sebagai sebuah klaster unik, mengidentifikasi puncak era merger perbankan sebagai fase konsolidasi yang intens, serta mengelompokkan bank-bank BUMN ke dalam segmen strategis yang berbeda berdasarkan profil modal dan profitabilitasnya. Studi ini mengonfirmasi bahwa K-Means adalah metode analisis eksplorasi yang efektif, mampu mentransformasi data historis yang kompleks menjadi wawasan terstruktur untuk mendukung perumusan kebijakan yang lebih informatif dan berbasis bukti.
KLASTERISASI DATA : ANALISIS KINERJA K-MEANS PADA SEKTOR PAJAK, EKSPOR, PERIKANAN, MODAL, DAN SUMBER DAYA Achmad S.W.A Nurba; Dharma Fathahillah; Muhamad Shafly Pratama; Muhammad Rivaldi Bachtiar; Muhammad Chesta Adabi Putra; Sartika Lina Mulani Sitio
JRIS : Jurnal Rekayasa Informasi Swadharma Vol 5, No 2 (2025): JURNAL JRIS EDISI JULI 2025
Publisher : Institut Teknologi dan Bisnis (ITB) Swadharma

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.56486/jris.vol5no2.883

Abstract

This study aims to cluster Indonesian economic data patterns from five sectors: tax, export, fisheries, capital markets, and resources, using the K-Means algorithm. Data were obtained from BPS, the Ministry of Finance, the Ministry of Marine Affairs and Fisheries, the Financial Services Authority (OJK), and UN Comtrade. Pre-processing was carried out through data cleaning and normalization. The optimal number of clusters was determined using the elbow and silhouette methods. Clustering evaluation used Inertia, Silhouette score, and the Davies-Bouldin Index. The results show variations in cluster patterns in each sector, with the fisheries and capital markets sectors providing the best results (high silhouette scores). Visualization using PCA supports cluster interpretation. These findings demonstrate that K-Means is effective in economic data analysis and helps support more adaptive and data-driven policies.Penelitian ini bertujuan mengelompokkan pola data ekonomi Indonesia dari lima sektor: pajak, ekspor, perikanan, pasar modal, dan sumber daya, menggunakan algoritma K-Means. Data diperoleh dari BPS, Kementerian Keuangan, KKP, OJK, dan UN Comtrade. Pra-pemrosesan dilakukan melalui pembersihan dan normalisasi data. Jumlah klaster optimal ditentukan menggunakan metode elbow dan silhouette. Evaluasi klasterisasi menggunakan Inertia, Silhouette score, dan Davies-Bouldin Index. Hasil menunjukkan variasi pola klaster di tiap sektor, dengan sektor perikanan dan pasar modal memberikan hasil terbaik (silhouette score tinggi). Visualisasi menggunakan PCA mendukung interpretasi klaster. Temuan ini menunjukkan bahwa K-Means efektif dalam analisis data ekonomi dan bermanfaat untuk mendukung kebijakan yang lebih adaptif dan berbasis data.
Comparison of Logistic Regression and Random Forest Performance in Student Dropout Prediction based on Multi Source Data Sartika Lina Mulani Sitio; Sunardi; Abdul Fadlil
JURIKOM (Jurnal Riset Komputer) Vol. 13 No. 3 (2026): Juni 2026
Publisher : Universitas Budi Darma

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

The high rate of student dropouts is one of the important challenges in higher education because it can affect academic quality, learning effectiveness, and the performance of educational institutions. This condition encourages the need for a prediction system that is able to identify students at risk of dropouts early so that preventive measures can be taken appropriately. This study aims to compare the performance of Logistic Regression and Random Forest algorithms in predicting student dropout based on multi-source. The dataset consists of 4,424 student data with 34 attributes covering academic, demographic, socioeconomic, and academic administration aspects. The research stages include data preprocessing, target transformation into binary classification, feature scaling, data sharing using an 80:20 scheme, and handling class imbalances using the Synthetic Minority Oversampling Technique (SMOTE). Furthermore, a modeling process was carried out using Logistic Regression and Random Forest algorithms to predict the risk of student dropout. Model evaluation was carried out using accuracy, precision, recall, F1-score, and Area Under Curve Receiver Operating Characteristic (AUC-ROC). The results showed that Random Forest performed better than Logistic Regression with an accuracy of 0.884, precision of 0.842, recall of 0.785, F1-score of 0.812, and AUC-ROC of 0.930. Meanwhile, Logistic Regression obtained an accuracy of 0.871, precision of 0.780, recall of 0.835, F1-score of 0.806, and AUC-ROC of 0.928. These results show that Random Forest is more effective in handling complex relationships in multi-source data for student dropout predictions
IMPLEMENTASI DATA MINING PREDIKSI KELULUSAN SISWA MENGGUNAKAN METODE DECISION TREE PADA SMK IPTEK TANGSEL Suryaningrat Suryaningrat; Sartika Lina Mulani Sitio; Ayni Suwarno Herry
Riau Jurnal Teknik Informatika Vol. 4 No. 1 (2025): Maret 2025
Publisher : Prodi Teknik Informatika Universitas Pasir Pengaraian

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30606/rjti.v4i1.3267

Abstract

Vocational High School (SMK) is a formal educational institution that prepares its graduates for the world of work. However, SMK Iptek Tangsel faces challenges in optimizing student data and overcoming labor shortages in the field of education administration. To overcome this problem, this study applies a prediction system using a data mining technique with the decision tree method. Aims to improve the accuracy of student graduation predictions. The methodology of this study adopts a quantitative approach with structured steps, including observation, interviews, data collection, and documentation. The results showed that the accuracy of Rapid miners in motorcycle and multimedia business engineering reached 98.49%, while in hospitality and accounting accommodation reached 99.05%. This graduation prediction system helps schools identify students at risk of not graduating and who are at potential to graduate, optimize data management, and enable appropriate interventions. This research makes a significant contribution to improving the decision-making process in the field of education through the use of data mining technology. Thus, the graduation prediction system can be an effective tool in supporting data management and increasing student graduation in vocational schools.
Rancang Bangun Alat Otomatis Pengganti dan Pengontrol Air dengan Deteksi Tingkat Kekeruhan dan PH Pada Akuarium Ikan Cupang Sartika Lina Mulani Sitio; Nardiono; Yuda Samudra
Telekontran : Jurnal Ilmiah Telekomunikasi, Kendali dan Elektronika Terapan Vol. 13 No. 2 (2025): TELEKONTRAN vol 13 no 2 Oktober 2025
Publisher : Program Studi Teknik Elektro, Fakultas Teknik dan Ilmu Komputer, Universitas Komputer Indonesia.

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.34010/telekontran.v13i2.16690

Abstract

Water quality that is not optimally maintained can have a negative impact on the health of ornamental fish, especially Bluerim betta fish that require an aquarium environment with a certain level of acidity and turbidity of the water. The problems that are often faced by ornamental fish hobbyists are delays in changing water and difficulties in monitoring water conditions manually. This study aims to design and build an automatic water replacement and control system in betta fish aquariums with the ability to detect turbidity levels and water pH in real-time. The system uses a pH-SEN0161 sensor to measure acidity, a turbidity-SEN0189 sensor to detect turbidity in NTU units, and an SRF05 ultrasonic sensor to measure water level. The software was developed using the Arduino IDE and implemented on the Arduino ATMega2560 microcontroller as well as the NodeMCU ESP8266 for data processing and automatic control. The test was carried out for 30 days with an ideal pH standard between 6–7 and a turbidity value below 400 NTU. The test results show that the system can work optimally in replacing and controlling the water conditions of the Bluerim betta fish aquarium, thus supporting the quality of life of the fish effectively and efficiently.
Comparison of the Performance of Machine Learning Classification Algorithms on Phishing URL Detection Ahmad; Sartika Lina Mulani Sitio
Jurnal Inotera Vol. 11 No. 2 (2026): July - December 2026
Publisher : LPPM Politeknik Aceh Selatan

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.31572/inotera.Vol11.Iss2.2026.ID697

Abstract

The development of internet technology has increased the intensity of digital activities, but it has also been followed by an increase in cybersecurity threats, one of which is phishing attacks through malicious URLs. Phishing is a fraudulent method that is carried out by manipulating users through fake websites that resemble official websites to obtain sensitive information, such as usernames, passwords, and financial data. Conventional blacklist-based detection methods are considered less effective in recognizing new phishing URLs that continue to develop dynamically. Therefore, this study aims to analyze and compare the performance of various machine learning and deep learning algorithms in accurately detecting phishing URLs. The dataset used was obtained from Kaggle with a total of 11,054 data points and 31 features that represent the characteristics of phishing and legitimate URLs. The methods used include Logistic Regression, K-Nearest Neighbor, Support Vector Machine, Naive Bayes, Decision Tree, Random Forest, Gradient Boosting, CatBoost, Extreme Gradient Boosting, and Multilayer Perceptron. The research stages include data preprocessing, exploratory data analysis, data visualization, separation of training and testing data, model training, and performance evaluation using accuracy, precision, recall, and F1-score. The results showed that the Gradient Boosting algorithm provided the best performance with an accuracy of 0.974, an F1-score of 0.977, a recall of 0.994, and a precision of 0.986. The results show that the ensemble learning method is able to detect phishing URLs effectively and can be used to improve artificial intelligence-based cybersecurity systems.
Penerapan Metode Fuzzy Tsukamoto dalam Sistem Pendukung Keputusan Penentuan Stok Mainan Berdasarkan Permintaan dan Musim Penjualan di Toko Vinca Mico Ferdian; Sartika Lina Mulani Sitio
Indonesian Journal of Multidisciplinary on Social and Technology Vol. 4 No. 3 (2026): Juli - Oktober
Publisher : PT Ilmu Data Indonesia

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.69693/ijmst.v4i3.13039

Abstract

Pengelolaan persediaan merupakan salah satu aspek penting dalam menjaga kelancaran operasional dan keberlangsungan usaha ritel. Toko Vinca sebagai toko yang bergerak dalam penjualan mainan anak menghadapi permasalahan berupa fluktuasi permintaan yang dipengaruhi oleh musim penjualan, sehingga dapat menyebabkan kelebihan maupun kekurangan stok. Penelitian ini bertujuan menerapkan metode Fuzzy Tsukamoto dalam Sistem Pendukung Keputusan untuk menentukan rekomendasi jumlah stok mainan berdasarkan tingkat permintaan dan musim penjualan. Data yang digunakan berupa data penjualan harian Toko Vinca pada periode 2023–2024. Metode Fuzzy Tsukamoto digunakan karena mampu mengolah data yang bersifat tidak pasti melalui tahapan fuzzifikasi, inferensi menggunakan aturan IF-THEN, dan defuzzifikasi untuk menghasilkan nilai rekomendasi stok. Sistem dikembangkan dalam bentuk aplikasi berbasis web menggunakan Streamlit yang dilengkapi fitur pengunggahan data, pengaturan musim, pengujian data, pencarian item, visualisasi grafik, serta rekapitulasi dan pengunduhan hasil. Hasil pengujian menunjukkan nilai Mean Absolute Error (MAE) sebesar 13,03, Root Mean Square Error (RMSE) sebesar 23,64, Mean Absolute Percentage Error (MAPE) sebesar 26,71%, dan tingkat akurasi sebesar 73,29%. Hasil tersebut menunjukkan bahwa sistem berada dalam kategori cukup baik dan dapat digunakan sebagai alat bantu pengambilan keputusan dalam menentukan jumlah stok mainan. Penerapan metode Fuzzy Tsukamoto juga dapat membantu toko mengurangi risiko overstock dan understock serta mendukung pengelolaan persediaan yang lebih sistematis dan sesuai dengan kondisi permintaan.
Co-Authors Abdul Fadlil Abed Neco Achmad S.W.A Nurba Ahmad Ahmad Arifin Ahmad Arifin Arifin Andika Gustiawan Anshar Daud Aries Saifudin Ariya Aritonang Ayni Suwarno Herry Bakri, Asri Ady Bayu Fadlan Rosid Bima Guntara Budi Apriyanto Darmawati Darmawati Delfi Yuliana Tanu Deny Setiawan Destin Mahardika Wijayanti Dharma Fathahillah Diki, Muhammad Asshidiqie Efronius Paduansi Entis Sutrisna Ester, Ria Faizi, Billy Nur Fajar Agung Nugroho Farida Nurlaila Fauzan, Wildan Tino Fazriansyah, Reza Fikri Alfiansyah Fiqih Wijaya Firman Aziz Saputra Fitri Miladiyah, Citra Gama, Fernando Hardiansyah hidayatullah Al Islami Ilham Pratama Ilham, Farizi Irpan Kusyadi Irpan Kusyadi Kusyadi Iwan Giri Waluyo Joko Suwarno Judijanto, Loso Julianus Alfario Junianto, Mochamad Bagoes Satria Khaidar, Ahmad Al Lely Panca Andriyanto Mahir, Shafa Mico Ferdian mohadib mohadib Muhamad Shafly Pratama MUHAMMAD AGIL Muhammad Al Fatih Muhammad Asshidiqie Diki Muhammad Chesta Adabi Putra Muhammad Rivaldi Bachtiar Nadiyanti, Ria Nanang Nanang Nardiono Nardiono Nardiono Nardiono Nardiono, Nardiono Nazar Fadhil Abdullah Nurhasanah Putra, Wahyu Aldi Ramadhan, Syahrul Ghufron Rausan Fikri, Genta Ridwan Rizki Maulana, Rizki Rizky Ramadhani Safitri, Andin Eka Sariadi, Slamet Solihin Solihin Sunardi Suryaningrat Suryaningrat Susanna Dwi Yulianti K Syaeful Machfud Syarif Hidayatullah Testarina Tatiana Hermansyah Teti Desyani Teti Desyani Desyani Widia Novita Sari Willis Puspitasi Sari Yuda Samudra Yulianti Yulianti Yulianti Yulianti Zakaria, Hadi