Yaqutina Marjani Santosa
Department of Informatics Engineering, Politeknik Negeri Indramayu

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Benchmarking Random Forest, Support Vector Machine, and XGBoost for Flood Risk Classification Using a Synthetic Dataset: A Case Study of Indramayu Regency Nur Budi Nugraha; Rendi Rendi; Yaqutina Marjani Santosa
Journal of Applied Informatics and Computing Vol. 10 No. 4 (2026): August 2026
Publisher : Politeknik Negeri Batam

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.30871/jaic.v10i4.13708

Abstract

Flood risk assessment plays a critical role in disaster mitigation planning, yet flood classification studies in Indonesia have largely relied on a single machine learning algorithm without systematically evaluating alternative classifiers under identical experimental conditions. This study benchmarks Random Forest (RF), Support Vector Machine (SVM), and XGBoost for three-class flood risk classification (Low, Medium, and High) in Indramayu Regency, Indonesia, using a literature-informed synthetic dataset of 4,500 samples generated from seven flood-related features. To ensure a fair comparison, all models were trained and evaluated under an identical preprocessing, hyperparameter optimization, and validation framework, with Logistic Regression and Decision Tree included as baseline classifiers. Experimental results show that SVM achieved the highest predictive performance with an accuracy of 88.56% and a macro F1-score of 86.83%, followed closely by XGBoost (88.44% accuracy, 86.76% macro F1), while RF obtained 85.00% accuracy and an 83.57% macro F1-score. Statistical significance testing confirmed that SVM and XGBoost significantly outperformed RF, whereas no significant difference was observed between SVM and XGBoost. Feature importance analysis consistently identified rainfall and river distance as the two most influential predictors across all models. Although SVM provided the strongest overall classification performance, RF demonstrated competitive predictive capability with better generalization characteristics than XGBoost, supporting its suitability for operational flood mitigation decision-support systems where model interpretability and robustness are important. Because the benchmark is based on a synthetic dataset, further validation using real observational flood data is recommended before operational deployment.