River water quality monitoring requires an analytical approach that can classify risk patterns from field observations, especially when independent official water quality labels are unavailable. Method: This study classified water quality risk patterns in the Bengawan Solo River using Naive Bayes and Random Forest based on K-Means cluster labels. The initial dataset consisted of 1,753 observations with ten attributes. After feature selection and missing-value removal, 1,751 observations with seven features were used, namely temperature, pH, electrical conductivity, total dissolved solids, water color, odor, and weather condition. K-Means with K=2 generated higher-risk and lower-risk pattern labels, which were then used as classification targets. Results: Naive Bayes achieved an accuracy of 0.988604, precision of 0.988880, recall of 0.988339, and F1-score of 0.988577. Random Forest achieved an accuracy of 0.980057, precision of 0.980499, recall of 0.979655, and F1-score of 0.980004. Discussion: The findings indicate that cluster-derived labels can be recognized consistently by supervised models, while Random Forest feature importance shows that total dissolved solids and electrical conductivity are the most dominant parameters. The results should be interpreted as cluster label based risk pattern classification, not as official pollution status prediction. The novelty of this study lies in integrating cluster-derived labels, supervised classification, and feature-importance analysis for unlabeled water-quality data, while the findings highlight the practical importance of total dissolved solids and electrical conductivity in preliminary water-quality monitoring.
Copyrights © 2026