This study aims to design and implement a clean water quality prediction system based on machine learning using the Random Forest algorithm. The background of this research is the limited public access to fast information regarding water feasibility, while laboratory testing requires significant time and cost. The data used is synthetic data constructed based on the value ranges and threshold limits of water quality in SNI 3553:2015, SNI 3553:2023, and the Ministry of Health Regulation No. 2 of 2023, so that it remains representative and aligned with real conditions. This dataset was created because field data is difficult to obtain completely and in a standardized form, but it still imitates real condition variations according to national standards, with feasibility labels determined based on official regulations. The system development uses the Agile method through the stages of dataset creation, preprocessing, training, evaluation, and application implementation. The Random Forest model is used to classify water into suitable, moderately suitable, and unsuitable categories. The test results show that the model with three physical parameters achieved an accuracy of 96.67%, while the model with ten chemical parameters achieved an accuracy of 100%, confirming that adding more parameters can improve prediction accuracy. This system is expected to help the public and environmental officers in conducting an initial assessment of water quality quickly before laboratory testing and can still be further developed to become more applicable in various regions.
Copyrights © 2026