Air quality is becoming an increasingly worrying global issue due to increasing air pollution. The increase in air pollution in the environment encourages the presence of innovative solutions in terms of countermeasures and prevention. This study compares regression algorithms (Random Forest Regression, Linear Regression, SVR, Decision Tree Regression, and KNN Regression) to find the best prediction model. Feature development is carried out using a clustering algorithm (K-means) to produce new features that are able to support the optimization of the model search process and prediction results. Model quality measurement was carried out by applying Mean Squared Error (MSE), Root Mean Squared Error (RMSE), Mean Absolute Error (MAE), and R2 metrics. The results showed that the best model was RFR, which excelled at the R2 approach = 0.75, MSE = 0.008823, RMSE = 0.093931, MAE = 0.060096, and ROC-AUC = 0.951. These findings suggest the effectiveness and quality of prediction models in supporting efforts to develop an early warning system based on ensemble learning dashboards. This research contributes practically to the application of machine learning in air pollution mitigation, as well as supporting the achievement of SDGs 3: Good Health and Well-Being through the provision of a healthier environment.
Copyrights © 2026