One of the important indicators that must be monitored to reduce the effects of pollution on health and the environment is air quality. Using the Beijing Multi-Site Air Quality dataset, this study investigated the influence of feature selection on the performance of the Random Forest algorithm in air quality classification. Orange Data Mining is used to process data through the stages of preprocessing, feature selection, model formation, and 10-Fold Cross Validation. The results of the feature selection resulted in five main attributes: PM10, CO, NO₂, SO₂, and O₃. The Random forest algorithm yielded an accuracy of 76.6%, AUC of 0.948, accuracy of 0.764, recognition of 0.766, F1 score of 0.764, and MCC of 0.689. The results show that feature selection has succeeded in simplifying the model by reducing the number of attributes, but it has not been able to improve the performance of Random Forest compared to the use of all attributes.
Copyrights © 2026