Tobacco quality is influenced by changes in volatile compounds that occur during the drying process and can be represented through the sensor response patterns on the electronic nose (E-Nose) device. Various machine learning algorithms have been used to classify tobacco based on E-Nose data, but the performance of each algorithm can vary depending on the characteristics of the data used. Therefore, this study aims to analyze and compare the performance of the C4.5 and Random Forest algorithms in classifying tobacco quality based on volatile compound data obtained using the E-Nose. A total of 375 tobacco sample data representing four drying conditions were used in this study. The preprocessing stage was carried out using the Interquartile Range (IQR) method to remove outliers and Moving Average to reduce noise in the MQ-4, MQ-7, and MQ-135 sensor signals. Furthermore, both algorithms were evaluated using stratified 10-fold cross-validation to obtain stable performance estimates. The results indicated that Random Forest performed better than C4.5 in all evaluation metrics. Random Forest achieved an accuracy of 94.44%, a Cohen's Kappa value of 0.9259, an MCC of 0.9262, a balanced accuracy of 0.945, and a cross-entropy log loss of 0.4291. These results indicate that Random Forest is more effective in classifying tobacco quality based on volatile compound data obtained using E-Nose.
Copyrights © 2026