Basic principles of public health form a crucial foundation in efforts to protect and improve public health, with a primary focus on disease prevention. In the era of transformation, a science capable of prediction is needed, such as the use of decision tree and random forest algorithms, which have the ability to generate accurate predictions through data processing by forming decision trees. This research utilizes a public dataset on "Peduli Sehat," collected from https://www.kaggle.com/datasets/krismonosadi/peduli-sehat-dataset, consisting of 4,921 records. The research problem is to determine which algorithm provides the best performance based on a dataset of 130 disease symptoms. Therefore, this study aims to explore data mining techniques using a comparative performance approach between the decision tree algorithm and the random forest algorithm in making predictions. Data analysis was conducted using a 70:30 validation split test to identify the best performance for disease prediction. The research results show that the decision tree algorithm achieved a performance of 94.17% accuracy, 95.04% precision, and 94.55% recall, while the random forest algorithm performance was 44.65% accuracy, 45.88% precision, and 46.67% recall. Therefore, this study proves that the decision tree algorithm is more effective for datasets with many symptom features but linear patterns compared to the random forest algorithm. Additionally, the results obtained can provide scientific contributions, particularly in the field of machine learning, and serve as a reference for scientific development in the health sector.
Copyrights © 2026