Oral cancer ranks as the sixth most common cancer worldwide and can be prevented if detected early. The lack of universally available, reliable screening methods has led to the use of machine learning as a promising alternative approach. This study aims to develop an oral cancer classification model using a Random Forest algorithm optimized with GridSearchCV, and to implement it into a Streamlit-based web application. The data is sourced from the Oral Cancer Prediction Dataset on Kaggle, using 1,500 samples from a total of 84,922 data points. Preprocessing included removing irrelevant columns, data cleaning, encoding categorical variables, and normalizing numerical features, resulting in 17 predictor features. The dataset was stratified into training and testing sets at a 70:30 ratio. GridSearchCV optimization yielded the best hyperparameters: n_estimators = 200, max_depth = 10, min_samples_split = 10, and min_samples_leaf = 1. The model achieved an accuracy of 79.91%, a precision of 85.8%, a recall of 68.72%, and an F1-Score of 76.32%. Feature importance analysis shows that the most influential variables, in order, are Treatment Type, Age, Diet, HPV Infection, Oral Hygiene, Alcohol Consumption, Gender, Unexplained Bleeding, Difficulty Swallowing, and Betel Nut Use. The model was successfully integrated into a Streamlit application that displays real-time predictions along with probability values and follow-up recommendations. This system can support early screening for oral cancer in primary healthcare facilities. Further research is recommended using the full dataset and exploring other algorithms to improve performance, particularly the recall value to minimize false negatives.
Copyrights © 2026