Cancer remains a major global health challenge, with increasing incidence and mortality rates worldwide. One promising approach in cancer immunotherapy is the identification of epitopes, particularly mutated epitopes that play a critical role in immune recognition. This study aims to classify cancer epitope mutations using machine learning approaches based on sequence-derived physicochemical features. A dataset consisting of 234 samples was used, comprising into tumor and mutated tumor epitopes. Feature extraction was performed using physicochemical properties such as molecular weight, isoelectric point, aliphatic index, aromaticity, and hydrophobicity. Two machine learning models, namely Support Vector Machine and Random Forest were selected due to their proven effectiveness in biological sequence classification tasks and their robustness in handling small-to-medium-sized datasets. The results show that Random Forest achieved the best performance with an accuracy of 83% and a macro average F1-score of 0.70, while consistently outperforming the Support Vector Machine model across all data partition scenarios. However, further analysis revealed that the model exhibits limitations in detecting mutated epitopes, as indicated by a relatively high false-negative rate. This issue is likely due to the use of global sequence-derived features, which may not effectively capture local variations caused by mutations. This study contributes to the field of computational immunology by providing a comparative evaluation of Random Forest and Support Vector Machine for mutation epitope classification using sequence-derived physicochemical features. In addition, the integration of machine learning analysis with structural bioinformatics interpretation offers further biological insight into mutation-associated epitopes and their potential relevance in cancer immunotherapy.
Copyrights © 2026