Scientific document classification in large-scale digital repositories requires an efficient process because manual categorization is time-consuming and may introduce inconsistencies. This study evaluates the performance of a Random Forest Classifier (RF) using Continuous Bag-of-Words (CBOW) for feature representation and Mutual Information (MI) for feature selection in scientific document classification. The dataset contains 5,560 titles and abstracts of nuclear-related scientific documents distributed across 10 categories. After data cleaning, 5,374 documents were used for experiments with an 80:20 training-test split. Experiments were conducted using CBOW vector dimensions of 100, 200, 300, 400, and 500 under two scenarios: RF+CBOW and RF+CBOW+MI. Performance was evaluated using precision, recall, F1-score, and accuracy. The RF+CBOW scenario achieved the highest accuracy of 73% at a vector dimension of 100, while the scenario with MI reached a maximum accuracy of 71% across several dimensions. Adding MI increased recall to 73% at dimension 100 but did not improve overall performance and reduced all reported macro metrics to 69% at dimension 300. These findings indicate that feature selection does not necessarily improve classification performance; dataset characteristics and representation dimensionality also influence model effectiveness.
Copyrights © 2026