The exponential growth of biomedical textual data presents new challenges for efficient classification, particularly under severe class imbalance where minority disease categories are underrepresented. This study proposes a novel framework that integrates Biomedical Bidirectional Encoder Representations from Transformers (BioMedBERT), a domain-specific transformer pretrained on biomedical corpora, with hybrid balancing strategies and a super-ensemble learning approach for medical abstract classification. The preprocessing pipeline includes label normalization, BioMedBERT-based tokenization, and fixed-length sequence alignment. To mitigate imbalance, oversampling, undersampling, and hybrid balancing are applied, and their outputs are combined in a probabilistic super-ensemble. Evaluation was conducted on a medical abstract dataset comprising 14.438 abstracts across five disease categories, using stratified 3-fold cross-validation. Experimental results demonstrate that the super-ensemble consistently outperforms individual balancing strategies, achieving an accuracy of 64.25% and a macro-F1 score of 64.12%. These results indicate improved robustness and sensitivity to minority classes compared to baseline transformer-based methods. The findings highlight that integrating domain-specific pretrained models with advanced resampling and ensemble techniques provides a promising solution for biomedical Natural Language Processing (NLP) applications, offering enhanced reliability for literature retrieval, clinical decision support, and biomedical research.
Copyrights © 2025