Hypertension remains a significant global risk factor for cardiovascular disease and related mortality, necessitating reliable early risk prediction models. Although boosting algorithms have demonstrated strong performance in structured medical data, limited studies have examined their consistency across heterogeneous datasets. This study aims to evaluate the cross-dataset performance and stability of three boosting models, such as XGBoost, LightGBM, and CatBoost, for hypertension prediction under multiple train–test split ratios. Two independent structured datasets were analyzed using 60:40, 70:30, 80:20, and 90:10 splits. To identify the optimal hyperparameters, grid search was performed using repeated stratified 5-fold cross-validation with three repetitions. Model effectiveness was measured using the evaluation metrics of accuracy, precision, recall, F1-score, and AUC. Results show that Dataset 1 gained consistently high predictive performance (accuracy > 0.98; AUC ≈ 1.00), indicating strong and well-separated predictive signals, whereas Dataset 2 demonstrated substantially lower discriminative ability (accuracy ≈ 0.71–0.72; AUC ≈ 0.50), suggesting limited predictive structure. Across both datasets, CatBoost consistently obtained the highest accuracy, particularly at the 90:10 split ratio. These findings demonstrate that dataset characteristics critically determine model effectiveness and that among the evaluated boosting algorithms, CatBoost delivered the strongest overall predictive performance.
Copyrights © 2026