Purpose – Wind power estimation is critical for grid stability. This study tests whether Bayesian-tuned gradient boosting, using a leakage-safe pipeline, can estimate turbine power without the Theoretical Power Curve (TPC), benchmarked against LSTM. Design/methods/approach – The study uses Esenkoy SCADA and 2018 MERRA-2 weather data (8,760 hourly observations), split before any transformation; outlier bounds are fitted on training data only, and this dataset needed no imputation. TPC is excluded as a redundant, deterministic function of wind speed. Four gradient boosting models are tuned via Optuna-TPE with 5-fold CV; an LSTM uses identical features and evaluation. Findings – LightGBM has the lowest full-test RMSE (364.17 kW), but CatBoost (385.88 kW) has significantly lower median error (p<0.0001); metrics disagree on the best model, and LightGBM shows a markedly larger train-test gap, consistent with overfitting. CatBoost and GBM outperform LSTM on normal-operation data (p<0.05), while AdaBoost and LightGBM do not. Permutation importance shows wind speed drives over 87% of predictive signal despite differing built-in measures. CatBoost beats LSTM by 18.3% NRMSE on normal-operation data, narrowing to 2.3-9.6% on full data. A seven-seed check confirms these full-test and normal-operation advantages, though intervals overlap. Research implications/limitations – Data cover 2018 at one Turkish site, limiting generalizability; a random split misses temporal shifts, and 50 trials may not fully explore hyperparameter space. Originality/value – This leakage-safe SCADA pipeline shows gradient boosting modestly but significantly outperforms LSTM, with gains depending on abnormal-condition inclusion. Future work should apply temporal cross-validation, test more sites, and tune LSTM more rigorously.
Copyrights © 2026