Muhammad Naufal Rustiawan
Universitas Negeri Semarang

Published : 1 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 1 Documents
Search

Gradient Boosting Models with Optuna Hyperparameter Optimization for Contemporaneous Wind Turbine Active Power Estimation at Esenkoy Wind Farm Muhammad Naufal Rustiawan; Yahya Nur Ifriza
Information Technology Education Journal Vol. 5, No. 3, August (2026)
Publisher : Jurusan Teknik Informatika dan Komputer

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.59562/intec.v5i3.13674

Abstract

Purpose – Wind power estimation is critical for grid stability. This study tests whether Bayesian-tuned gradient boosting, using a leakage-safe pipeline, can estimate turbine power without the Theoretical Power Curve (TPC), benchmarked against LSTM. Design/methods/approach – The study uses Esenkoy SCADA and 2018 MERRA-2 weather data (8,760 hourly observations), split before any transformation; outlier bounds are fitted on training data only, and this dataset needed no imputation. TPC is excluded as a redundant, deterministic function of wind speed. Four gradient boosting models are tuned via Optuna-TPE with 5-fold CV; an LSTM uses identical features and evaluation. Findings – LightGBM has the lowest full-test RMSE (364.17 kW), but CatBoost (385.88 kW) has significantly lower median error (p<0.0001); metrics disagree on the best model, and LightGBM shows a markedly larger train-test gap, consistent with overfitting. CatBoost and GBM outperform LSTM on normal-operation data (p<0.05), while AdaBoost and LightGBM do not. Permutation importance shows wind speed drives over 87% of predictive signal despite differing built-in measures. CatBoost beats LSTM by 18.3% NRMSE on normal-operation data, narrowing to 2.3-9.6% on full data. A seven-seed check confirms these full-test and normal-operation advantages, though intervals overlap. Research implications/limitations – Data cover 2018 at one Turkish site, limiting generalizability; a random split misses temporal shifts, and 50 trials may not fully explore hyperparameter space. Originality/value – This leakage-safe SCADA pipeline shows gradient boosting modestly but significantly outperforms LSTM, with gains depending on abnormal-condition inclusion. Future work should apply temporal cross-validation, test more sites, and tune LSTM more rigorously.