Claim Missing Document
Check
Articles

Found 1 Documents
Search

Comparative Evaluation of Machine Learning Models for Heavy Crude Oil Viscosity Prediction Using Repeated Nested Cross-Validation and Independent Holdout Testing Enggie Hendrawan Saputra; Ilham Ari Elbaith Zaeni; Didik Dwi Prasetya; Azlan Mohd Zain; Welly Antonius; I Made Wirawan
Indonesian Journal of Data and Science Vol. 7 No. 2 (2026): Indonesian Journal of Data and Science
Publisher : yocto brain

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.56705/ijodas.v7i2.455

Abstract

Introduction: Accurate prediction of heavy crude oil viscosity is important for reservoir engineering, production planning, and flow assurance because viscosity strongly affects fluid mobility and transport behavior. This study comparatively evaluates established machine learning models under a rigorous validation protocol rather than proposing a new predictive framework. Method: A published Middle Eastern heavy crude-oil dataset containing 196 development measurements and 47 independent holdout measurements was used. Linear Regression, Support Vector Regression, Random Forest, Gradient Boosting, and the Beggs–Robinson correlation were evaluated using repeated nested cross-validation with five outer folds repeated twice and five inner folds. Preprocessing and hyperparameter selection were embedded within the validation pipeline, while the untouched holdout set was used only for final evaluation. Results and Discussion: Gradient Boosting achieved the best internal performance with R² = 0.99313 and RMSE = 11.41 cP. On the independent holdout set, it achieved R² = 0.99308, RMSE = 8.43 cP, MAE = 6.64 cP, and MAPE = 0.78%, outperforming Random Forest and Support Vector Regression. Residual diagnostics showed no detectable heteroscedasticity, while permutation importance identified temperature and C7+ as the dominant predictors. Conclusion: Gradient Boosting provides highly accurate viscosity predictions within the sampled domain; however, the absence of row-level oil identifiers and external reservoir data limits conclusions regarding oil-disjoint and field-level generalization.