Indonesian Journal of Data and Science
Vol. 7 No. 2 (2026): Indonesian Journal of Data and Science

Comparative Evaluation of Machine Learning Models for Heavy Crude Oil Viscosity Prediction Using Repeated Nested Cross-Validation and Independent Holdout Testing

Enggie Hendrawan Saputra (Universitas Negeri Malang)
Ilham Ari Elbaith Zaeni (Universitas Negeri Malang)
Didik Dwi Prasetya (Universitas Negeri Malang)
Azlan Mohd Zain (Universiti Teknologi Malaysia)
Welly Antonius (PT Kilang Pertamina Internasional)
I Made Wirawan (Universitas Negeri Malang)



Article Info

Publish Date
31 Jul 2026

Abstract

Accurate prediction of heavy crude oil viscosity supports reservoir engineering, production planning, and flow assurance. This study presents a rigorous comparative evaluation rather than a new machine-learning framework. The published dataset, derived from Kamel et al. and reproduced by Li et al., represents 28 Middle Eastern heavy crude-oil samples through 196 development measurements and 47 independent holdout measurements at 20–80 °C. The supplied workbook contains API gravity, temperature, C1, C2, C3, C4–C6, C7+, and viscosity, with a development viscosity range of 632.88–1267.65 cP; however, it does not include row-level oil or reservoir identifiers. Linear Regression, the Beggs–Robinson correlation, SVR, Random Forest, and Gradient Boosting were evaluated. Imputation, scaling, and hyperparameter selection were embedded in repeated nested cross-validation with five outer folds repeated twice and five inner folds. Final models were evaluated once on the untouched 47-record holdout, with bootstrap confidence intervals, corrected pairwise tests, residual diagnostics, and held-out permutation importance. Gradient Boosting achieved an internal R² of 0.99313 (95% CI: 0.99077–0.99549) and RMSE of 11.41 cP (10.03–12.80). On the independent holdout, it achieved R² = 0.99308 (bootstrap 95% CI: 0.98772–0.99637), RMSE = 8.43 cP, MAE = 6.64 cP, and MAPE = 0.78%. Corrected comparisons showed lower RMSE than Random Forest and SVR. Holdout diagnostics detected no statistically significant heteroscedasticity for Gradient Boosting, and extreme-value sensitivity produced similar performance. Temperature and C7+ were the dominant held-out predictors. These results are encouraging within the sampled domain, but the absent oil identifiers prevent group-disjoint validation, and no external reservoir dataset supports field-level generalization.

Copyrights © 2026






Journal Info

Abbrev

ijodas

Publisher

Subject

Computer Science & IT Decision Sciences, Operations Research & Management Mathematics

Description

IJODAS provides online media to publish scientific articles from research in the field of Data Science, Data Mining, Data Communication, Data Security and Data ...