This study looks at using Small Area Estimation (SAE) to assess per capita spending in subdistricts of Jambi Province. Direct estimators in small areas usually show high variance because of small sample sizes, so model-based methods are needed. The Fay–Herriot model based on EBLUP–FH was used as the primary indirect estimation method, while XGBoost served as a flexible machine learning benchmark. These approaches show fundamentally different methods. EBLUP-FH acts as a model-based estimator with a clear sampling-error framework, while XGBoost functions as a prediction model. Its uncertainty is evaluated through bootstrap resampling, so their standard errors can't be directly compared. Before modeling, data was cleaned and winsorized to lessen outlier impact. The normality assumption for the sampling error in EBLUP-FH was breached, and a logarithmic transformation was used which improved the model components' distribution. XGBoost achieved the smallest bootstrap-based standard error, indicating reduced variability in predictions across resamples rather than greater estimation accuracy. Subdistrict-level analysis revealed that XGBoost tends to produce homogeneous estimates that fail to capture contextual socioeconomic variation, confirming that a small bootstrap standard error does not necessarily indicate accurate small-area estimates.
Copyrights © 2026