BPJS Kesehatan Indonesia's national health insurance system, faces significant challenges in managing large-scale claim data, particularly in identifying cost inefficiencies such as abnormal claims, duplication, and misuse. Although machine learning methods have been widely applied in healthcare analytics, limited studies address both large-scale data processing and class imbalance issues in inefficiency detection. This study aims to develop a scalable classification model using XGBoost integrated with Vaex for efficient big data processing. The dataset was obtained from the JKN Healthkathon, consisting of millions of claim records with predefined inefficiency labels. Data preprocessing was performed using Vaex to enable out-of-core computation, followed by model development using XGBoost with various data splitting strategies and imbalance handling techniques, including random oversampling and undersampling. Model performance was evaluated using accuracy, precision, recall, and F1-score. The results indicate that the balanced split with random oversampling achieved the most stable performance, with precision, recall, and F1-score of 0.91. In contrast, stratified splitting without imbalance handling yielded higher accuracy but lower recall. This study contributes by proposing a scalable and efficient framework that integrates XGBoost with Vaex for large-scale healthcare claim analysis and systematically evaluates the impact of imbalance handling strategies on model performance. However, the use of predefined labels with undisclosed generation processes may introduce potential bias in the results.
Copyrights © 2026