The pervasive problem of diagnostic coding inaccuracies significantly impacts the financial integrity and efficiency of Indonesia's National Health Insurance (JKN) system in Type B hospitals. This study aims to assess the efficacy of utilizing large-scale BPJS Health claims data to improve coding accuracy and identify its key determinants. A quantitative, retrospective secondary data analysis was conducted on 150,000 claim records spanning 2020–2024. Big Data analytics employing Random Forest (RF) and Classification and Regression Tree (CART) models successfully detected coding discrepancies, achieving an overall accuracy of 87.2% for primary diagnoses. Statistical analysis indicated that the maturity of the Electronic Medical Record (EMR) system (p<0.01) and staff ICD-10 training (p<0.05) are highly significant determinants. Crucially, the application of this predictive analysis resulted in a 12% reduction in coding errors compared to historical methods. In conclusion, the utilization of BPJS claim Big Data substantially enhances coding accuracy and reliability, confirming the necessity of integrating data-driven technology with simultaneous investments in digital infrastructure and continuous human capacity building for the sustainable quality improvement of the Indonesian health system.
Copyrights © 2025