ICD-10 coding errors remain a major challenge in healthcare facilities, affecting the validity of morbidity data, the accuracy of BPJS Kesehatan claims processing, and the quality of health information management. This study aimed to identify the main factors contributing to ICD-10 coding errors and evaluate the effectiveness of machine learning algorithms in detecting such errors. Electronic medical record data were analyzed through data cleaning, preprocessing, descriptive analysis, and machine learning modeling. Three algorithms were applied: Random Forest, Support Vector Machine (SVM), and Neural Network. Model performance was evaluated using accuracy, precision, and sensitivity metrics. The findings revealed an ICD-10 coding error rate of 33.8%, primarily caused by nonspecific diagnoses and insufficient clinical information. Among the tested models, the Neural Network achieved the highest accuracy (72%), followed by SVM (68%) and Random Forest (60%). These results suggest that machine learning techniques can effectively support the early detection of ICD-10 coding errors and enhance the quality of health data management. The adoption of machine learning–based predictive models may improve coding accuracy and facilitate evidence-based decision-making in health information management.
Copyrights © 2026