Dian Hafidh Zulfikar
Universitas Islam Negeri Raden Intan Lampung

Published : 2 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 2 Documents
Search

Leakage-Aware and Explainable Machine Learning for Healthcare Claim Fraud Detection Using Imbalanced Medical Insurance Data Dian Hafidh Zulfikar; Ery Setiyawan Jullev Atmadji; Widya Wisanti
International Journal of Artificial Intelligence in Medical Issues Vol. 4 No. 1 (2026): International Journal of Artificial Intelligence in Medical Issues
Publisher : Yocto Brain

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.56705/ijaimi.v4i1.433

Abstract

Healthcare insurance fraud is a critical challenge in health systems because fraudulent claims may cause financial losses, increase administrative burden, and reduce trust in healthcare services. This study proposes an explainable machine learning approach for detecting fraudulent healthcare insurance claims using imbalanced medical claim data. The dataset consisted of 10,000 healthcare insurance claim records with 20 attributes, including patient information, provider characteristics, claim-related financial variables, medical codes, temporal features, and fraud labels. Fraudulent claims represented only 8.29% of the dataset, indicating a clear class imbalance problem. Several machine learning models were evaluated, including Logistic Regression, Decision Tree, Random Forest, Extra Trees, and AdaBoost, under different imbalance handling strategies, namely baseline learning, class weighting, and SMOTE. In addition, two feature scenarios were compared: a full-feature scenario and a leakage-aware scenario that excluded potentially post-decision variables such as claim status and approved amount. The experimental results showed that the best full-feature model was Logistic Regression without additional imbalance handling, achieving an accuracy of 0.9900, precision of 0.9740, recall of 0.9036, F1-score of 0.9375, ROC-AUC of 0.9989, and PR-AUC of 0.9896. The model correctly detected 150 out of 166 fraudulent claims in the test set. However, the best leakage-aware model achieved a lower F1-score of 0.6983, indicating that potentially leaked variables may substantially affect model performance. Feature importance analysis showed that claim amount, approved amount, claim submission delay, claim status, and provider-related variables were among the most influential predictors. These findings demonstrate that explainable machine learning can support healthcare claim fraud detection, but careful attention must be given to class imbalance, data leakage, and operational deployment context
Predicting Cardiovascular Disease Using Machine Learning: A Feature Engineering and Model Comparison Approach Bagus Satrio Waluyo Poetro; Dian Hafidh Zulfikar; I Made Sunia Raharja; Nicodemus Mardanus Setiohardjo
International Journal of Artificial Intelligence in Medical Issues Vol. 3 No. 2 (2025): International Journal of Artificial Intelligence in Medical Issues
Publisher : Yocto Brain

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.56705/ijaimi.v3i2.363

Abstract

Cardiovascular disease (CVD) remains one of the leading causes of mortality globally, emphasizing the need for early detection and effective risk stratification. With the increasing availability of clinical and lifestyle-related health data, machine learning (ML) has become a powerful tool to support data-driven diagnosis and decision-making in healthcare. This study aims to develop and evaluate multiple supervised ML models to predict the presence of cardiovascular disease based on non-invasive features obtained from routine medical checkups. The dataset, comprising 69,301 individual records, includes variables such as age, gender, blood pressure, cholesterol, glucose levels, body measurements, and lifestyle habits. Following comprehensive data cleaning and feature engineering such as the derivation of BMI, Mean Arterial Pressure (MAP), and Pulse Pressure four classifiers were applied: Logistic Regression, Random Forest, Gradient Boosting, and Support Vector Machine (SVM). Model performance was evaluated using metrics including accuracy, precision, recall, F1-score, and ROC-AUC. Among all models tested, the Gradient Boosting Classifier achieved the highest performance, with a ROC-AUC score of 0.8060 and a balanced precision-recall tradeoff, indicating strong discriminatory power. Visualizations such as ROC curves and confusion matrices confirmed the superior capability of Gradient Boosting in differentiating between patients with and without CVD. These findings demonstrate the viability of ML-driven risk assessment models as decision-support tools in clinical settings, potentially aiding in earlier diagnosis and more personalized intervention strategies.