Klasifikasi depresi berbasis machine learning semakin penting untuk mendukung deteksi dini gangguan kesehatan mental. Namun, banyak model modern bersifat black-box sehingga sulit dipahami oleh profesional kesehatan mental. Model berbasis aturan seperti CN2 Rule menawarkan interpretabilitas melalui aturan if–then. CN2 standar merupakan algoritma pembelajaran aturan yang membentuk rule secara iteratif dari seluruh dataset tanpa reduksi instance, sehingga rentan terhadap overfitting dan penurunan performa pada data besar atau tidak seimbang. Penelitian ini mengusulkan kombinasi CN2 Rule dengan modified Partial Instance Reduction (mPIR). Metode mPIR melakukan pengurangan sebagian instance yang memiliki kontribusi rendah terhadap pembentukan aturan, sehingga data menjadi lebih representatif dan efisien dalam proses pembelajaran. Eksperimen dilakukan pada dataset depresi sebanyak 2.556 instances dari Kaggle. Hasil menunjukkan bahwa model usulan meningkatkan kinerja sebesar 7,6% secara konsisten pada akurasi, presisi, dan recall dibandingkan CN2 standar. Peningkatan ini menunjukkan bahwa reduksi instance yang terarah mampu menghasilkan aturan yang lebih general dan mengurangi overfitting. Selain itu, model tetap mempertahankan interpretabilitas, yaitu aturan yang dihasilkan dapat menjelaskan hubungan antara gejala dan tingkat depresi secara jelas. Hal ini penting karena sistem deteksi dini berbasis machine learning sering sulit dipahami akibat kompleksitas model atau kurangnya penjelasan keputusan. Dengan demikian, pendekatan ini tidak hanya meningkatkan performa, tetapi juga mendukung transparansi bagi profesional kesehatan mental dalam proses pengambilan keputusan. Abstract Machine learning-based depression classification has become increasingly important in supporting the early detection of mental health disorders. However, many modern models are black-box in nature, making them difficult to interpret for mental health professionals. Rule-based models such as the CN2 Rule offer interpretability through if–then rules. The standard CN2 algorithm is a rule learning method that constructs rules iteratively from the entire dataset without instance reduction, making it prone to overfitting and performance degradation when applied to large or imbalanced datasets. This study proposes a combination of the CN2 Rule with a modified Partial Instance Reduction (mPIR) technique. The mPIR method reduces instances with low contribution to rule formation, resulting in a more representative and efficient dataset for the learning process. Experiments were conducted on a depression dataset consisting of 2.556 instances obtained from Kaggle. The results show that the proposed model consistently improves performance by 7.6% in terms of accuracy, precision, and recall compared to the standard CN2 Rule. This improvement indicates that targeted instance reduction can produce more generalizable rules and reduce overfitting. Furthermore, the model maintains interpretability, as the generated rules clearly describe the relationship between symptoms and levels of depression. This is important because machine learning-based early detection systems are often difficult to understand due to model complexity or a lack of decision transparency. Therefore, the proposed approach not only enhances performance but also supports transparency for mental health professionals in decision-making processes.
Copyrights © 2026