Journal of Systems Engineering and Information Technology
Vol. 3 No. 3 (2024)

Quantifying Target Leakage in Rule-Derived Mental Health Screening Labels

Rozi Meri (Politeknik Negeri Padang)
Ikhsan (Politeknik Negeri Padang)
Dian Eka Putra (Politeknik Negeri Padang)
Sofia Yosse (Politeknik Negeri Padang)
Ismael (Politeknik Negeri Padang)



Article Info

Publish Date
01 Dec 2024

Abstract

Publicly hosted tabular datasets increasingly package screening-questionnaire items alongside computed, not clinically observed, risk and severity labels. This creates a risk of target leakage: if a label is a deterministic function of the same features given to a classifier, reported predictive performance measures rule reconstruction, not screening validity. We audit a recent IEEE DataPort dataset (n = 2,005) scoring 27 Likert items across five conditions – ADHD, Autism Spectrum Disorder (ASD), Social Pragmatic Communication Disorder (SPCD), Depression, and Anxiety – each with a binary risk flag and a three-level severity label. The dataset's documentation discloses that every label is generated from a threshold or tertile rule on the summed items of its own condition, yet its usage instructions recommend fitting Random Forest or XGBoost with SHAP directly on the 27 raw items. We verify the documented rule empirically, then compare a trivial single-feature decision stump (the summed score) against a 300-tree Random Forest trained on all 27 items, under identical 5-fold stratified cross-validation, across all ten condition-by-task combinations (5 conditions × {risk, severity}). The single-feature stump matches or exceeds the 27-feature ensemble in every combination, reaching 100% cross-validated accuracy on nine of ten targets. Random Forest feature importance further shows 72–88% of predictive weight concentrated on each condition's own items. These results indicate the pipeline the dataset recommends reconstructs a disclosed arithmetic rule rather than learning a clinically meaningful pattern; any predictive-validity study on this dataset should exclude a condition's own summed items from that condition's feature set.

Copyrights © 2024






Journal Info

Abbrev

JOSEIT

Publisher

Subject

Computer Science & IT

Description

International Journal of Systems Engineering and Information Technology (JOSEIT) is an international journal published by Ikatan Ahli Informatika Indonesia (IAII / Association of Indonesian Informatics Experts). The research article submitted to this online journal will be peer-reviewed. The ...